Senior MLOps Engineer
Summary
Senior MLOps Engineer builds and runs scalable LLM inference and training pipelines, deploys large open-weight models, and maintains a cost-aware multi-provider gateway for production ML services.
We are seeking a Senior MLOps Engineer to own end-to-end Machine Learning infrastructure, with a strong focus on high-scale LLM inference serving, distributed fine-tuning pipelines, and multi-provider gateway routing. In this role, you will bridge the gap between ML models and production reliability—deploying self-hosted open-weight models (ranging from ~7B to ~376B parameters), optimizing multi-GPU/distributed topologies, managing cloud cost governance, and treating the internal engineering platform as a product.
Details
Role: Senior MLOps Engineer / ML Infrastructure Engineer
Seniority: Senior / Lead (5+ years in production MLOps/ML Infrastructure)
Allocation: Full-Time (1 FTE)
Rate Cap: Max 145 PLN/h net B2B
Work Model: Remote / Cloud-native across major hyperscalers (AWS, GCP, Azure)
Responsibilities
-
High-Scale Inference Serving
Deploy, scale, and optimize self-hosted open-weight models (~7B to ~376B parameters) using engines such as vLLM, Triton Inference Server, or TGI.
Apply continuous batching, tensor/pipeline parallelism, and quantization strategies (AWQ, GPTQ, FP8) tailored to model size and strict SLA constraints.
-
Multi-Provider Gateway & Cost Governance
Operate and enhance the intelligent routing layer spanning self-hosted models and external APIs (OpenAI, Anthropic, OpenRouter).
Build token accounting, rate limiting, budget controls, and cost/latency/quality-aware routing logic.
-
Distributed Training & Fine-Tuning Infrastructure
Build and maintain automated pipelines for fine-tuning, evaluation, versioning, and continuous delivery (MLflow, SageMaker Pipelines, or Kubeflow).
Manage distributed training workloads using DeepSpeed, FSDP, or Accelerate.
-
Reliability, Observability & Production Ownership
Own production reliability, monitoring, logging, and incident response for ML services (handling GPU OOMs, degraded inference, and latency spikes).
Participate in on-call rotation for the inference serving platform.
-
Evaluation Harnesses & Platform Engineering
Stand up automated evaluation and verification harnesses to catch quality/performance regressions before deployment.
Deliver infrastructure-as-code (IaC) and CI/CD pipelines to provide self-serve tooling for internal engineering teams.
Requirements
Production Experience: 5+ years of hands-on experience in MLOps, ML Infrastructure, or ML Engineering owning end-to-end model lifecycles in production.
On-Call & Incident Handling: Proven experience being on-call for live ML services and resolving real-world incidents (e.g., GPU OOMs, memory leaks, routing failures, cost blowups).
Inference & Quantization Depth: Deep understanding of quantization and parallelism trade-offs under latency, throughput, and hardware cost constraints.
-
Core Tech Stack:
Languages: Advanced Python (C/C++ for performance-sensitive paths is a plus).
Containerization & Orchestration: Docker, Kubernetes (EKS/GKE), Helm, and IaC tools (Terraform/CloudFormation).
Cloud Ecosystems: Deep experience with hyperscaler ML services (AWS SageMaker, EC2 GPU instances, Lambda).
MLOps & Training Frameworks: MLflow, Kubeflow, PyTorch, DeepSpeed, FSDP, or Hugging Face Accelerate.
Nice to Have / Bonus
Prior experience building multi-provider LLM gateways with token billing and cost controls.
Track record of building programmatic LLM evaluation/verification frameworks.