Point your AI agent at freehire and let it find you a job.

Get the CLI →

Telcovas Solutions & Services

NewBe an early applicant

Senior MLOps & DevOps Engineer

Posted
Discussion

Summary

Senior MLOps & DevOps Engineer with 8+ years of experience to architect and scale AI/ML platforms on-prem. Key tasks include deploying ML/LLM models on GPU infrastructure, managing Kubernetes/OpenShift, and building CI/CD pipelines using Python, Docker, and Jenkins.

Role Overview

We are looking for a Senior MLOps + Dev Ops Engineer (8+ years) to architect, build, and scale AI/ML platforms in an on-prem enterprise environment. This role requires end-to-end ownership of ML systems, infrastructure, CI/CD, and production reliability, enabling scalable deployment of machine learning and Gen AI solutions.

Key Responsibilities

  • 1. Platform Architecture & Ownership- Design and own end-to-end ML platform architecture (data - training - deployment - monitoring)- Define and enforce best practices for scalable and secure ML systems- Standardize MLOps + Dev Ops frameworks and processes
  • 2. Model Deployment & Serving- Deploy and manage ML/LLM models on GPU-based on-prem infrastructure- Optimize inference performance (latency, throughput, batching)- Implement model versioning, A/B testing, and rollback strategies
  • 3. CI/CD & Automation- Design and implement CI/CD pipelines for ML models, APIs, and data workflows- Enable automated testing, deployment, and release management
  • 4. Infrastructure & Containerization- Manage Linux-based (RHEL preferred) on-prem infrastructure- Containerize applications using Docker- Deploy and orchestrate workloads using Kubernetes / Open Shift- Operate within restricted or air-gapped environments
  • 5. Data & System Integration- Build pipelines integrating structured databases and high-volume logs/streaming data- Support batch and real-time inference architectures
  • 6. Monitoring, Observability & Reliability- Implement end-to-end observability (model + infra)- Use tools like Prometheus, Grafana, ELK stack- Ensure high availability, SLA adherence, and incident response
  • 7. Gen AI & Advanced ML Systems- Deploy RAG pipelines and vector databases- Manage LLM serving frameworks- Work with agent orchestration frameworks
  • 8. Leadership & Collaboration- Mentor engineers on MLOps and Dev Ops best practices- Collaborate with cross-functional teams- Drive design reviews and production readiness

Required Skills

  • Strong Python and scripting (Bash)
  • Deep understanding of ML lifecycle and productionization
  • Experience deploying ML/LLM systems in production
  • Linux, Docker, Kubernetes/Open Shift
  • CI/CD tools (Jenkins/Git Lab CI)
  • SQL and data pipeline experience

Good to Have

  • GPU optimization knowledge
  • MLflow / Kubeflow
  • Terraform / Ansible
  • Experience in on-prem or restricted environments

Experience

8+ years in MLOps / Dev Ops / Platform Engineering- Proven experience scaling production ML systems

Ideal Candidate

A hands-on platform architect who can operate across ML systems and infrastructure, driving automation, scalability, and reliability.

Skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available