freehire launches on Product Hunt on 26 August.

Follow →

MLOps & Devops Engineers

Summary

The MLOps & DevOps Engineer will design and maintain scalable cloud infrastructure and automated CI/CD pipelines on GCP to support AI platforms and RAG architectures. The role focuses on bridging data science and engineering by managing ML lifecycles, monitoring model performance, and ensuring system reliability.

We are seeking a highly skilled DevOps MLOps Engineer to bridge the gap between data science research and AI engineering to ensure smooth SDLC delivery You will be responsible for designing implementing and maintaining the infrastructure and automated pipelines that allow our teams to build deploy and monitor modern AI platforms and RAG pipelines at scale This role requires a deep understanding of cloud infrastructure CI CD practices and the unique challenges associated with the machine learning lifecycle

Key Responsibilities

Infrastructure and Automation

  • Design and manage scalable cloud infrastructure on Google Cloud Platform GCP using Terraform with a focus on GKE clusters firewalls and network policies
  • Develop and maintain full SDLC CI CD pipelines using GitHub Actions integrating Renovate for dependency management Sonar for code quality and Artifactory for binary management
  • Optimize system performance and implement cost-saving measures across cloud environments
  • Build and automate end-to-end ML pipelines on Vertex AI specializing in RAG architectures and automated data ingestion into Qdrant databases
  • Implement and manage evaluation pipelines to measure and improve the performance of LLM-based systems and agentic workflows
  • Establish automated deployment strategies for ML models e g A B testing Canary deployments

Monitoring and Reliability

  • Develop comprehensive monitoring and alerting systems to ensure the health of production models and infrastructure
  • Implement data and model drift detection to maintain the accuracy of deployed models over time
  • Collaborate with security teams to ensure compliance and data privacy throughout the ML lifecycle
  • Integrate and maintain observability tools such as Langfuse OpenTelemetry and Prometheus to enhance system transparency and debugging for ML pipelines and LLM applications
  • Utilize distributed tracing and logging to identify bottlenecks and optimize performance across microservices and agentic workflows

Experience

  • Experience: 3+ years in DevOps, SRE, or MLOps roles.
  • Cloud Platforms: Proficiency in Google Cloud Platform (GCP).
  • Containerization: Advanced knowledge of Docker and Kubernetes (GKE).
  • Automation: Expertise in Python, GitHub Actions, Renovate, Sonar, Artifactory, Argo CD, and Helm charts.
  • Data Tools: Experience with SQL, NoSQL databases, and data orchestration (Airflow).

Preferred Skills

  • Preferred Skills: Experience with Vertex AI, RAG pipelines, Qdrant, and LLM orchestration (LangChain or LlamaIndex).
  • Contributions to open-source DevOps or MLOps projects.
  • Relevant certifications (e.g., AWS Certified DevOps Engineer, CKA).

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available