MLOps Platform Engineer (Chennai / Pune)
Summary
Build and maintain scalable MLOps infrastructure for AI/ML services, including Kubernetes clusters, GPU optimization, and LLM inference platforms to support Money Forward’s fintech products.
Responsibilities
As an MLOps platform engineer, you will play a critical role by enabling our team of ML engineers to develop, train and deploy ML projects efficiently using the latest technologies in container orchestration, cloud services, CI/CD pipelines for data collection, model training and monitoring in production
Building and maintaining a scalable infrastructure to execute ML projects, while committed to results and user value
Develop, design, maintain and manage container orchestration using Kubernetes
Design and execute strategies for GPU optimization, prediction servers, data and training pipelines while ensuring efficient use
Design and build inference platforms while ensuring reliability and high performance
Provision and monitor infrastructure resources
Build and maintain ML workflows and pipelines
Deploy and maintain monitoring services for observability
Ensure compliance with security best practices
Manage and expand LLM serving clusters using stacks like vLLM
Requirements
Qualification
Bachelor's degree in Computer Science, engineering or related field
3+ years building core infrastructure for ML projects
Demonstrated background in DevOps, Platform Engineering, SRE, cloud-based infrastructure, or managing production operations
Experience supporting Generative AI, LLM, production-level AI/ML, or platforms focused on data-intensive workloads
Deep understanding of the AI application lifecycle, including MLOps, LLMOps, model monitoring, and deployment strategies
Hands-on experience deploying and providing support for AI services, inference endpoints, and APIs
Experience in managing, designing, implementing and maintaining robust ML infrastructure to support development and inference workloads, ML workflows, training pipelines and versioning
Experience building and scaling machine learning infrastructure
Experience with AWS cloud services
Experience with Kubernetes to deploy and manage containerized applications with high availability and performance
Experience in running and scaling inference clusters
Experience with TerraGrunt or TerraForm, IaC and CI/CD practices
Comfortable taking over legacy projects for operation and maintenance
Proficiency in programming Python
Excellent problem-solving skills and ability to work in a dynamic environment
Effective communication skills to collaborate with technical and nontechnical members
Nice-to-have
Master’s degree in Computer Science, engineering or related field
Production experience operating LLM inference servers such as vLLM (or equivalent serving stacks)
Experience with LLM observability, including the detection of hallucinations, toxicity, and model drift, alongside implementing tracing through OpenTelemetry protocols
Experience with RayServe
Proficiency on KubeFlow and MLFlow for workflows and pipelines
Experience in designing, developing and operating large-scale AI/ML systems
Certifications in AWS(MLS-C01), Kubernetes(CKA) or relevant technologies
Experience with additional cloud services
Contributions to open-source projects
Experience in working to improve model performance, including AI/ML model refinement and fine-tuning
Knowledge of data security standards such as handling personal information, financial/accounting data, PCI DSS, etc., and experience in designing, developing, and operating systems by these requirements.