Lead ML Ops Engineer
Summary
Lead the design and scaling of ML infrastructure and MLOps platforms to deploy, monitor, and improve AI models for a gaming company, using Python, Kubernetes, and cloud services.
We are looking for a Lead ML Platform / MLOps Engineer to build and scale the infrastructure that powers AI across PlaySimple.
- Design and build scalable ML infrastructure supporting training, deployment,monitoring, and governance.
- Develop CI/CD pipelines for Machine Learning workflows.
- Build reliable model serving and inference platforms.
- Develop feature stores, experiment tracking, model registries, and automated deployment pipelines.
- Optimize large-scale data processing pipelines using Apache Spark.
- Build infrastructure supporting LLM applications and Agentic AI workflows.
- Partner with Data Scientists to productionize Machine Learning models.
- Improve platform observability, reliability, scalability, and developer productivity.
- Drive engineering best practices across the ML platform.
Requirements
- 5–8 years of experience building Machine Learning platforms or MLOps infrastructure.
- Demonstrated experience deploying and operating production Machine Learning systems.
- Strong software engineering skills in Python.
- Experience with Apache Spark and distributed computing.
- Strong understanding of modern MLOps practices
Experience with:
- Familiarity with Agentic AI concepts and orchestration platforms such as n8n.
- Experience with AWS, Azure, or GCP.
- Product experience in Gaming, Consumer Internet, E-commerce,Marketplace, or Decision Science platforms.
- Strong communication and collaboration skills.
- An avid gamer with a passion for understanding player behaviour and game mechanics.
- Strong engineering fundamentals.
- Platform-first mindset.
- Excellent problem-solving ability.
- Ownership and accountability.
- Curiosity for modern AI technologies.
- Ability to collaborate across cross-functional teams.
- Comfortable working in a fast-paced product environment.
- Passion for gaming.