Senior AI Platform Engineer
Our client is a technology consulting company providing operational and engineering services to the high-tech sector. They support platform and infrastructure teams across cloud and distributed environments, helping operate scalable, reliable, and production-ready technology platforms.
The Role
Our client is looking for a Senior AI Platform Engineer to operate, maintain, and continuously improve production AI platforms running on Kubernetes across on-premise, AWS, and GCP environments.
The role combines AI platform engineering, MLOps, Kubernetes, Python, observability, and production operations, working with environments similar to AI on EKS and Kubeflow-based machine learning platforms.
You will also play a senior role in improving engineering practices, mentoring team members, and helping shape the platform roadmap.
Key Responsibilities
- Deploy platform releases and configuration changes using GitOps and DevOps practices.
- Monitor AI platform and service health through logs, metrics, monitoring, and observability tools.
- Improve platform reliability through automation, operational tooling, observability, and self-service capabilities.
- Participate in incident response, root cause analysis, and 24/7 operational rotations.
- Investigate and resolve user, platform, integration, and configuration-related issues.
- Promote strong standards across platform security, reliability, and operational engineering.
- Mentor junior engineers in Python fundamentals and help develop their MLOps capabilities.
- Drive the adoption of MLOps best practices across the engineering team.
- Identify gaps in tooling, technical capabilities, and processes required to support production-grade AI systems.
- Contribute to the technical direction and ongoing development of the AI platform.
Qualifications
- 3+ years of experience supporting production AI, ML, or data platforms using technologies such as Ray, Jupyter, AWS SageMaker, Kubeflow, or similar platforms.
- 5+ years of experience across the AI/ML lifecycle, including development, deployment, DevOps, or MLOps.
- 5+ years of hands-on Python experience supporting AI/ML workflows, applications, or data engineering pipelines.
- Strong practical experience with Kubernetes, including managed platforms such as AWS EKS or Google GKE.
- Good understanding of microservices architectures and service communication patterns.
- Strong troubleshooting skills across application crashes, resource contention, service latency, performance, and scaling issues.
- Experience analysing logs, metrics, monitoring systems, and service-level KPIs within production environments.
Nice to Have
- Exposure to additional AI and data platforms such as Flyte, Hugging Face, Vertex AI, LangChain, Claude Code, or other AI agent platforms.
- Hands-on automation or scripting experience using Bash or Python.
- Relevant Kubernetes or cloud certifications such as CKAD or AWS certifications.