AI/ML Platform Engineer (Kubernetes & MLOps)
Summary
Build and scale production AI/ML services on Kubernetes, setting up CI/CD, model serving, monitoring, and MLOps pipelines to ensure reliable, high-performance, and cost-efficient AI deployments.
We are seeking an MLOps/AI Platform Engineer to build, deploy, and scale production AI/ML services on Kubernetes. You will partner with ML engineers and software teams to operationalize models with strong reliability, performance, and cost efficiency using automated pipelines, monitoring, and end-to-end MLOps practices.
Job PurposeOwn the deployment and operational excellence of AI/ML services by establishing robust CI/CD, model serving, observability, and lifecycle management on Kubernetes—optimizing inference performance (latency/throughput), ensuring reliable releases, and reducing infrastructure and operating costs.
Job Duties and Responsibilities- Kubernetes
- Docker
- CI/CD (GitHub Actions/Jenkins)
- Python
- REST/gRPC model serving
- MLflow
- TensorFlow or PyTorch
- Prometheus/Grafana
- Kafka/RabbitMQ
- Monitoring and alerting
- Automated MLOps practices
- Model performance optimization
- Cost optimization
- Kubernetes
- Docker
- Python
- MLflow
- TensorFlow or PyTorch
- REST/gRPC model serving
- CI/CD (GitHub Actions/Jenkins)
- Prometheus/Grafana
- Kafka/RabbitMQ
- AWS/GCP/Azure basics
- Monitoring and incident response
- MLOps/production AI experience