MLOps Engineer (m/f/d)
- Implement and operate CI/CD pipelines, automated testing and release processes for AI/ML workloads.
- Build and maintain model registry, model serving and AI gateway integrations for LLM APIs and internal applications.
- Configure and maintain observability for model usage, cost, token consumption, latency, reliability and quality signals using tools such as Prometheus, Grafana, logging and alerting platforms.
- Support the transition of workloads from sandbox or PoC environments into production by following defined standards, runbooks and support models.
- Implement reusable technical components for LLM API integration, RAG pipelines, evaluation pipelines and integration with business applications.
- Execute infrastructure-as-code for platform environments across container and cloud infrastructure, including Docker, Kubernetes and Helm-based deployment patterns.
- Maintain runbooks, operating procedures, technical documentation and operational dashboards for platform components.
- Support incident analysis, reliability improvements, cost optimization and lifecycle maintenance for production AI workloads.
- Work with nearshore, system integration or cloud partners on specific implementation tasks as directed by the AI Platform Engineer.
- Collaborate with data engineering, application development, cloud platform and security teams on integration, identity, access and deployment requirements.