AI DevOps Engineer
Posted
Overview
In this role you will own the deployment, operations, and reliability of the AI-SDLC platform as it moves to a centralized hosted service. You will build and run the infrastructure, CI/CD, monitoring, and observability to keep the platform available, secure, and scalable across teams. You’ll automate deployments, maintain pipelines, and support productionization to meet enterprise SLAs and governance. This is a hands-on role with a focus on reliability, security, and cost-conscious scalability in an enterprise context.
Responsibilities- Own AI-SDLC deployment, operations, and platform reliability (uptime, performance, incident response)
- Manage infrastructure, monitoring, observability, and on-call support
- Build and maintain CI/CD pipelines and automation for platform releases
- Drive the shift to a centralised, hosted AI-SDLC service — scalability, multi-tenancy, cost, and resilience
- CI/CD pipelines experience (GitLab Pipeline / Jenkins & Harness)
- Infrastructure as Code (Terraform)
- Cloud platform experience (AWS ECS/EKS, Lambda, S3, IAM, networking) with container orchestration (Docker + Kubernetes)
- Observability and monitoring (OpenTelemetry, CloudWatch, logging/tracing)
- Scripting/automation (Python and/or Bash)
- Security & compliance operations (secrets management, IAM/least privilege, PCI DSS-aligned controls, auditability)
- Reliability engineering fundamentals (SLIs/SLOs, incident management, cost and capacity management)
- CI/CD pipelines: GitLab, Jenkins, Harness
- Infrastructure as Code: Terraform
- AWS cloud services: ECS/EKS, Lambda, S3, IAM, networking
- Containerization: Docker, Kubernetes
- Observability: OpenTelemetry, CloudWatch
- Scripting: Python, Bash