AI DevOps Engineer
Summary
Deploy and maintain an AI platform across cloud GPU, managed services, and on-premise setups using Kubernetes, CI/CD, and cloud-native tools.
- Own the deployment of the AI platform into production across cloud GPU infrastructure, managed services, and future on‑premise or self‑hosted environments, maintaining stability, scalability, and the flexibility to adapt as business needs evolve.
- Build and maintain CI/CD pipelines, monitor system performance, respond to incidents, uphold SLAs, and drive cost and security optimization across all environments.
- Skilled in cloud platforms (e.g., AWS, Azure) and containerization (e.g., Kubernetes), with the ability to deploy across both cloud and on‑premise setups and select the best‑fit solution based on business needs.
- Proven experience operating AI platforms or large‑scale systems, including support for multi‑region and multi‑environment deployments.
Job Responsibilities
- Own the deployment of the AI platform into production across cloud GPU infrastructure, managed services, and future on‑premise or self‑hosted environments, maintaining stability, scalability, and the flexibility to adapt as business needs evolve.
- Build and maintain CI/CD pipelines, monitor system performance, respond to incidents, uphold SLAs, and drive cost and security optimization across all environments.
- Skilled in cloud platforms (e.g., AWS, Azure) and containerization (e.g., Kubernetes), with the ability to deploy across both cloud and on‑premise setups and select the best‑fit solution based on business needs.
- Proven experience operating AI platforms or large‑scale systems, including support for multi‑region and multi‑environment deployments.
Job Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Proven experience in DevOps engineering, with a focus on AI/ML workflows and infrastructure.
- Strong proficiency in cloud platforms such as AWS, Azure, or Google Cloud, including experience with cloud‑native services.
- Hands‑on experience with containerization and orchestration tools like Docker and Kubernetes.
- Proficiency in scripting and programming languages such as Python, Bash, or Go.
- Experience with CI/CD tools such as Jenkins, GitLab CI/CD, or CircleCI.
- Familiarity with AI/ML frameworks and tools such as TensorFlow, PyTorch, or Scikit‑learn.
- Strong understanding of networking, security, and system administration principles.
- Excellent problem‑solving skills and the ability to work collaboratively in a fast‑paced environment.
- Strong communication skills to effectively collaborate with cross‑functional teams.