AI Cloud Engineer
Summary
Design and automate cloud infrastructure for AI workloads on AWS and GCP, deploy models via CI/CD, and monitor performance and security.
Duties and Responsibilities
- Cloud Infrastructure & Automation Design, deploy, and maintain scalable cloud infrastructure for AI workloads on the Cloud and On-Prem Infrastructure.
- Leverage Infrastructure as Code (e.g., Terraform, Helm) to ensure efficient, repeatable, and reliable deployments.
- Continuously optimize cloud resource usage and costs while ensuring infrastructure scalability and reliability.
- AI Integration & Model Deployment
- Work closely with AI teams to build and manage cloud-based environments for AI model training, validation, and deployment.
- Develop and maintain CI/CD pipelines tailored for AI applications, ensuring smooth and automated deployment of models and services.
- Ensure high availability and scalability for AI inference endpoints.
- Performance Monitoring & Optimization
- Implement robust monitoring and alerting solutions (CloudWatch, Stackdriver, Prometheus, Grafana) for AI workloads and cloud infrastructure.
- Proactively identify and address performance bottlenecks, optimizing resources to reduce latency and improve system efficiency.
- Regularly review performance metrics and recommend improvements for operational excellence.
- Security & Compliance
- Implement and manage cloud security best practices, including IAM, encryption, vulnerability scanning, and network security.
- Ensure compliance with industry standards and regulations (e.g., GDPR, HIPAA) specific to AI and data management.
- Collaborate with security teams to conduct regular assessments and integrate security controls into deployment pipelines.
- Collaboration & Communication
- Partner with cross-functional teams including backend engineers, data scientists, and DevOps teams to ensure cohesive AI platform management.
- Communicate effectively across teams to translate complex infrastructure requirements into actionable plans and solutions.
- Actively participate in agile methodologies, contributing to sprint planning, retrospectives, and continuous improvement initiatives.
Deliverables
- Design and implement scalable, secure, and cost-effective cloud infrastructure on AWS and GCP tailored specifically for AI workloads.
- Develop and maintain Infrastructure as Code (IaC), ensuring consistent, repeatable, and automated cloud deployments.
- Create and optimize CI/CD pipelines for efficient AI model deployment, ensuring seamless integration, continuous availability, and scalability.
- Implement comprehensive monitoring and alerting systems, proactively resolving performance bottlenecks and minimizing system downtime.
- Ensure robust security and compliance measures for AI cloud environments, aligning with industry standards and regulations (e.g., GDPR, HIPAA).
Skills
- Solid proficiency in Infrastructure as Code for infrastructure automation.
- Strong hands‑on experience with cloud services in GCP (Compute Engine, Kubernetes Engine, BigQuery, Cloud Storage) and AWS (EC2, S3, Lambda, ECS/EKS).
- Knowledge of containerization (Docker) and orchestration (Kubernetes).
- Familiarity with CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI.
- Understanding of AI model lifecycle, model serving, and data‑intensive infrastructure.
- Exposure to data pipeline technologies (Pub/Sub, Kafka) and familiarity with analytics services.
- Portfolio evidence of successful cloud infrastructure deployments, ideally showcasing Terraform expertise and cloud automation.
- Relevant certifications such as AWS SysOps Administrator – Associate, AWS Solutions Architect – Associate, GCP Associate Cloud Engineer, GCP Professional Cloud Architect or GCP Professional DevOps Engineer, Terraform Certified Associate.
Requirements
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related fields (or equivalent experience).
- Prior experience in cloud infrastructure design, implementation, and management, preferably within AI‑focused environments.
- Experience 2–5 years in cloud engineering, with hands‑on experience in GCP and AWS.
- Demonstrable expertise with infrastructure as code tools, especially Terraform.
- Experience deploying, managing, and optimizing cloud infrastructure for high‑availability applications.
- Strong communication and collaboration skills to effectively interact with cross‑functional teams.
- Demonstrated ability to adapt quickly to changing technologies, continuously learning and integrating new tools and frameworks.
- Relevant certifications such as AWS Certified Cloud Practitioner or GCP Cloud Digital Leader.
Equal Opportunity Employer
Globe is an equal opportunity employer. Our hiring process promotes equal opportunity to all applicants. Any form of discrimination is not tolerated throughout the entire employee lifecycle, including the hiring process.