AI/ML Platform Cloud Infrastructure Engineers
Summary
Build and maintain a scalable AI/ML platform on Google Cloud using Terraform, focusing on secure, automated infrastructure for AI workloads.
Start Date: 03.08.2026
End Date: 30.07.2027
Long-term: Yes
Work Model: Remote
Summary
The primary goal of this role is to develop and maintain a scalable AI/ML platform infrastructure on Google Cloud Platform (GCP) to support AI/ML workloads, ensuring security, automation, and scalability.
Main Responsibilities
Design, implement, and maintain scalable infrastructure using Terraform (HCL).
Develop reusable infrastructure modules with dynamic blocks and lifecycle rules.
Integrate IaC standards with tools like Terragrunt, tflint, and tfsec.
Manage infrastructure across multiple GCP projects and enforce governance policies.
Design self-hosted GitHub Actions runner infrastructure on GCP.
Architect advanced GCP networking solutions including VPC design and hybrid connectivity.
Provision and manage GPU infrastructure for AI/ML workloads.
Operate autonomously with a security-first mindset, ensuring effective communication and planning.
Key Requirements
Deep expertise in Terraform (HCL) for large-scale environments.
Strong experience with Google Cloud Platform (GCP) in multi-project settings.
Proven design experience with self-hosted CI/CD platforms.
Advanced knowledge of cloud networking.
Experience with CI/CD automation integration.
Nice to Have
Experience with GPU types like A100, H100, L4, and T4.
Understanding of Slurm clusters and Kubernetes orchestration.
Details
Location: Remote
Team Structure: Cross-functional team
Tools: Terraform, GCP, GitHub Actions
Start Date: 03.08.2026
End Date: 30.07.2027
Long-term: Yes
Work Model: Remote