Azure DevOps / Platform Engineer with GPU 8418 M
Summary
Designs and operates GPU-enabled Kubernetes platforms on Azure and hybrid infrastructure to support AI/ML workloads, using Infrastructure as Code and GitOps practices.
Azure DevOps / Platform Engineer with GPU 8418 M
Join Akvelon — build products used by millions!
Akvelon is an IT company with 20+ years of experience and 1,200+ engineers across 15+ locations worldwide.
We work with both well-known global tech companies, including Microsoft, Facebook, Airbnb, Dropbox, and Pinterest, and with growing startups.
Our teams are involved in different types of engineering projects, from cloud solutions and AI/ML systems to big data, web, and mobile applications.
Since we are remote-first, our engineers work in distributed teams with flexible hours. We value ownership, clear communication, and the ability to take responsibility for your part of the work.
About the role
We are building a cloud-native platform that enables research teams to run GPU-intensive AI/ML workloads efficiently across Azure and hybrid infrastructure.
This role focuses on designing, automating, and operating production-grade GPU infrastructure, Kubernetes platforms, and hybrid cloud environments. You'll build scalable GPU clusters, automate infrastructure provisioning, and improve the developer experience through Infrastructure as Code, GitOps, and platform engineering best practices.
Requirements
- Strong hands-on experience with Azure and hybrid infrastructure, including Azure Arc, Azure Networking, IAM Bicep
- Experience with GPU infrastructure on Kubernetes (Volcano Scheduler, NVIDIA GPU Operator, GPU scheduling)
- Production experience with Infrastructure automation and GitOps (Ansible, ArgoCD, Helm, Terraform or Bicep)
- Experience with k8s platform engineering (control plane, cluster administration, production troubleshooting, Linux)
- Understanding of Bare-metal and infrastructure operations (bring-up, SSH, storage and node lifecycle)
- Experience building AI/ML platforms
- Experience with Slurm
- Experience with GitHub Actions or Azure DevOps
- Python or Bash scripting
- Experience with Dev Containers
Responsibilities
- Build, configure, and operate production Kubernetes clusters for GPU workloads
- Manage Azure and hybrid infrastructure, including Azure Arc, networking, RBAC, and Azure VM integration
- Automate infrastructure using Bicep, Ansible, Helm, and Argo CD
- Implement GPU scheduling, platform security, patching, scaling, and infrastructure automation
- Improve observability through monitoring, Grafana dashboards, telemetry, and alerting
- Support storage, backup, disaster recovery, and CI/CD pipelines
- Troubleshoot production issues and continuously improve platform reliability and performance
- Be available for 2 syncs per week until 7 PM CET
- Ability to accommodate occasional ad hoc meetings with UK- and PDT-based full-time team members, with 1–2 hours of working time overlap as needed