Platform Engineer (Kubernetes)
Summary
Build and maintain a GitOps-native Kubernetes platform in Go, using ArgoCD, Terraform, and Helm to streamline AI workload deployments and cluster management.
About the Role
Join an early-stage, venture-backed AI infrastructure startup building a next-generation, GitOps-native distributed operating system on top of Kubernetes. As a Platform Engineer, you'll work directly alongside the founding team to design and build the core platform that abstracts Kubernetes complexity while preserving its power — enabling teams to go from bare metal to production AI clusters in days, not months.
This is a high-impact, hands-on role at a small (2–10 person) company with meaningful open-source components. You'll shape architecture decisions, establish engineering patterns, and contribute to a product used by GPU-intensive AI inference and training workloads at scale.
What You'll Do
Build and extend core platform features in Go, including custom Kubernetes operators and controllers.
Design and implement GitOps workflows with ArgoCD to make continuous deployment seamless and automatic.
Develop infrastructure-as-code patterns using Terraform and Helm to provision and manage clusters.
Work on distributed storage solutions using Ceph and WEKA for high-performance, scalable cluster storage.
Create observability and monitoring systems with Prometheus and Grafana to surface cluster health and performance.
Build and optimize container networking with Cilium for network security and observability.
Design and implement federated Kubernetes architectures for multi-cluster management.
Build automation tooling that reduces operational overhead for developers running production workloads.
Contribute to open-source components and establish platform architecture patterns as an early team member.
What We're Looking For
Required:
Strong foundational engineering talent — we prioritize aptitude and hunger to learn over years of specific experience.
Proficiency in Go and hands-on Kubernetes experience (operators, CRDs, controllers).
Experience with distributed storage solutions such as Ceph and/or WEKA.
Experience with ArgoCD and CI/CD pipeline automation.
Comfort with Terraform, Helm, and related infrastructure provisioning tools.
Self-directed, strong communicator, and able to prioritize independently in a fast-moving environment.
Ability to work on-site in San Francisco, CA.
Nice to Have:
Familiarity with Ansible or Kubespray.
Background with service mesh technologies (Istio or Linkerd) or CNI plugins.
Experience with cloud platforms (AWS, GCP, Azure) and their managed Kubernetes offerings.
Contributions to Kubernetes ecosystem tools or other open-source infrastructure projects.
2+ years of software development experience in a platform, SRE, or infrastructure role.
Compensation & Benefits
Salary: $150,000 – $200,000 USD annually
Early-stage equity
Visa sponsorship available
Location
This role is on-site in San Francisco, CA. Candidates must be willing and able to work in-person. Visa sponsorship is available for qualified candidates.