freehire launches on Product Hunt on 26 August.

Follow →

Senior AI Platform Engineer

Summary

Run and harden Kubernetes-based AI infrastructure, own SLOs and observability, automate deployments with Helm/Argo CD, and debug incidents from API edge to model workers.

You will operate and harden Kubernetes-based MaaS production environments across CPU nodes, edge ingress, and regional GPU tiers. You will own SLOs, observability, runbooks, and incident response; improve rollout safety; plan GPU and regional capacity; automate operations with Helm, Argo CD, operators, and scripts; and debug incidents from the public API edge to model workers.

Responsibilities

  • Operate and harden Kubernetes-based MaaS production environments
  • Define and own SLOs, alerting, dashboards, runbooks, and incident response
  • Improve rollout safety with canaries, fallback, health-aware routing, maintenance mode, and rollback
  • Plan capacity for GPU utilization, burst traffic, quotas, rate limits, latency, and customer growth
  • Automate operations with Helm, Argo CD, operators, scripts, and self-healing workflows
  • Debug incidents from the public API edge to model workers
  • Partner with runtime and performance engineers on incident resolution

Requirements

  • 6+ years in SRE, platform engineering, or infrastructure engineering for production cloud services
  • Deep Kubernetes experience
  • Experience with Helm, Argo CD, GitOps, CNI, ingress, secrets, storage, and workload scheduling
  • GPU, AI infrastructure, or HPC workload experience strongly preferred
  • Observability experience with Prometheus, VictoriaMetrics, OpenTelemetry, logs, and traces
  • Go, Python, Bash, Linux networking, and production automation experience
  • Ability to design reliable systems with SLOs and operational ownership

Benefits

  • Welfare benefits

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available