Senior AI Platform Engineer
Summary
Run and harden Kubernetes-based AI infrastructure, own SLOs and observability, automate deployments with Helm/Argo CD, and debug incidents from API edge to model workers.
You will operate and harden Kubernetes-based MaaS production environments across CPU nodes, edge ingress, and regional GPU tiers. You will own SLOs, observability, runbooks, and incident response; improve rollout safety; plan GPU and regional capacity; automate operations with Helm, Argo CD, operators, and scripts; and debug incidents from the public API edge to model workers.
Responsibilities
- Operate and harden Kubernetes-based MaaS production environments
- Define and own SLOs, alerting, dashboards, runbooks, and incident response
- Improve rollout safety with canaries, fallback, health-aware routing, maintenance mode, and rollback
- Plan capacity for GPU utilization, burst traffic, quotas, rate limits, latency, and customer growth
- Automate operations with Helm, Argo CD, operators, scripts, and self-healing workflows
- Debug incidents from the public API edge to model workers
- Partner with runtime and performance engineers on incident resolution
Requirements
- 6+ years in SRE, platform engineering, or infrastructure engineering for production cloud services
- Deep Kubernetes experience
- Experience with Helm, Argo CD, GitOps, CNI, ingress, secrets, storage, and workload scheduling
- GPU, AI infrastructure, or HPC workload experience strongly preferred
- Observability experience with Prometheus, VictoriaMetrics, OpenTelemetry, logs, and traces
- Go, Python, Bash, Linux networking, and production automation experience
- Ability to design reliable systems with SLOs and operational ownership
Benefits
- Welfare benefits