Platform Engineer
Summary
Owns cloud infrastructure reliability, scalability, and developer experience for an AI/ML platform startup, focusing on AWS, Kubernetes, and observability to ensure high availability and cost efficiency.
About the Role
We're a fast-growing AI/ML platform startup building infrastructure for training, evaluating, and aligning AI models within reinforcement learning environments. Our engineering team of ~15 includes competitive programming medalists, serial AI startup founders, and researchers published at top venues — and we're looking for a Platform Engineer to own the reliability, scale, performance, and developer experience of our core infrastructure.
This is a backend-architecture-heavy role with high ownership. Your work will directly determine how fast, reliable, and cost-effective our platform is to build on and run.
What You'll Do
Own production uptime, latency, provisioning speed, infrastructure cost, and incident response for core platform services.
Build and maintain AWS infrastructure using Terraform, Kubernetes/EKS, Helm, Docker, EC2, CodeBuild, ECR, S3, IAM, networking, and secrets management.
Design and improve backend and platform systems for scale — capacity planning, autoscaling, queueing, backpressure, cleanup jobs, retries, and rollback paths.
Define and improve dashboards, alerts, logs, traces, SLOs, runbooks, and on-call workflows so failures are detected, debugged, and resolved quickly.
Build reliable CI/CD pipelines, release automation, environment management, and deployment workflows that improve developer productivity and reduce production risk.
Write clean, maintainable production code to automate systems, improve backend services, and create internal developer tooling.
What We're Looking For
Required:
2–4 years of experience owning production cloud infrastructure for a high-availability, user-facing platform, with accountability for uptime, performance, deployment safety, and cost.
Deep hands-on experience with AWS and containerized systems; strong familiarity with Terraform, Kubernetes/EKS, Docker, EC2, load balancers, networking, and secrets management.
Track record of building or operating CI/CD, release automation, observability, alerting, and incident response systems.
Strong backend engineering judgment — able to reason about service architecture, APIs, databases, async systems, queues, scaling limits, and production failure modes.
Ability to write production-quality code to automate infrastructure, improve backend services, and build internal tooling.
Nice to Have:
Experience designing systems for bursty workloads, long-running jobs, sandboxed execution, distributed workers, or high-concurrency services.
Background operating infrastructure for data-heavy, ML/AI, workflow, marketplace, developer-tools, or enterprise platforms.
Demonstrated focus on reducing cloud spend through better architecture, autoscaling, workload placement, caching, or cleanup systems.
Experience building internal platforms or developer tools that improve engineering productivity without hiding complexity.
We prioritize technical aptitude, ownership, and learning potential over years of experience.
Location
San Francisco, CA (on-site): US-based candidates must be located in San Francisco.
Singapore (on-site): Southeast Asia-based candidates must be located in Singapore.
Fully remote (contractor): Candidates based elsewhere — particularly in Europe — may be considered as fully remote independent contractors.
Visa sponsorship is available.
Compensation & Benefits
Salary: $150,000 – $250,000 USD annually (for full-time roles)
Opportunity to have significant ownership and direct impact at an early-stage, well-funded AI infrastructure company.
Work alongside a world-class technical team building foundational infrastructure for AI alignment and post-training data.