Infrastructure Engineer and SRE
You will design and operate secure, scalable infrastructure for AI workloads. You will build systems for untrusted code execution, parallel AI-agent runtimes, multi-cloud and customer VPC deployments, observability, and enterprise integrations. You will help define how production AI systems run safely.
Responsibilities
- Design systems for large-scale untrusted code execution and sandboxing
- Build massively parallel AI-agent runtime and scheduling systems
- Design multi-cloud and customer VPC deployment architecture
- Develop highly reliable distributed systems with strict security and data guarantees
- Implement observability across AI workflows and infrastructure
- Build enterprise-grade integration and metadata platforms
- Define how AI systems run safely in production
Requirements
- Strong background in distributed systems, scalability, multi-cloud architecture, and security
- Experience with operating systems, containers, Kubernetes, cloud networking, and event-driven runtimes such as Knative or KEDA
- Experience designing, building, and operating large-scale infrastructure