Cloud Infrastructure Engineer
You will design, deploy, and continuously improve the infrastructure powering a blockchain developer platform serving 100+ chains, billions of daily requests, and over $150B in annual transactions. You'll provide the infrastructure, tooling, and expertise needed to allow engineers to ship, scale, and operate high-quality products in a fast, safe, and cost-efficient manner.
Responsibilities
- Architect and operate scalable, self-healing infrastructure leveraging Kubernetes, Terraform, and cloud-native tools across multi-region deployments.
- Drive AI enablement across engineering, ensuring repos, tooling, and workflows are optimized for agentic development with tools like Claude Code, Cursor, and Codex.
- Build AI-powered infrastructure tooling and automation such as automated K8s upgrades, IaC plan analysis, cost optimization advisors, MCP servers, and n8n workflows.
- Build and maintain internal developer platform capabilities for self-service deployments, observability, and reliability.
- Develop observability frameworks using Prometheus and Grafana for metrics, dashboards, and alerting.
- Lead incident management with blameless post-mortems and define and enforce SLIs, SLOs, and error budgets across services.
- Design and manage multi-cloud, multi-region network architecture including VPC design, IPAM, DNS, cross-cloud connectivity, security groups, and edge-proxy/istio gateway configuration.
- Collaborate with security teams to embed compliance into infrastructure, including IaC scanning and runtime protection.
- Provide technical leadership and mentorship to elevate the team's operational capabilities.
Requirements
- 5+ years as an Infrastructure Engineer focused on reliability (SRE, Production Engineer, Platform Engineer).
- Experience driving company-wide reliability efforts, including SLO frameworks and error budget policies.
- Strong proficiency with observability stacks: OpenTelemetry, Prometheus/Grafana.
- Deep experience with cloud infrastructure (AWS/GCP), Kubernetes, and multi-region architectures.
- Skilled with Terraform, Helm, and GitOps workflows (e.g., ArgoCD) with an automation-first mindset.
- Experience leveraging agentic development tools (Claude Code, Cursor, Codex) and workflow automation (n8n) is a strong plus.
- Solid networking fundamentals — VPC design, DNS, IPAM, security groups, cross-cloud connectivity, and service mesh (e.g., Istio) experience is a plus.
- Strong cross-functional communicator across SRE, security, and product engineering.
- Blockchain infrastructure, distributed systems, or high-throughput RPC experience — not required but a plus.
Benefits
- Medical, Dental, & Vision
- Gym Reimbursement
- Home Office Build-out Budget
- In-Office Group Meals
- Wellbeing & Mental Health Perks
- Learning & Development Stipend
- Company Sponsored Conferences & Events
- HSA and FSA Plans
- Fertility Benefits
- 401k
- Unlimited flexible time off
Y Combinator
a16z