Staff Site Reliability Engineer
Summary
Staff SRE to design and maintain highly available GCP infrastructure for voice AI services handling millions of interactions, focusing on reliability, automation, and observability.
The Opportunity
What You'll Do
- Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
- Architect and automate CI/CD pipelines to ensure rapid, reliable deployments.
- Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues.
- Partner with engineering teams to optimize performance, cost, and reliability of backend services.
- Drive incident response, post-mortem analysis, and long-term remediation efforts.
- Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
- Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
- Lead department wide compliance (PCI, SOC) initiatives.
What You'll Bring
- 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
- Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
- Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
- Deep experience with Kubernetes, container orchestration, and service mesh architectures.
- Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
- Experience designing and managing high-throughput, distributed systems.
- Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs.
- Excellent communication skills and a demonstrated ability to mentor engineers.
Preferred Qualifications
- Experience working in a high-velocity, customer-focused environment.
- Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript).
- Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
- Experience implementing security and compliance best practices in the cloud.