Staff Engineer - Cloud DevOps/ SRE
Summary
Staff Engineer designs and maintains GCP-based Kubernetes infrastructure handling 1B+ daily requests, automates deployments with Terraform/Helm, and ensures reliability for global automotive and media services.
About the role:
We are seeking a Staff DevOps Engineer to manage cloud infrastructure, enable delivery of product features, and design scalable solutions to complex operational challenges across our portfolio. Our team manages multiple products with global reach and demanding scale requirements. Key highlights include: AutoStage, a global API servicing the automotive industry with enhanced radio, audio, and video content for vehicles worldwide; and Music Metadata, which supplies music-related data to industry-leading clients including Pandora, Google, and Microsoft to enrich their customers' search and listening experiences. Built on a modern microservices architecture deployed on Google Cloud Platform, our infrastructure leverages Kubernetes, Helm, Terraform, GitHub Actions, and comprehensive observability tooling. Our systems handle over 1 billion daily requests. You'll collaborate closely with development teams across products in an Agile environment to ensure reliability, performance, and scalability. What you’ll get to do:- Manage multi-region GCP infrastructure at scale (1B+ requests/day)
- Diagnose high- and low-level performance issues in a highly distributed environment
- Manage and develop GKE Kubernetes clusters
- Migrate other projects hosted on-prem or on different solutions to our stack
- Use and improve existing monitoring systems
- Identify time-consuming tasks and automate them
- Maintain standards of security, reliability, performance and quality
- Research best practices and technologies and implement them into your day-to-day work
- Participate in follow-the-sun on-call rotation, only after you are up-to-speed with our products.
- 5+ years of experience in DevOps, SRE, or infrastructure engineering roles
- 3+ years of experience with Linux and Networking in production environments
- 2+ years of experience with running Kubernetes in a production environment
- Experience managing infrastructure with a major cloud provider, GCP most preferably
- Experience with time-series-based monitoring systems (Prometheus, Datadog,
- Experience with coding in Terraform
- Experience managing Kubernetes configurations with Helm
- Experience writing automations in Python or similar scripting language
- Experience with CI/CD platforms (GitHub Actions, Jenkins, GitLab CI)
- Experience with database operations (PostgreSQL, MySQL, BigQuery)
- Container registries and image security scanning
- Experience with terragrunt
- Experience with incident management and SRE practices (SLIs/SLOs, blameless post-mortems)
- Image building tools (Packer, Cloud Build)
- The ability to propose, design and develop solutions that scale
- Keen troubleshooting skills
- Excellent written and oral communication skills
- Expert problem-solving skills
- Strong documentation practices
- Competitive compensation (salary, equity and bonuses) and comprehensive benefits designed to foster work-life balance, care for your health, protect your finances and help you save and invest for the future.
- Generous paid time away from work, including flexible time off, holidays and sick time, health and wellness initiatives, and a charitable match program to help you give back to your community.
- Great perks, which vary by location and can be site-specific: employee discounts, transportation reimbursements, subsidized cafes and fitness facilities.
- A flexible, hybrid work environment combining the best of in-office collaboration and community-building along with the benefits of working from home.