Software Engineering Manager, Site Reliability Engineering, Gemini Enterprise Agent Platform, Google Cloud
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.
To learn more: check out our books on Site Reliability Engineering or read a career profile about why a Software Engineer chose to join SRE.
Gemini Enterprise supports platforms for third-party models, including popular open-source models from Model Garden and enterprise-grade models from key partners like Anthropic.
In this role, you will be at the forefront of ensuring the rock-solid reliability and performance of the infrastructure underpinning these third-party model deployments, with a focus on our GKE-based serving stack.
Poland: zł480000 - zł492000 (PLN) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
- Lead a team of site reliability engineers (SREs), guide their professional development, and ensure their success.
- Own end-to-end availability and performance of Gemini Enterprise Agent Platform, and build automation to prevent problem recurrence. Automate response to all non-exceptional service conditions.
- Lead by example, mentor the team, and establish credibility through quality technical execution, including code reviews.
- Extend site reliability engineering (SRE) best practices across the Cloud AI development organization to ensure operational reliability at scale.
- Take part in and manage on-call rotations across continents, using a follow-the-sun model.
Minimum qualifications:
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 8 years of experience with software development in one or more programming languages.
- 3 years of experience managing people or teams.
- 3 years of experience leading projects.
- 3 years of experience designing, analyzing, and troubleshooting distributed systems.
Preferred qualifications:
- Master's degree in Computer Science or Engineering.
- Experience working in an agile software development environment.
- Familiarity with Site Reliability Engineering (SRE) or Production Engineering practices, such as managing system reliability, uptime, and performance, as well as on-call and incident-response.
- Familiarity with high-scale container orchestration fleet operations (e.g., GKE/Kubernetes) or specialized AI/ML compute infrastructure management involving GPUs, TPUs, or model-serving performance optimization.
- Familiarity with security engineering, platform hardening, or data privacy protocols in large-scale environments.