Site Reliability Engineer
Summary
Site Reliability Engineer at Google Cloud in Bengaluru: design, code, and run large-scale, fault-tolerant distributed systems, focusing on reliability, uptime, capacity, and automation of critical enterprise applications. Core work spans software development, algorithms, and large-scale system design on Unix/Linux infrastructure.
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.
- Design, code and execute on projects to improve the reliability posture of critical enterprise applications.
Minimum qualifications:
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 1 year of experience with software development in one or more programming languages.
- 1 year of experience with data structures and algorithms.
Preferred qualifications:
- Experience in an engineering or operations role in large-scale enterprise space.
- Expertise in Unix/Linux systems, IP networking, performance and application issues.
- Expertise in problem solving and analyzing complex enterprise systems.
- Proficiency in navigating enterprise software, deployment and management of workloads.
- Ability to work with multiple global stakeholders.