Site Reliability Engineer
Summary
Builds and maintains reliable cloud infrastructure using Kubernetes and GCP, analyzing incident logs to improve system uptime and performance.
Seeking candidates with strong experience in Python, MTTR, incident response, SRE, Kubernetes, and GCP or cloud infrastructure.
Job description:
Project Outline:
We are looking for a Site Reliability Engineer with experience in incident response (Must have). There will be a focus on the intersection of systems engineering and data science, building the tooling and culture necessary to transform raw incident logs into actionable reliability strategies.
Skill Requirements:
- Engineering Background: 4+ years in SRE, DevOps, or Systems Engineering roles managing production environments at scale.
- Data Proficiency: Strong experience with SQL and data analysis
- Coding Skills: Expertise in one or more programming languages such as Golang, Java, Python, or C++.
- Observability Expertise: Deep understanding of alerting systems, distributed tracing, structured logging, and metrics collection.
- Systems Design: Experience with container orchestration (Kubernetes) and cloud infrastructure (GCP).
- Experience Requirements:
- Statistical Mindset: Experience applying statistical methods (e.g., outlier detection, regression analysis) to system performance data.