Senior Site Reliability Engineer (AWS Cloud)
Summary
Design and maintain highly reliable, scalable AWS cloud infrastructure and automation pipelines, ensuring low-latency, high-throughput systems while reducing operational toil and single points of failure.
- Architect & Drive the reliability, scalability, and performance of our multi-cloud provisioning platform across all production stacks.
- Architect and implement end-to-end automation pipelines to eliminate manual intervention, actively identifying and reducing technical toil.
- Define, monitor, and improve critical system health indicators (SLIs/SLOs), including latency, throughput, error rates, and capacity usage, making data-driven architectural recommendations.
- Lead Collaboration with product and cross-functional engineering teams to embed reliability and security considerations early into the software development lifecycle (SDLC).
- Own Incident Response Management: Design robust detection mechanisms, triage critical incidents, lead deep Root Cause Analysis (RCA), and implement long-term preventative engineering solutions.
- Simplify Complex Systems: Continually audit platform operations to identify bottlenecks, eliminate single points of failure, and reduce structural complexity.
- 8+ years of experience working within the cloud environment, in roles such as SRE (Site Reliability Engineer) or Cloud Platform/Reliability Engineer
- Strong experience in cloud development and multi-cloud environments, preferably with a strong exposure to AWS cloud
- Knowledge of cloud architecture, scalability, and high-availability design
- Hands-on experience with Kubernetes and container orchestration
- Experience with Terraform and Infrastructure as Code (IaC)
- Experience designing automation and CI/CD pipelines to reduce operational toil
- Strong understanding of SRE principles, SLIs/SLOs, monitoring, and observability
- Proven experience with Incident Management, Root Cause Analysis (RCA), and reliability engineering
- Ability to identify performance bottlenecks, single points of failure, and architectural risks