Sr. Engineer – Site Reliability Engineering
Summary
Ensures high availability, performance, and resilience of cloud and infrastructure platforms using SRE principles, automation-first practices, observability, and incident management in an on-site role in Chennai.
1. Role Purpose (1–3 lines): Ensures high availability, performance, scalability, and resilience of cloud and infrastructure platforms by applying SRE engineering principles, automation-first practices, observability, and continual reliability improvements across services and platforms.
2. Key Responsibilities:
· Implement SRE frameworks, SLIs/SLOs/SLAs, error budgets, performance engineering, and reliability guardrails across cloud platforms & services.
· Drive automation for provisioning, deployment, configuration management, drift control, patching, recovery, and operations workflows.
· Build observability stack, dashboards, anomaly detection, synthetic tests, runbooks, incident readiness and RCA automation.
· Partner with DevOps, Platform Engineering, Cloud Engineering & application squads to define reliability patterns, capacity planning & scalable workload landing models.
· Manage incident response, major incident coordination, postmortem improvement actions, resiliency testing, fault injection, chaos engineering initiatives.
· Ensure infra security alignment, vulnerability remediation, compliance & secure configuration baselines in cloud infrastructure.