Lead Site Reliability Engineer
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
Job responsibilities
- Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
- Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
- Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, scalability, and solutions in their applications
- Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
- Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
- Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customers
- Supports the adoption of site reliability engineering best practices within your team
- Production 24*7 support for business-critical applications
- Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience
- Proficient in site reliability engineering (SRE) culture and principles, with experience implementing SRE practices within applications and platforms; strong observability background including white/black-box monitoring, SLO-based alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and similar.
- Proficient in at least one programming language (e.g., Python, Java/Spring Boot,.NET) with strong knowledge of software applications and technical processes within a technical discipline such as cloud, artificial intelligence, Android, or related areas.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
- Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and infrastructure automation using Terraform to build reliable, repeatable delivery pipelines.
- Strong familiarity with containers and orchestration platforms (Docker, Kubernetes, ECS), including deploying, scaling, and operating containerized services in production.
- Proven ability to troubleshoot and resolve common networking issues (DNS, TCP/IP, routing, TLS, load balancing), applying structured debugging to restore service quickly.
- Collaborative, proactive team contributor: communicates clearly and persuasively with minimal supervision, identifies roadblocks early, learns new technologies quickly, and has experience with event streaming platforms such as Kafka.
- Ability to identify new technologies and relevant solutions to ensure design constraints are met by the software team
- Proven track record of initiating and executing ideas that address complex business challenges
- Deep expertise in networking and systems, including TCP/IP, DNS, load balancing, firewalls, and VPN technologies; strong Linux performance tuning and system-level troubleshooting skills
- Certifications a plus: AWS Certified SysOps Administrator or AWS Professional, Certified Kubernetes Administrator (CKA), Terraform Associate (or equivalent)
- Collaborative leader with a proven track record mentoring junior engineers, driving SRE best-practice adoption across teams, and communicating clearly to both technical and non-technical stakeholders (including presentations)
- Experience in handling critical incident and change management – be part of critical incident taskforce call.
- Familiarity of agile practices – preferably, scrum and Kanban
