SRE Lead
Posted Updated
You will own the reliability, scalability, performance, and operational maturity of production infrastructure. You will operate Kubernetes and cloud environments, automate delivery and provisioning, establish observability and reliability standards, lead incident response, strengthen security, and set technical direction while remaining hands-on.
Responsibilities
- Own production infrastructure reliability, availability, scalability, and performance
- Design resilient infrastructure for distributed workloads
- Establish capacity planning, backup, recovery, and disaster-recovery practices
- Own and optimize Kubernetes environments
- Build and maintain infrastructure using Terraform, Helm, Ansible, or equivalent IaC tooling
- Automate provisioning, deployments, configuration, testing, and operational workflows
- Design and improve CI/CD pipelines and release controls
- Build monitoring, logging, tracing, dashboards, and alerting infrastructure
- Define and implement SLIs, SLOs, SLAs, and error budgets
- Lead technical response to critical production incidents
- Establish incident management, escalation, RCA, post-mortem, and on-call practices
- Embed cloud and infrastructure security practices
- Set SRE technical direction, standards, and best practices
- Review infrastructure architecture, mentor engineers, and troubleshoot systems
Requirements
- 7+ years of experience in SRE, DevOps, Platform Engineering, Infrastructure Engineering, or similar roles
- Experience owning production infrastructure at scale
- Hands-on Kubernetes expertise
- Experience with Terraform and Infrastructure as Code
- Experience designing and operating production-grade CI/CD pipelines
- Knowledge of Linux, networking, DNS, load balancing, storage, containers, and distributed systems
- Experience with Prometheus, Grafana, VictoriaMetrics, ELK/EFK, or similar observability stacks
- Scripting or programming skills in Python, Go, Bash, or similar
- Experience with SLIs, SLOs, SLAs, alerting, and error budgets
- Experience leading production incidents, RCA, and post-mortems
- Knowledge of cloud-native and infrastructure security
- Experience making architectural decisions and driving technical standards
Benefits
- Performance-based incentives
- 24 days annual leave plus public holidays
- Health insurance