Senior SRE
Summary
Senior SRE designs, builds, and operates resilient Kubernetes-based cloud infrastructure for AI and developer-tool clients, focusing on automation, observability, and SLO-driven reliability.
- Manage and maintain Kubernetes clusters across cloud platforms, including OpenShift, Amazon EKS, Azure AKS, and Google GKE.
- Implement and manage CI/CD pipelines using tools such as Jenkins, GitHub Actions, Argo CD, or GitLab CI/CD.
- Design and maintain observability stacks with tools including Prometheus, Grafana, Loki, OpenTelemetry, and related technologies. Be part of the team who support open source projects like Prometheus, Thanos, Mimir, CloudNativePG, Istio and more.
- Optimize system performance and resolve production issues. Be part of the on call roster to provide 24x7 coverage for the critical production systems.
- Implement SRE principles, including Service Level Indicators (SLIs) and Service Level Objectives (SLOs), to uphold system reliability.
- Automate infrastructure and operational tasks using programming languages such as Go or Python, and Infrastructure as Code (IaC) tools like Terraform.
- Apply agentic AIto automate the SDLC lifecycle, AIOps and automation.
- Learn about emerging technologies, including AI, GPU Infrastructure
- Contribute to knowledge sharing through technical writing and presentations.
- Bachelor’s degree in Computer Science, Information Technology, or a related field.
- 5+ years of experience in SRE, Platform Engineering, or DevOps Engineer.
- Strong expertise in Kubernetes, cloud-native technologies, on-premise and major cloud platforms (AWS, Azure, GCP).
- Proficiency in programming languages such as Python or Go or Node.js.
- Familiarity with CI/CD tools and modern deployment practices.
- Proficiency in one or more open source observability stacks and Infrastructure as Code (Terraform/Pulumi).
- CKA/CKAD Certified (Brownie points!)
- Excellent problem-solving abilities and communication skills.
- Inclination toward open-source contributions is advantageous.