Senior Site Reliability Engineer
Summary
Own and automate global EKS clusters powering authentication services using GitOps, Kubernetes, and AI agents while improving monitoring, CI/CD, and secrets management.
You will own infrastructure powering authentication and authorization services across global production EKS clusters. You will automate deployments with GitOps, build scalable platform foundations, integrate AI agents into engineering workflows, support identity infrastructure, and improve monitoring, alerting, secrets management, and CI/CD processes.
Responsibilities
- Drive deployment automation across production and test EKS clusters using GitOps
- Own infrastructure requirements and coordinate the maintenance backlog
- Collaborate with delivery engineers, product owners, and software developers
- Build and maintain an automated platform foundation across production and development environments
- Integrate and extend AI agents for incident triage and deployment pipelines
- Support infrastructure for authentication and authorization services
Requirements
- Professional experience as an SRE or DevOps Engineer
- Experience with infrastructure automation and configuration management using Helm, Terraform, or Crossplane
- Working knowledge of Kubernetes, Istio, and Flux GitOps workflows
- Experience deploying and operating container workloads on EKS
- Experience with monitoring and alerting using New Relic, Splunk, or PagerDuty
- Experience using AI agents and LLM-based tooling in engineering work
- Familiarity with SOPS and AWS KMS
- Knowledge of CI/CD tools, ideally CircleCI
- Python or Go scripting experience is beneficial
- Background in identity and security infrastructure is beneficial
Benefits
- Generous PTO and holiday schedule
- Parental leave
- Progressive healthcare options
- Retirement programs
- Education reimbursement
- Commuter offset for specific locations
- Employee Resource Groups
- Regular company and team bonding events
- Global volunteering and community initiatives