Senior Site Reliability Engineer
You will lead projects, mentor team members, automate infrastructure, build and maintain AWS production infrastructure, design deployment pipelines for global infrastructure, analyze system behavior and performance, develop observability and runbooks, plan capacity, manage traffic routing and security policies, and participate in on-call operations.
Responsibilities
- Lead projects technically
- Mentor team members
- Automate infrastructure and operations
- Build and maintain AWS production infrastructure
- Design and build pipelines for global infrastructure deployment and management
- Analyze complex system behavior, performance, and application issues
- Develop observability, alerts, and runbooks
- Perform capacity analysis and planning
- Manage traffic routing and security policies
- Participate in an on-call rotation
Requirements
- 6–9 years of experience in software engineering focused on SRE or DevOps
- Experience provisioning large cloud environments with CloudFormation or Terraform
- Experience with Kubernetes and microservice architectures
- Scripting skills in Python, Ruby, Bash, Go, or similar languages
- Experience with Jenkins, GitLab CI/CD, or similar CI/CD automation tools
- Experience with Puppet, Chef, Salt, or similar configuration management tools
- Knowledge of GitOps and ChatOps methodologies
- Knowledge of SLOs, SLIs, error budgets, post-mortems, and incident response
- Experience with secure, scalable, and resilient cloud-native applications
- Experience in high-volume, mission-critical production environments
- Knowledge of New Relic, Grafana, and CloudWatch
Benefits
- Generous PTO and holiday schedule
- Parental leave
- Progressive healthcare options
- Retirement programs
- Education reimbursement
- Commuter offset for specific locations
- Employee Resource Groups
- Regular company and team bonding events
- Global volunteering and community initiatives