Senior Site Reliability Engineer
You will build and scale internal platform offerings covering compute, storage, and networking services to ensure reliability and performance. You will design and implement monitoring, alerting, and incident response systems, collaborate with application software engineers to guide scalable designs, and act as an agent of change to incrementally improve systems as the platform expands globally. You will work with a stack of Python, Java, Terraform, gRPC, Docker, Kubernetes, and Postgres running on AWS, and you will use AI tools daily to reduce toil.
Responsibilities
- Build and scale internal platform offerings including compute, storage, and networking services
- Design and implement monitoring, alerting, and incident response systems
- Collaborate with application software engineers to guide scalable design
- Drive incremental improvements to systems as the company expands globally
Requirements
- Extensive experience with cloud platforms such as AWS, Google Cloud Platform, or Azure, including EC2, S3, RDS, and Lambda
- Experience with Kubernetes or other container orchestration preferred
- Proficiency with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation
- Experience with networking concepts including Container Network Interface (CNI) and network policy implementations
- Experience with proxies and service mesh is a plus
- Strong knowledge of monitoring tools such as Prometheus, Grafana, ELK Stack, or Datadog
- Proficiency in Python with the ability to write efficient, maintainable, and scalable code
- Experience designing, deploying, and maintaining API services with RESTful and/or GraphQL design principles
- Fluency with AI tools, including building agents to reduce toil
- Experience operating CI/CD and its best practices appreciated
a16z