Senior Site Reliability Engineer - Data Infrastructure (San Jose)
Summary
Senior Site Reliability Engineer responsible for automating operational tasks, building monitoring systems, conducting postmortems, and ensuring availability of data center and AI infrastructure. Core technologies include distributed data systems, SLO management, and incident response.
- Automate operational tasks
- Build monitoring and alerting
- Conduct blameless postmortems
- Define and maintain SLOs
- Ensure data center and AI infrastructure availability
- Execute performance tuning
- Implement postmortem follow up actions
- Improve deployment safety
- Lead incident triage and resolution
- Maintain runbooks
- Manage error budgets
- Mentor junior SREs
- Operate distributed data systems
- Optimize resource management
- Perform capacity planning
- Perform change management for production
- Respond to production incidents
- Troubleshoot and resolve production incidents
- Vet production ready deployments
Perks/Benefits:
- 401k match
- Life insurance
- Long-term disability
- Medical, dental, and vision insurance
- Paid Holidays
- Paid parental leave
- Paid personal time
- Paid sick days
- Rotational on call coverage
- Short-term disability
- Wellbeing benefits