Senior Site Reliability Engineer
Summary
Own GPU fleet reliability, observability, incident response, and automation to improve system uptime and operational clarity across sites.
You will own work across GPU fleet reliability, observability, incident response, and automation. You will partner with engineering, infrastructure, operations, and customer teams while improving reliability, delivery speed, and operational clarity across sites.
Responsibilities
- Own high-impact work across GPU fleet reliability, observability, incident response, and automation
- Partner with engineering, infrastructure, operations, and customer teams
- Improve reliability, delivery speed, and operational clarity across sites
Requirements
- Strong experience in site reliability engineering or an adjacent technical field
- Clear written communication
- Comfort working with infrastructure teams
- Pragmatic judgment in production environments where reliability matters