Site Reliability Engineer owning reliability of xAI's Memphis/Southaven data center campus: designing fleet-scale monitoring and alerting, commanding SEV incidents, running blameless postmortems, building playbooks and SLOs/error budgets, and joining on-call. Core skills are observability design, incident leadership, and Python/Bash scripting plus a systems language.
Sign in to see your match