Site Reliability Engineer
You will operate the infrastructure supporting research compute, container registries, and dashboards. You will improve compute access and resource visibility, enable autoscaling, manage access controls, create reproducible deployments, and automate operational processes.
Responsibilities
- Ensure efficient access to compute resources
- Provide visibility into resource utilization and cluster health
- Enable automatic scaling of compute resources
- Manage access to infrastructure resources
- Drive deterministic deployments and reproducible research environments
- Automate operational processes
- Operate infrastructure using Ansible, Kubernetes, Docker, Tailscale, Python, Grafana, Prometheus, and Talos Linux
Requirements
- Take accountability for cluster health and capacity
- Understand interactions among schedulers, containers, networking, storage, and hardware
- Design systems with predictable failure modes
- Apply observability, reproducibility, and clear operational boundaries
- Support experimental research workloads pragmatically