Point your AI agent at freehire and let it find you a job.

Get the CLI →

Astera Institute

NewBe an early applicant

Site Reliability Engineer

Posted 1 view
Discussion

You will operate the infrastructure supporting research compute, container registries, and dashboards. You will improve compute access and resource visibility, enable autoscaling, manage access controls, create reproducible deployments, and automate operational processes.

Responsibilities

  • Ensure efficient access to compute resources
  • Provide visibility into resource utilization and cluster health
  • Enable automatic scaling of compute resources
  • Manage access to infrastructure resources
  • Drive deterministic deployments and reproducible research environments
  • Automate operational processes
  • Operate infrastructure using Ansible, Kubernetes, Docker, Tailscale, Python, Grafana, Prometheus, and Talos Linux

Requirements

  • Take accountability for cluster health and capacity
  • Understand interactions among schedulers, containers, networking, storage, and hardware
  • Design systems with predictable failure modes
  • Apply observability, reproducibility, and clear operational boundaries
  • Support experimental research workloads pragmatically

Skills

Apply

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available