Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Leads a team of site reliability engineers to maintain and secure Okta's high-compliance federal cloud environments, ensuring 99.999% availability while meeting FedRAMP, IL4, and IL5 requirements.
Lead a team managing AWS-based infrastructure, Kubernetes, observability, and edge networking to keep Okta’s identity platform running at 99.999% availability.
Senior Site Reliability Engineer at Okta to ensure high availability and scalability of Auth0's identity infrastructure, using cloud platforms, observability tools, and automation.
Senior SRE building and maintaining highly reliable infrastructure for Okta's security SaaS and Snowflake data systems, automating with Terraform, Kubernetes, and Spinnaker while participating in on-call incident response.
Senior SRE building and operating highly reliable, scalable, FedRAMP-compliant cloud services for Okta's Emerging Products Group using Kubernetes, Go, Python, and Terraform on AWS/GCP.
Senior Site Reliability Engineer on Okta's Emerging Products Group in Bengaluru, building and operating large-scale cloud services. Day to day: on-call and incident response, defining SLIs/SLOs, improving observability, and automating away toil with Go, Python, Terraform, Kubernetes, and CI/CD/GitOps.
Designs, deploys, and operates the high-performance parallel/distributed storage layer (WEKA, VAST Data, Ceph, DDN/Lustre) behind Bitdeer's GPU cloud, tuning it for AI training/inference I/O like checkpoint bursts and multi-tenant isolation, and instruments storage telemetry to feed AIOps fault prediction and runbook automation.
Design and lead the global network architecture for AI data centers, including GPU cluster interconnects, DCI, and backbone networks to support large-scale AI workloads.
Design and operate a Kubernetes control plane for GPU workloads, integrating Nvidia operators, topology-aware scheduling, and AIOps-driven remediation to run Bitdeer’s AI-operated GPU cloud.
Senior SRE at Pragmatike running production Kubernetes clusters across bare-metal, virtualized, and on-prem environments for its cloud computing projects. Day to day: Linux administration (Debian/Ubuntu), automation with Ansible/Bash/Python and GitOps, networking design, observability (Prometheus/Grafana, ELK, Loki), incident response, and on-call.
Fully remote Senior Site Reliability Engineer who operates and scales Linux/Kubernetes infrastructure across bare-metal, virtualized and on-prem environments. Day to day involves cluster lifecycle management, network architecture, automation with Ansible/Bash/Python, observability stacks (Prometheus/Grafana, ELK), incident response, and defining SLOs/SLIs.
Senior Infrastructure SRE at a healthcare tech company in Mississauga (hybrid), designing and operating resilient multi-cloud infrastructure across Azure, AWS, GCP, and on-prem. Day-to-day work centers on IaC (Terraform/Pulumi), Kubernetes, service mesh, identity/SSO, observability, SLIs/SLOs, on-call incident response, and automation to reduce toil.
Senior Site Reliability Engineer at Ericsson in Athlone, Ireland, architecting, deploying, and operating private cloud-native infrastructure — Kubernetes, CI/CD pipelines, monitoring, data-center networking, and Ceph storage — to enable Ericsson's 5G and cloud-native development. Core stack: Linux, Go/Python, Docker, Ansible/Terraform.
About Bitdeer: Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing…
Senior Site Reliability Engineer helping build and run Betashares Direct, treating reliability as a software problem. Day to day: building automation and tooling on AWS, improving observability and CI/CD practices, and partnering with software engineers to ship changes safely at scale.
The Principal Platform Engineer will manage and optimize compute resources for Elastic's cloud and serverless workloads. The role involves developing capacity models, operating autoscaling frameworks, and collaborating with cross-functional teams to ensure seamless scalability across global cloud regions.
We couldn't check your fit for this role — add a CV to your profile to see it next time.