Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Design and build large-scale platform infrastructure—Kubernetes clusters, IaC frameworks, SDKs, and platform APIs—for a global ad-tech exchange processing 700B+ daily real-time auctions, using Go/Python on bare-metal and cloud environments.
Leads technical strategy for a global SRE team at Splunk Cloud, owning incident response, incident prevention, and large-scale infrastructure decisions. Mentors engineers, escalates critical production issues, and shapes operational processes for enterprise-scale cloud environments.
The Staff Site Reliability Engineer will manage and automate infrastructure across AWS, containers, and physical servers to ensure the reliability and security of a high-volume payment processing platform. The role involves incident response, infrastructure-as-code development, and maintaining PCI DSS compliance.
The Staff Site Reliability Engineer will serve as the senior technical owner for production issues within the OpenBlue smart building ecosystem, managing infrastructure as code and leading root cause analysis. The role requires deep expertise in Azure, AWS, Kubernetes, and Terraform to ensure the reliability of enterprise SaaS platforms and data pipelines.
The Staff SRE will design, build, and maintain large-scale, distributed systems at Google, focusing on reliability, performance, and automation throughout the service lifecycle. The role involves leading complex technical projects and applying expertise in coding and system architecture to ensure high availability.
The Staff SRE for Google Home will lead reliability, performance, and scalability efforts for Google's smart home infrastructure, focusing on distributed systems, ML model serving, and automation. This role involves architectural oversight and cross-functional collaboration to ensure high-availability services for millions of connected devices.
The Staff Software Engineer for Health SRE will design, scale, and maintain reliable infrastructure for Google Health products, including AI-powered tools and wearable tech. The role involves optimizing distributed systems, automating service lifecycles, and leading cross-functional technical initiatives.
Designs AI-powered incident investigation platforms and scalable backend systems to correlate production data, automate root-cause analysis, and improve software reliability for engineering teams. Focuses on Java-based distributed systems, observability, and AI-driven workflows for incident resolution.
Senior Staff SRE at NVIDIA building and operating large-scale Kubernetes, KubeVirt, and bare-metal compute platforms in Bengaluru, focusing on automation, observability, SLOs, and incident response for global engineering workloads.
Staff SRE responsible for building, scaling, and improving Carrier's cloud-native SaaS platform for building automation systems, with deep focus on AWS infrastructure, Infrastructure as Code, observability, CI/CD, and incident response.
Senior Staff Site Reliability Engineer builds and maintains Ping Identity’s cloud-based identity platform, designing resilient infrastructure, optimizing CI/CD pipelines, and ensuring high availability for enterprise customers.
Staff Site Reliability Engineer at Ping Identity designs, deploys, and maintains cloud-based identity infrastructure using Go, Kubernetes, and GCP, ensuring high availability and security for enterprise clients.
The Staff Site Reliability Engineer will lead infrastructure strategy, focusing on observability, system resilience, and scaling for a global learning platform. The role involves designing SLO/SLI frameworks, mentoring engineers, and optimizing cloud infrastructure using AWS, Kubernetes, and IaC tools.
Staff SRE leading platform reliability and operational excellence at WEX, with strong emphasis on applying AI agents to automate operational workflows, reduce TOIL, and improve system scalability using Kubernetes, observability stacks, and cloud platforms.
Why UKG: At UKG, the work you do matters. The code you ship, the decisions you make, and the care you show a customer all add up to real impact. Today, tens of millions of workers start and end their days with our…
Lead infrastructure transformation at a high-scale digital identity provider, moving from monoliths to microservices while ensuring reliability and guiding engineering teams toward best practices.
The Staff Site Reliability Engineer will lead technical strategy and large-scale projects to ensure the reliability, scalability, and performance of Google's Semantic Understanding Platform. This role involves collaborating with development teams, managing complex distributed systems, and participating in on-call rotations.
Senior SRE ensuring reliability and health of ServiceTitan's cloud platform — designing SLO-based observability, operating Kubernetes at scale across AWS/Azure, automating incident response, and leveraging AI-assisted tooling.
Define and strengthen reliability, scalability, and operational excellence for a healthcare reimbursement platform using cloud infrastructure, observability, and automation.
Staff SRE leads Plaid’s release engineering, designing reliability frameworks (SLOs, error budgets) and progressive delivery systems to enable safe, high-velocity deployments across fintech platforms.
We couldn't check your fit for this role — add a CV to your profile to see it next time.