Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Senior SRE responsible for designing observability, maintaining Kubernetes/AWS/Azure infrastructure, resolving production incidents, and automating ops to improve reliability and scalability for a cloud-native SaaS platform.
The Senior Staff Site Reliability Engineer will lead the architecture and development of NVIDIA's enterprise AI runtime platform, focusing on Kubernetes-based systems for deploying and scaling AI applications and databases. The role involves building control-plane services, optimizing GPU scheduling, and establishing standards for high-performance AI inference infrastructure.
About the team Sight Machine is built on the shoulders of a unique, robust and highly scalable Infrastructure as Code model. This enables the creation and operation of customer instances in our ecosystem in a…
Senior SRE focused on improving production resilience, security, and operational excellence for a fintech company’s AWS/Kubernetes-based systems. Designs scalable infrastructure, automates CI/CD pipelines, and defines reliability standards (SLIs/SLOs) while mentoring engineers.
Staff SRE builds and maintains the reliability of Grafana Cloud’s managed databases (Mimir, Loki, Tempo) on Kubernetes across AWS, GCP, and Azure, defining SLOs and leading incident response for high-SLA customers.
Staff SRE to ensure reliability of Grafana Cloud’s managed databases (Mimir, Loki, Tempo, Pyroscope) on AWS/GCP/Azure, owning SLOs, automation, and incident response for high-value customers.
Staff SRE to ensure reliability of Grafana Cloud’s managed databases (Mimir, Loki, Tempo, Pyroscope) on AWS/GCP/Azure, defining SLOs, automating scaling, and leading incident response for high-SLA customers.
Staff SRE engineer maintaining Grafana Cloud’s managed databases (Mimir, Loki, Tempo) on AWS/GCP/Azure, ensuring high reliability and SLOs for top-tier customers through automation, incident response, and cross-team collaboration.
Staff SRE to ensure reliability of Grafana Cloud’s managed databases (Mimir, Loki, Tempo, Pyroscope) on AWS/GCP/Azure, owning SLOs, automation, and incident response for high-SLA customers.
A Staff SRE role focused on eliminating systemic reliability issues and shaping SRE strategy at a large-scale engineering organization, with influence over OKRs, standards, and AI-driven incident management.
Leads a team of SREs to design, implement, and maintain automation solutions for financial infrastructure, ensuring high availability, scalability, and reliability of critical banking systems using cloud-native tools and DevOps practices.
Staff engineer on a team that designs and runs cloud database platforms, focusing on reliability, performance, and operational consistency for a large-scale retail/media company.
Designs and maintains cloud infrastructure and deployment systems using IaC, Kubernetes, and CI/CD tools to ensure scalable, secure, and reliable operations.
Staff SRE builds and scales Kubernetes/EKS and AWS infrastructure for Skydio’s autonomous-drone cloud platform, ensuring 24/7 reliability for critical operations like inspections and disaster response.
Staff SRE designs and implements scalable, reliable systems for Attentive’s AI-powered marketing platform, handling billions of daily customer interactions across SMS, email, and push notifications.
Overview Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for…
Staff Engineer in Site Reliability Engineering for CVS Health’s retail and pharmacy systems, ensuring high availability and performance of critical healthcare infrastructure.
Lead the reliability and scalability of BeyondTrust’s cybersecurity SaaS platform, designing resilient cloud and on-prem systems, CI/CD pipelines, and observability stacks while mentoring engineers and shaping SRE strategy.
Lead SRE initiatives to improve reliability and scalability of NVIDIA's AI-powered enterprise systems, using automation, observability, and AI-assisted incident response.
Leads a team of SREs to ensure reliability and availability of NBCUniversal’s live linear playout systems (e.g., NBC, Telemundo) using cloud, Kubernetes, and broadcast tech like HEVC/HLS.
We couldn't check your fit for this role — add a CV to your profile to see it next time.