Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Job Location MANILA NET PARK OFFICE Job Description Overview of the job As the Senior SRE Lead in the Warehousing IT Operations – Incident Response Team, you will be responsible for leading incident response efforts,…
Leads incident response and ensures system reliability for P&G’s warehousing IT operations, troubleshooting critical issues, optimizing infrastructure, and collaborating cross-functionally to maintain high availability and performance.
Staff DB SRE building and operating large-scale database infrastructure (MySQL, MSSQL, Oracle, vector/graph DBs) for NVIDIA's enterprise AI platforms, with a focus on automation, high availability, and CI/CD-integrated database lifecycle management using Python, Go, and Kubernetes.
The Senior Site Reliability Engineer will manage and scale AI hardware infrastructure, focusing on automation, monitoring, and incident resolution. The role involves developing Python-based tooling and maintaining high availability for distributed systems within a global cloud platform.
SRE Compute role at S3NS (Thales/Google Cloud JV) operating and maintaining GCP-equivalent compute platforms (GCE, GKE, Vertex AI) with 24/7 on-call, incident management, SLI/SLO monitoring, and infrastructure automation in a SecNumCloud-certified sovereign cloud environment.
The application window is expected to close on: Job posting may be removed earlier if the position is filled or if a sufficient number of applications are received . Meet the Team We are CloudOps — the team that keeps…
The SRE Engineer will design, code, and deploy reliable product features while championing automation and observability. The role involves working with CI/CD pipelines, managing incident resolution, and ensuring system scalability using Python and various DevOps tools.
Senior SRE designing and optimizing cloud infrastructure (IaC with Terraform on AWS) for video game production tools and applications within the R&D team.
The SRE Technical Lead will ensure the operational excellence of critical production systems by driving automation, optimizing performance, and leading technical incident response. The role requires deep expertise in cloud platforms, Kubernetes, and observability tools to build scalable, self-healing systems.
Senior Associate SRE ensuring reliability, availability, and performance of systems in Hyderabad — monitoring, incident response, automation, and deployment support using cloud platforms, scripting, and tools like Prometheus, Grafana, and Terraform.
As an SRE Engineer at Goldman Sachs, you will ensure the availability and reliability of critical platform services by defining SLOs, managing incidents, and automating systems. The role involves using languages like Go, Python, or Java to build scalable, fault-tolerant systems and operate observability platforms.
This hybrid role combines technical support and SRE responsibilities to maintain and scale an AI security SaaS platform. The engineer will monitor system performance, troubleshoot customer incidents, and develop automation to improve platform reliability.
The Software Engineering Manager leads development teams and application support functions within the Site Reliability Engineering Center, overseeing production services and incident response. The role requires managing projects, coaching staff, and providing after-hours leadership for critical system events.
Senior SRE leading the build and transformation of a cloud-native technology stack across the full SDLC, working with AWS, Terraform, GitLab CI/CD, Kubernetes, and Python at a state pension fund.
Site Reliability Engineer owning system reliability, availability, and performance on AWS, designing multi-AZ/multi-region architectures, leading incident response, and standardizing observability using IaC (Terraform/CloudFormation/CDK) and monitoring tools like Dynatrace and OpenTelemetry.
Site Reliability Engineer I focused on production operations (AppOps) for NCR Atleos Financial Services products, maintaining availability, performance, and security using cloud platforms, CI/CD, IaC, and observability tooling.
Sobre a vaga: No time de Engenharia de Malha Batch nosso objetivo é gerar autonomia para BUs e áreas de negócio, promovendo melhores práticas, eficiência e excelência operacional. Portanto, automatizar para agilizar e…
Designs and maintains cloud-based systems for reliability, scalability, and fault tolerance; defines SLIs/SLOs, automates deployments, and ensures high availability through observability and incident response.
About Lunit From screening to research and development, Lunit is transforming how the world advances cancer detection — boldly, intelligently, and with humanity at the core. We build trusted partnerships with…
Senior SRE to maintain and scale a high-throughput, event-driven ticketing platform on AWS, ensuring reliability during live events and optimizing CI/CD and observability.
We couldn't check your fit for this role — add a CV to your profile to see it next time.