Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Senior Site Reliability Engineer designs and maintains scalable, containerized infrastructure for a healthcare AI platform, focusing on Kubernetes, cloud services, and automation to ensure uptime and performance.
Designs, automates, and maintains scalable cloud infrastructure for a healthcare SaaS platform, focusing on Kubernetes, observability, and reliability engineering.
Optimizes compute resources for Elastic’s cloud-hosted and serverless workloads, ensuring seamless scaling across 60+ regions via capacity modeling, autoscaling frameworks, and cross-team collaboration.
The Principal Platform Engineer will manage and optimize compute resources for Elastic's cloud and serverless workloads. The role involves developing capacity models, operating autoscaling frameworks, and collaborating with cross-functional teams to ensure seamless scalability across global cloud regions.
Senior SRE embedded with teams to migrate and onboard Grafana for monitoring, setting up dashboards, alerts, and guiding users on best practices.
Site Reliability Expert driving observability, reliability, and operational excellence across cloud-native environments using Dynatrace (or similar), AWS, Kubernetes, Terraform, and Python.
The Site Reliability Engineer will define SLOs, automate deployment processes, and manage containerized infrastructure for web applications. The role involves collaborating with DevOps teams to ensure system reliability and performance within a secure, high-quality environment.
The Site Reliability Engineer will manage and evolve the internal infrastructure and observability stack at smartclip, focusing on automation, platform stability, and open-source integration. The role involves defining SLOs, improving monitoring systems, and fostering a culture of proactive reliability and security engineering.
Site Reliability Engineer building and operating cloud infrastructure (GCP/AWS, Kubernetes, Terraform) for a French payment platform, driving an AWS-to-GCP migration, IaC optimization, and 24/7 on-call operations.
This role involves working within quantitative trading firms to design infrastructure, write software, and ensure the reliability of critical live trading systems. It is a hybrid position bridging software engineering and production infrastructure.
SRE managing large geo-distributed production environments with Kubernetes (on-prem and AWS), Linux, and data systems like Kafka and Cassandra, automating with Python/Golang.
Manage a team of SREs on Google's Data Cloud, owning end-to-end availability and performance of large-scale distributed systems while building automation and mentoring engineers.
Salary: £70,000 - 90,000 per year Requirements: Deep expertise in Kubernetes and/or OpenShift Experience working in multi-cloud or hybrid cloud environments Strong understanding of SRE principles, including SLAs, SLOs,…
hackajob is collaborating with J.P. Morgan to connect them with exceptional professionals for this role. JOB DESCRIPTION Drive reliability at scale - join a team where your engineering expertise shapes the resilience…
The SRE Ordonnancement will manage, maintain, and evolve IT scheduling tools and infrastructure to ensure the reliability and performance of production batch processes. The role involves working within the production tools team to handle administration, monitoring, and disaster recovery planning.
Senior Site Reliability Engineer at Taboola builds, scales, and maintains high-scale infrastructure across on-premise, public cloud, and AI/ML Kubernetes environments, optimizing performance and reliability for a global ad-tech platform.
Senior SRE at Lodgify ensures reliability, scalability, and observability of cloud/Kubernetes-based travel platform services. Focuses on defining SLIs/SLOs, reducing operational toil, improving incident response, and automating workflows to enhance production excellence.
The Site Reliability Engineer will manage and modernize critical IT infrastructure, focusing on automation, system reliability, and security across hybrid and cloud environments. The role involves designing monitoring solutions, managing CI/CD pipelines, and ensuring compliance with energy sector standards.
This role involves leading a team of engineers to maintain the reliability, availability, and performance of Google's large-scale infrastructure. The manager will oversee on-call rotations, drive automation initiatives, and design software to optimize system scalability and efficiency.
The SRE Lead will manage an engineering team to ensure 99.95% availability for the Moi MTS product. The role involves defining reliability strategies, implementing SLI/SLO/SLA processes, and managing incident response and disaster recovery planning.
We couldn't check your fit for this role — add a CV to your profile to see it next time.