Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Design and implement observability pipelines using OpenTelemetry, Kubernetes, and Grafana/Prometheus/ELK to standardize logs, metrics, and traces for Meritis’ clients across industries.
Build and maintain large-scale streaming and storage systems using Kafka, Ceph, and Hadoop to handle billions of daily events across hybrid infrastructure.
Build and own a custom AWS environment for a fast-moving AI team deploying production systems in manufacturing plants, using Terraform, Kubernetes, and GitOps to enable rapid iteration and robust observability.
Design and implement scalable, reliable infrastructure and observability systems for a fintech platform, using AWS, Kubernetes, and Terraform to set SLOs, automate monitoring, and lead incident response.
Build and scale Honeycomb’s backend infrastructure using AWS, Kubernetes, and observability tools to ensure reliability for high-volume customers while improving developer experience.
Salary: Additional leave, medical insurance, phone + more About us Born in Christchurch in 2009, OneLaw emerged from a powerful insight: New Zealand law firms were craving a modern, locally supported practice…
Build and maintain Akamai’s global network infrastructure and routing software, define SLOs, and automate deployments to keep the world’s largest edge platform fast and reliable.
Provide 24/7 production support for DBS’s in-house trading and FX applications, troubleshoot incidents, and automate operational tasks using AWS, OpenShift, and DevOps tooling.
Build and maintain highly available, scalable cloud systems using Kubernetes, Azure, and observability tools like Prometheus and Grafana to ensure reliability and reduce operational toil.
Manage and improve the reliability of P&G’s IT systems by implementing monitoring, automation, and incident response to keep services highly available and efficient.
Manage and improve P&G’s IT infrastructure reliability by implementing observability tools, automation, and SLOs to ensure high availability and fast incident response.
Maintain and scale cloud-based financial services infrastructure, automate deployments, and ensure high availability using SRE practices, Kubernetes, and Azure/AWS/GCP.
Ensure reliability and performance of cloud apps on AWS/Azure, automate ops, and lead incident response using Python, PowerShell, and CI/CD tools.
Senior SRE building and scaling global AI hardware infrastructure, automating provisioning with Python, and ensuring 24/7 reliability via Prometheus/Grafana and AI-driven anomaly detection.
Senior Site Reliability Engineer optimizing Akamai’s global network by tuning distributed systems, automating monitoring, and troubleshooting performance issues using Linux, Python, and big data tools.
Lead SRE practices for large-scale platforms, defining SLAs/SLOs, embedding reliability into design, and driving automation and observability in hybrid cloud environments.
Build and maintain CI/CD pipelines and DevOps automation for Mastercard’s BizOps team, ensuring high availability and reliability of payment systems through scripting, monitoring, and incident response.
Build and own the telemetry stack for a SaaS reliability platform powering critical energy infrastructure, defining SLOs, dashboards, and synthetic monitors to ensure real-time system health.
Senior SRE responsible for keeping NVIDIA’s Digital Marketing Services fast and reliable using Akamai CDN, Kubernetes, and AWS, with on-call incident response and AI/ML workloads.
Maintains and improves the reliability of TransUnion’s credit bureau systems by implementing SRE practices, observability tools, and automation to minimize downtime and enhance performance.
We couldn't check your fit for this role — add a CV to your profile to see it next time.