freehire launches on Product Hunt on 26 August.

Follow →

Senior Infrastructure and Operations Engineer Kubernetes Platform Reliability

Open 24d

You will design, deploy, and maintain production Kubernetes clusters and own their reliability, security, upgrades, and performance. You will build monitoring, logging, alerting, observability, CI/CD, and infrastructure automation systems. You will investigate incidents, strengthen resilience and recovery, optimize Cloudflare, define reliability targets, and improve operational practices.

Responsibilities

  • Design, deploy, and maintain production Kubernetes clusters
  • Own cluster reliability, upgrades, security, and performance
  • Build and operate monitoring, logging, and alerting pipelines
  • Ensure full-stack observability across infrastructure and services
  • Design and maintain reliable CI/CD pipelines
  • Improve deployment strategies, including rollouts, canaries, and rollbacks
  • Automate infrastructure provisioning and configuration
  • Investigate and resolve production incidents
  • Improve system resilience, redundancy, and recovery strategies
  • Define SLOs and SLIs and track reliability targets
  • Optimize and maintain Cloudflare caching, routing, security, and edge behavior
  • Improve operational practices with engineering teams
  • Identify and remove single points of failure

Requirements

  • Senior-level experience operating production infrastructure
  • Deep hands-on Kubernetes expertise, including cluster internals, networking, storage, and security
  • Networking fundamentals, including TCP/IP, routing, DNS, TLS, and load balancing
  • Experience debugging distributed systems and network-related issues
  • Experience optimizing CDN and edge setups, including Cloudflare
  • Experience building monitoring and observability systems
  • Experience with metrics, logs, traces, and alerting pipelines
  • Experience designing reliable CI/CD pipelines
  • Linux fundamentals
  • Experience with infrastructure as code and automation
  • Experience debugging issues across the entire stack
  • Experience handling incidents and conducting postmortems
  • Experience with multi-cluster or multi-region setups
  • Experience with high-throughput or data-heavy systems
  • Experience with Elasticsearch or large-scale data infrastructure
  • Experience with service meshes
  • Experience with cost optimization and capacity planning
  • Experience in regulated or reliability-focused environments

Benefits

  • Meaningful equity upside
  • Remote-first culture
  • Bi-yearly international off-sites
  • Global conference travel and ecosystem engagement
  • Health and wellness benefits

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available