freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineering Lead

Summary

Lead the design, automation, and scaling of Zenith’s cloud infrastructure and SRE practices, ensuring reliability for blockchain validator operations and client environments.

You will own Zenith’s infrastructure systems and lead reliability, performance, and operational functions across internal systems, client environments, and validator infrastructure. You will build cloud environments, establish SRE processes, automate provisioning and deployment, support blockchain protocol operations, and grow an SRE team as the organization scales.

Responsibilities

  • Own production reliability and resilience practices
  • Build, automate, and maintain scalable cloud systems
  • Define and implement incident response, observability, monitoring, alerting, on-call rotation, runbooks, and reliability KPIs
  • Establish versioning, release management, configuration, secrets, environment orchestration, and reproducibility systems
  • Support deployment, scaling, and reliability of client-tailored Zenith Stack environments
  • Support validator operations, network participation, upgrades, and secure infrastructure
  • Automate provisioning, deployment, and system testing
  • Grow and mentor an SRE team
  • Provide hands-on leadership across infrastructure decisions, cloud architecture, and reliability challenges

Requirements

  • Previous experience owning site reliability for a company
  • Experience building or leading an SRE team
  • 7+ years in site reliability engineering, DevOps, platform engineering, or infrastructure-focused software engineering
  • Experience owning reliability, uptime, and operational excellence for distributed systems or mission-critical infrastructure
  • Deep understanding of Linux systems, networking, cloud platforms, containerization, and infrastructure automation
  • Previous expertise with Kubernetes
  • Experience building monitoring, observability, performance, alerting, incident response, and on-call systems
  • Knowledge of Infrastructure-as-Code such as Terraform or Pulumi
  • Knowledge of CI/CD systems
  • Knowledge of secrets management
  • Knowledge of logging and metrics stacks
  • Ability to design and execute scalable, fault-tolerant architectures
  • Ability to coordinate with engineering, product, and protocol teams
  • Based in Europe strongly preferred
  • Experience operating blockchain, validator, L1, L2, consensus, or distributed ledger infrastructure
  • Familiarity with EVM-based blockchain architectures
  • Experience with high-throughput, low-latency, or cryptographic workloads
  • Background supporting enterprise clients, regulated environments, or high-availability financial systems

Benefits

  • Remote work

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available