Point your AI agent at freehire and let it find you a job.

Get the CLI →

Zenith

Open 53d

Site Reliability Engineering Lead

Posted Updated 8 views
Discussion

Summary

Lead the design, automation, and scaling of Zenith’s cloud infrastructure and SRE practices, ensuring reliability for blockchain validator operations and client environments.

Own Zenith’s infrastructure systems and lead reliability, performance, and operational functions across internal systems, client environments, and validator infrastructure. Build cloud environments, establish SRE processes, automate provisioning and deployment, support blockchain protocol operations, and grow an SRE team as the organization scales.

Responsibilities

  • Own production reliability and resilience practices
  • Build, automate, and maintain scalable cloud systems
  • Define and implement incident response, observability, monitoring, alerting, on-call rotation, runbooks, and reliability KPIs
  • Establish versioning, release management, configuration, secrets, environment orchestration, and reproducibility systems
  • Support deployment, scaling, and reliability of client-tailored Zenith Stack environments
  • Support validator operations, network participation, upgrades, and secure infrastructure
  • Automate provisioning, deployment, and system testing
  • Grow and mentor an SRE team
  • Provide hands-on leadership across infrastructure decisions, cloud architecture, and reliability challenges

Requirements

  • Previous experience owning site reliability for a company
  • Experience building or leading an SRE team
  • 7+ years in site reliability engineering, DevOps, platform engineering, or infrastructure-focused software engineering
  • Experience owning reliability, uptime, and operational excellence for distributed systems or mission-critical infrastructure
  • Deep understanding of Linux systems, networking, cloud platforms, containerization, and infrastructure automation
  • Previous expertise with Kubernetes
  • Experience building monitoring, observability, performance, alerting, incident response, and on-call systems
  • Knowledge of Infrastructure-as-Code such as Terraform or Pulumi
  • Knowledge of CI/CD systems
  • Knowledge of secrets management
  • Knowledge of logging and metrics stacks
  • Ability to design and execute scalable, fault-tolerant architectures
  • Ability to coordinate with engineering, product, and protocol teams
  • Based in Europe strongly preferred
  • Experience operating blockchain, validator, L1, L2, consensus, or distributed ledger infrastructure
  • Familiarity with EVM-based blockchain architectures
  • Experience with high-throughput, low-latency, or cryptographic workloads
  • Background supporting enterprise clients, regulated environments, or high-availability financial systems

Benefits

  • Remote work
  • Remote-first team
  • Work-life fit
  • Rest is part of performance

Skills

What Lead SRE jobs ask for — and how much of it you have →
Apply

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available