Lead SRE (Site Reliability Engineer)
Summary
Design, build, and optimize high-availability blockchain data infrastructure at Subsquid, focusing on CI/CD, orchestration, monitoring, and incident management.
You will design, build, and optimize high-availability blockchain data infrastructure. You will own CI/CD, orchestration, infrastructure as code, monitoring, alerting, incident management, and on-call operations. You will improve reliability, observability, automation, fault tolerance, and cost efficiency while collaborating on blockchain integrations.
Responsibilities
- Design, build, and optimize a high-availability blockchain data ingestion pipeline
- Own CI/CD, orchestration, and infrastructure-as-code layers
- Identify and implement tools for running and monitoring blockchain nodes
- Assess infrastructure trade-offs to optimize performance, reliability, and cost
- Build and maintain a public status page and incident management process
- Define and maintain SRE metrics, logging, and alerting
- Improve observability, automation, and fault tolerance
- Contribute to incident response, troubleshooting, and on-call rotations
Requirements
- 3+ years of experience as an SRE, DevOps Engineer, or similar role
- Experience running production services against real SLAs, including on-call, incident response, and post-mortems
- Experience defining and implementing metrics, logging, and alerting
- Proficiency in Kubernetes, Terraform, Prometheus, Grafana, or equivalent monitoring tools
- Strong understanding of distributed systems and streaming data pipelines
- Deep knowledge of cloud infrastructure, including AWS, GCP, or bare metal setups
- Experience balancing performance, reliability, and cost
- Programming skills in Python, Go, Rust, or Bash
- Experience monitoring and running blockchain nodes or working with node providers
- Willingness to learn blockchain node internals, EVM/SVM data, and new chain integration
- Previous Web3 experience preferred
Benefits
- Token incentives
- Fully remote work
- Flexible hours