freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer

Summary

Own production deployment, observability, and incident response for low-latency trading systems; build tooling and SLOs while troubleshooting C++/Linux infrastructure.

You will develop expertise in your assigned technology area, own production deployment and release processes, and improve performance, reliability, and operability. You will build production tooling, define observability metrics and SLOs, lead incident response and post-mortems, manage operational risk, document procedures, and mentor peers.

Responsibilities

  • Develop technical expertise in the assigned product area
  • Own production deployment, configuration, and release processes
  • Build and maintain production tooling
  • Define observability, SLI, SLO, and performance metrics
  • Use metrics and capacity planning to support scalability and uptime
  • Troubleshoot and resolve production incidents
  • Lead incident response, root cause analysis, and post-mortems
  • Align with global SRE teams on architecture and best practices
  • Document processes and procedures
  • Provide mentorship and cross-training
  • Manage operational risk for production changes

Requirements

  • Degree in Computer Science or a related field, or equivalent professional experience
  • At least 5+ years of relevant IT operations experience
  • Expert-level proficiency in C++
  • Linux operating system knowledge
  • Knowledge of network and system configuration
  • Knowledge of kernel internals, scheduling, and performance tuning
  • Networking knowledge including routing, multicast, LLDP, VLANs, and Ethernet
  • Ability to handle shared operational and periodic on-call duties
  • Reliable and predictable availability

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available