freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer

Summary

Own and improve the production environment for a high-frequency trading firm, monitoring and troubleshooting trading systems, building DevOps tooling, and managing incidents.

You will own the production environment and improve its performance, reliability, and operability. You will monitor and troubleshoot trading systems and exchange connectivity, build production operations tooling, coordinate changes and incidents, reconcile trades and position breaks, manage operational risk, document procedures, and mentor other technical operations SREs.

Responsibilities

  • Own the production environment
  • Monitor and troubleshoot large-scale trading systems and exchange connectivity
  • Build and maintain the DevOps toolkit
  • Improve scalability and system performance using firm-wide metrics
  • Analyze and troubleshoot complex system problems
  • Coordinate changes and manage incidents
  • Communicate technology changes with traders
  • Reconcile trades and position breaks
  • Assess operational risk of production changes
  • Define and document processes and procedures
  • Provide mentorship and cross-training

Requirements

  • Degree in Computer Science or a related field, or equivalent professional experience
  • At least 5+ years of relevant IT operations experience
  • At least 3+ years of experience with Python and shell scripting
  • Familiarity with C++
  • Linux operating system knowledge
  • Knowledge of network and system configuration
  • Knowledge of kernel internals, scheduling, and performance tuning
  • Networking knowledge including routing, multicast, LLDP, VLAN tagging, and Ethernet
  • Ability to handle shared operational and periodic on-call duties
  • Reliable and predictable availability

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available