Lead Site Reliability Engineer
Summary
Lead a team of engineers to design and maintain high-performance, low-latency infrastructure for a proprietary trading firm, focusing on monitoring, automation, and incident response.
You will lead and mentor engineers while contributing hands-on to production reliability initiatives. You will design monitoring, alerting, packet analysis, and automation systems; improve incident and change management; eliminate operational toil; investigate low-level performance issues; collaborate with global engineering and trading teams; and shape infrastructure and tooling strategy.
Responsibilities
- Manage and mentor engineers across teams
- Architect and implement monitoring and alerting systems
- Build real-time packet and flow analysis tooling
- Develop automation frameworks for production infrastructure
- Oversee incident management and change management
- Improve post-incident review processes
- Eliminate operational toil through automation and tooling
- Investigate low-level performance issues
- Optimize software stacks for low latency and high throughput
- Collaborate with engineering, networking, and trading teams
- Shape production tooling, infrastructure scaling, and vendor partnership strategy
Requirements
- Leadership experience managing people across distributed teams
- Experience solving reliability challenges in large-scale production environments
- Strategic thinking and problem-solving maturity
- Strong programming skills in Python, Go, or an equivalent language
Benefits
- Discretionary bonus eligibility
- Medical, dental, and vision insurance
- HSA, FSA, and Dependent Care options
- Employer Paid Group Term Life and AD&D Insurance
- Voluntary Life & AD&D insurance
- Paid vacation plus paid holidays
- Retirement plan with employer match
- Paid parental leave
- Wellness Programs