freehire launches on Product Hunt on 26 August.

Follow →

Senior Lead Site Reliability Engineer, Electronic Colo Trading

Open 23d reposted 2× · 2 open copies

Summary

Lead reliability engineering for low-latency Linux platforms powering JPMorgan’s electronic colo trading systems, focusing on automation, observability, and incident response in high-availability data centers.

Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.

As a Principal Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms, Electronic Trading Services, you work with your fellow stakeholders to define non-functional requirements (NFRs) and availability targets for the services in your application and product lines. You will be responsible for the reliability, performance, and operational excellence of Linux-based compute platforms that support electronic, colocated (colo) trading. This role will focus on building and operating highly resilient, low-latency infrastructure in data centers where milliseconds matter, with an emphasis on automation, standardization, and disciplined incident management. The ideal candidate combines strong Linux systems administration skills with a production engineering mindset and comfort working close to hardware and networks in a high-availability environment.

Job responsibilities

  • Own the build, configuration, and lifecycle management of Linux server fleets supporting colo trading workloads, including provisioning, patching, hardening, and performance tuning.
  • Engineer and maintain automation for OS deployment, configuration management, and continuous compliance, with a bias toward reducing manual touch and improving repeatability.
  • Partner with network, trading technology, and data center teams to optimize latency, throughput, and stability, including kernel, IRQ, CPU isolation, NUMA, and NIC tuning where appropriate.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate reliability design and operational decisioning (e.g., incident/post-incident analysis and requirements traceability), validating outputs and handling operational data according to sensitivity and security requirements.
  • Operate and improve observability across the stack, including metrics, logs, and alerting, and translate signals into actionable runbooks and service-level improvements.
  • Lead incident response for Linux/compute-related events, including rapid triage, mitigation, root-cause analysis, and corrective/preventative actions; drive measurable reduction in recurring incidents.
  • Manage hardware-adjacent responsibilities typical of colo environments, including server break/fix coordination, remote hands engagement, firmware alignment, and standardized rack-level practices.
  • Implement and maintain secure access patterns, secrets handling, and least-privilege controls aligned to enterprise security and audit expectations.
  • Contribute to capacity planning and reliability engineering, including failure-mode thinking, maintenance windows, upgrade strategies, and resiliency testing.
  • Produce clear operational documentation, including build standards, runbooks, incident reports, and environment-specific procedures for colo constraints.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., testing/validation automation and production readiness), ensuring traceability/auditability, resiliency, and security controls.

Required qualifications, capabilities, and skills

  • Bachelor’s Degree in Computer Science, Cybersecurity, Data Science, or related disciplines
  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience
  • Professional experience administering Linux in production (e.g., RHEL-derived, Debian/Ubuntu) in a mission-critical environment.
  • Strong operational competency in troubleshooting performance and reliability issues across OS, hardware, and basic networking layers.
  • Strong expertise in Linux internals, performance tuning, and troubleshooting using enterprise-standard diagnostic tools, with the ability to identify and resolve complex reliability and latency issues in production environments.
  • Proficiency in infrastructure automation using Bash and Python or Go, with hands-on experience in configuration management, CI/CD practices, monitoring, observability, and root-cause analysis to drive operational excellence and service reliability.
  • Demonstrated experience automating system administration tasks using scripting and/or configuration management.
  • Hands-on incident management experience, including participation in on-call rotations and executing structured post-incident reviews.
  • Ability to communicate clearly with both engineering and non-engineering stakeholders, including translating technical issues into business impact.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity.
  • Ability to set team practices for safe AI usage in operations (e.g., review/approval expectations and escalation paths) while maintaining resiliency, security, and auditability outcomes.
Preferred qualifications, capabilities, and skills
  • Experience supporting electronic trading, market connectivity platforms, or similarly latency-sensitive, high-availability environments.
  • Exposure to colocation data centers, including working with remote hands, controlled maintenance windows, and strict change discipline.
  • Experience with performance tuning for low-latency Linux environments (e.g., kernel/CPU pinning strategies, interrupt tuning, time sync discipline).
  • Familiarity with infrastructure-as-code patterns and building standardized “golden” server builds.
  • Experience operating at scale, including fleet management practices, automation-driven patching, and configuration drift control.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available