Site Reliability Engineer Low Latency Trading Systems
You will own reliability, observability, and performance for real-time trading services. You will debug production incidents and latency regressions, operate multi-region Kubernetes infrastructure, harden market-data ingestion, build reconciliation tooling, improve deployment safety, and join an on-call rotation covering equity-market and crypto trading operations.
Responsibilities
- Own production reliability for trading engines, execution gateways, market-data ingestion, and PnL and reconciliation pipelines
- Operate and evolve multi-region AWS EKS clusters using Flux and SOPS
- Build observability through Prometheus metrics and alerting, Datadog logs and dashboards, and SLOs
- Improve deployment safety through progressive rollouts, configuration reload behavior, and deployment guardrails
- Debug stale feeds, rate limits, WebSocket disconnects, order-lifecycle desynchronization, and latency regressions
- Harden market-data ingestion with staleness detection, failover, and replay
- Build reconciliation and data-integrity tooling across gauges, Postgres, and S3 parquet data
- Participate in on-call coverage for US equity-market hours and 24/7 crypto venues
Requirements
- 5+ years of SRE, production engineering, or infrastructure experience
- Experience supporting real-time or latency-sensitive systems
- Go or Rust programming
- Kubernetes
- AWS
- Production experience with stateful latency-sensitive workloads
- PromQL
- Structured-log analysis
- Alert design
- Linux internals
- Networking
- Incident communication
Benefits
- Medical, vision, and dental benefits
- Flexible vacation policy
