Senior Observability Platform Engineer
Summary
Build and run a shared observability platform—collecting, storing, and visualizing telemetry while improving reliability and debugging workflows for engineering teams.
You will design, build, and operate a shared observability platform covering telemetry collection, storage, querying, visualization, alerting, diagnostics, and service health. You will improve reliability, scalability, developer experience, instrumentation, and production investigation workflows.
Responsibilities
- Design, build, and operate shared observability platform components
- Build software, services, APIs, integrations, libraries, dashboards, automation, and reusable patterns
- Improve the scalability, reliability, performance, cost-effectiveness, and operational quality of telemetry systems
- Improve developer and operator experience through self-service workflows, documentation, investigation tooling, and platform abstractions
- Work with engineering, infrastructure, trading systems, research, and regional operations teams
- Own the reliability and operational quality of platform components
- Improve telemetry quality, instrumentation, alerting, dashboards, diagnostic workflows, and service health
Requirements
- Strong engineering experience in SRE, software engineering, platform engineering, infrastructure, observability, developer tooling, or distributed systems
- Production experience with failure modes, debugging workflows, service reliability, and operational impact
- Technical understanding of logs, metrics, traces, events, alerting, dashboards, telemetry pipelines, diagnostics, instrumentation quality, and service health
- Experience designing, building, or operating reliable services, platforms, pipelines, tools, or automation used by engineering teams
- Ability to make technical trade-offs across performance, scalability, reliability, complexity, cost, and maintainability
- Ability to turn ambiguous platform problems into practical solutions
- Experience with Kafka, Grafana, ELK/OpenSearch, ClickHouse, VictoriaMetrics, InfluxDB, Telegraf, Vector, OpenTelemetry, Prometheus-style systems, or custom telemetry collectors
Benefits
- Performance-based bonus
- Daily breakfast, lunch, and snacks
- Gym membership
- Weekly in-house chair massages
- Regular social events
- Company trip every two years
- Relocation package
- Visa sponsorship