Tech jobs
Job listings
Staff Site Reliability Engineer
Staff SRE leading reliability and infrastructure architecture for Legora's AI-native legal workspace platform, defining SLI/SLO frameworks, incident management, and operational excellence across distributed systems—NYC-based, 5 days in-office.
Service Manager & Site Reliability Engineer
Oversees real-time incident response, impact assessment, and cross-team coordination for global production systems while ensuring SLA compliance and software quality via CI/CD integration and reliability best practices.
Staff Observability Platform Engineer
Lead the design and operation of Nscale's observability platform for GPU-based AI infrastructure, ensuring deep visibility into metrics, logs, and traces while partnering with SRE and AI teams.
Senior Observability Platform Engineer
The Senior Observability Platform Engineer will design and scale observability systems for Nscale's GPU cloud infrastructure. The role involves building metrics, logging, and tracing pipelines to support AI/ML workloads using tools like Prometheus, Grafana, and OpenTelemetry.