Software Engineer
Summary
Builds and maintains automation for a large VR application’s CI/CD, build health, and incident response across cloud-streamed surfaces, using AI-assisted code repair and self-healing systems.
This is a remote position.
Summary:
We are looking for Software Engineer cover to build and maintain the automation that keeps a large UGC application healthy in production across every surface it ships on.
In this role you will build and operate the systems that keep a large UGC application healthy in production. It ships as a single product across multiple surfaces: natively on VR headsets, and on mobile and PC via cloud streaming (Horizon Cross Screen). You will work alongside the runtime and KTLO (keep-the-lights-on) team, focused on reducing manual operational and on-call work through automation, keeping the cloud rendering and deployment pipelines healthy, and keeping CI/CD functional as upstream dependencies change.
Responsibilities
- Build and maintain automation that keeps a large VR application healthy in production (release pipelines, build health, crash triage, incident detection)
- Maintain and improve an AI-assisted code repair system that creates and lands fix diffs autonomously
- Develop tooling to automatically identify broken builds, pinpoint root-cause changes, and recommend or execute fixes
- Monitor production quality metrics and respond to regressions and outages
- Reduce manual on-call burden through automation, with a goal of cutting recurring operational work by 80–90%
- Complete required infrastructure migrations to keep CI/CD pipelines functional as upstream dependencies are retired
Requirements
- 8+ years of professional software engineering experience, or equivalent
- Proven experience building and operating CI/CD, build, release, and cloud deployment pipelines at scale
- Experience operating cloud services and server-side fleets in production, including reliability, capacity, and latency
- Experience building or operating AI-assisted developer tooling or agents that generate or repair code
- Experience building tooling that detects broken builds and traces failures to their root-cause change
- Experience with production monitoring, crash triage, and incident response for a large-scale, multi-surface application
- A track record of reducing operational and on-call load through automation
- Experience completing infrastructure or dependency migrations without breaking downstream CI/CD
Desirable
- Experience with cloud game or application streaming, or remote rendering
- Experience with asset delivery or CDN pipelines at scale
- Experience with capacity, latency, or session-orchestration monitoring for streamed workloads
- Experience operating live-service or large-scale production applications (Live Ops)
- Familiarity with large monorepo build systems and dependency management
- Experience designing self-healing or auto-remediation systems
Benefits
- Competitive salary
- Healthcare contribution and inclusion in company pension scheme
- Work laptop and phone
- 25 days annual leave (pro-rata) plus paid bank holidays
- Expanding workforce with potential for career progression for top performers