Site Reliability Engineer, Studios
Who We Are:
IMG is a leading global sports marketing agency, specializing in media rights management and sales, multi-channel content production and distribution, brand partnerships, strategic consulting, digital services, and event management. It powers growth of revenues, fanbases and IP for more than 250 federations, associations, events, and teams, including the National Football League, English Premier League, International Olympic Committee, National Hockey League, Major League Soccer, ATP and WTA Tours, the AELTC (Wimbledon), Euroleague Basketball, CONMEBOL, World Rugby, DP World Tour, and The R&A, as well as UFC, WWE, and PBR. IMG is a subsidiary of TKO Group Holdings, Inc. (NYSE: TKO), a premium sports and entertainment company.

TKO Group Holdings, Inc. (NYSE: TKO) is a premium sports and entertainment company. TKO owns iconic properties including UFC, the world’s premier mixed martial arts organization; WWE, the global leader in sports entertainment; and PBR, the world’s premier bull riding organization. Together, these properties reach 1 billion households across 210 countries and territories and organize more than 500 live events year-round, attracting more than three million fans. TKO also services and partners with major sports rights holders through IMG, an industry-leading global sports marketing agency; and On Location, a global leader in premium experiential hospitality.

Working Conditions
Permanent Position, Mon-Fri, 9am-5pm
This role is based at our facilities in Stockley Park, Uxbridge, with hybrid working options where applicable.
You may be required to work unsociable hours, including occasional weekends or on-call rotations, to support live operations and critical systems.
Occasional travel may be required depending on project and client needs.
IMG is looking for a Site Reliability Engineer to help design, build, operate, and continuously improve resilient, secure, and highly available platforms that underpin our digital, cloud, and broadcast-adjacent services. This role is suited to someone who combines strong infrastructure and software engineering capability with an operational mindset, and who can help embed reliability engineering practices across systems that support live, business-critical environments.
The successful candidate will play a key role in improving service reliability, observability, incident response, automation, and disaster recovery readiness across IMG platforms, while working closely with engineering, operations, and project stakeholders.
Key Responsibilities and Accountabilities
Design, build, and maintain reliable, scalable infrastructure and platform services across on-premises and cloud environments.
Improve service availability, latency, performance, and operational efficiency through engineering-led reliability practices.
Build and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards, and service health indicators.
Define and maintain SLIs, SLOs, alerting standards, and operational runbooks for critical services.
Automate infrastructure provisioning, configuration, deployment, and recovery processes using Infrastructure as Code and scripting.
Partner with software, platform, broadcast engineering, and operational teams to improve release quality, resilience, and supportability.
Act as an escalation point for production incidents, leading or supporting rapid diagnosis, mitigation, communication, and post-incident follow-up.
Drive root cause analysis and corrective actions following incidents, with a focus on prevention and continuous improvement.
Support the design, testing, and documentation of high availability, backup, failover, and disaster recovery arrangements.
Help enforce security, access control, patching, and operational best practices across infrastructure and services.
Optimise system capacity, cost, and performance across environments.
Produce and maintain clear technical documentation, operational procedures, and support handover materials.
Support live event and critical operational workflows where reliability, rapid response, and stakeholder communication are essential.
Contribute to technical planning for new services, migrations, and platform enhancements, ensuring resilience is designed in from the start.
Improve reliability, stability, and recoverability of IMG’s platform services.
Reduced mean time to detect and resolve incidents through better observability and response processes.
Higher levels of automation across provisioning, deployment, remediation, and operational support.
Clearer operational ownership, documentation, and service standards across critical environments.
Stronger resilience for live and client-facing workflows through tested failover and recovery approaches.
Knowledge and Experience
Mandatory
Proven experience in a Site Reliability Engineer, DevOps Engineer, Platform Engineer, or similar role.
Strong knowledge of Linux and operating system fundamentals.
Strong hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
Experience with containerisation and orchestration technologies such as Docker and Kubernetes.
Strong experience with CI/CD tooling and modern software delivery practices.
Hands-on experience with Infrastructure as Code tools such as Terraform or CloudFormation.
Experience with monitoring, logging, and alerting tooling, and with designing actionable observability solutions.
Solid understanding of networking, security, system architecture, and distributed systems principles.
Strong scripting or programming capability in Python, Bash, or similar languages.
Experience working in high-availability, live production, or other business-critical operational environments.
Strong troubleshooting skills, calm decision-making under pressure, and a continuous improvement mindset.
Excellent communication and collaboration skills, including the ability to work effectively with technical and non-technical stakeholders.
Desirable
Experience supporting media, broadcast, streaming, or live event platforms.
Familiarity with incident management, postmortem practice, and error-budget based operational models.
Experience with resilience engineering, multi-site failover, and disaster recovery testing.
Exposure to event-driven or low-latency systems, media transport, or hybrid on-prem/cloud architectures.
Understanding of compliance, operational risk management, and support processes in client-facing environments.
Personal Attributes
Proactive and ownership driven.
Methodical, analytical, and detail oriented.
Comfortable operating in fast-moving, high-pressure environments.
Pragmatic in balancing engineering excellence with operational needs.
Collaborative, service oriented, and committed to raising reliability standards across teams.
In addition, success at IMG is driven by four core competencies that apply to all employees:
Business Acumen – Understanding financial drivers, interpreting business data, aligning decisions to strategic outcomes
Operational Excellence – Driving efficiency, governance, and continuous improvement in delivery & operations
Innovation Mindset – Cultivating curiosity, experimentation, and forward-looking capability development
Leadership & Collaboration – Inspiring others, building trust, and enabling collaboration across teams
TKO EEO Statement
TKO is an Equal Opportunity Employer and complies with all applicable federal, state, and local laws regarding non-discrimination in employment. TKO makes employment decisions based on merit and qualifications, without considering an employee’s or applicant’s race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, marital status, veteran status, or any other basis prohibited under federal or local laws governing non-discrimination in employment in every location in which the Company has facilities. TKO also provides reasonable accommodations for qualified individuals with disabilities in accordance with the Americans with Disabilities Act (ADA) and applicable state or local laws. For information about Privacy and Information Security for TKO employment candidates, please review our Privacy Policy. For information regarding Terms of Use for this and other TKO websites, please review our Terms of Use.


TKO EEO Statement:
TKO is an Equal Opportunity Employer and complies with all applicable federal, state, and local laws regarding non-discrimination in employment. TKO makes employment decisions based on merit and qualifications, without considering an employee’s or applicant’s race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, marital status, veteran status, or any other basis prohibited under federal, state or local laws governing non-discrimination in employment in every location in which the Company has facilities. TKO also provides reasonable accommodations for qualified individuals with disabilities in accordance with the Americans with Disabilities Act (ADA) and applicable state or local laws. For information about Privacy and Information Security for TKO employment candidates, please review our Privacy Policy. For information regarding Terms of Use for this and other TKO websites, please review our Terms of Use.