Senior Platform Engineer (SRE)
Summary
A senior platform engineer (SRE) at a high-scale AI-powered sports media platform, designing and operating distributed systems with petabyte-scale media storage and millions of video minutes monthly, using AWS, Kubernetes, infrastructure-as-code, and Kafka to build highly available, cost-efficient infrastructure and developer platforms.
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Platform Engineer (SRE) based in France.
This is a senior infrastructure role at the heart of a high-scale, AI-powered sports media platform.
You will help design and operate distributed systems handling petabyte-scale media storage and millions of minutes of video every month.
The role combines deep AWS expertise, Kubernetes, infrastructure-as-code, distributed systems, developer experience, and platform reliability.
You will work as one of the core members of a lean Platform team, with significant autonomy over architecture, technical decisions, and delivery.
A major focus will be building highly available infrastructure capable of supporting critical live sports workflows across multiple regions.
You will also shape internal developer platforms, cost-efficient storage strategies, security practices, and multi-cloud capabilities.
It is an opportunity to solve complex infrastructure challenges at meaningful scale while working in a fast-moving, international, engineering-led environment.
Accountabilities:
- Design, build, operate, and continuously improve AWS infrastructure across multiple regions.
- Develop and maintain Kubernetes-based infrastructure managed through Terraform, Terragrunt, and Helm, with infrastructure defined and reviewed entirely as code.
- Design and operate highly available, distributed architectures capable of supporting critical live-event workloads without conventional maintenance windows.
- Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
- Build and operate event-driven infrastructure using technologies such as Kafka, supporting high-volume media ingestion, processing, and distribution.
- Drive infrastructure cost optimization, particularly around storage and data egress, treating cloud economics as an architectural consideration.
- Contribute to multi-cloud and Bring Your Own Storage architectures that enable customers and partners to retain media in their own cloud environments while maintaining efficient platform access.
- Build and evolve an internal developer platform that enables engineers to develop locally, deploy through reliable pipelines, and provision infrastructure through self-service workflows.
- Establish clear engineering standards and processes for infrastructure definition, modification, review, deployment, and operational ownership.
- Partner with product engineers to reduce platform friction, respond effectively to infrastructure needs, and prevent the Platform function from becoming a delivery bottleneck.
- Strengthen infrastructure security through vulnerability analysis, patching, hardening, and continuous security improvements in support of SOC 2 and broader compliance requirements.
- Participate in on-call responsibilities for critical infrastructure, sharing operational coverage across the Platform team.
- Own incidents through resolution and contribute to detailed postmortems that document technical causes, business impact, costs, and preventative actions.
- Reduce alert noise, operational toil, and recurring incidents through proactive reliability engineering and platform improvements.
- Lead projects autonomously from discovery and scoping through architectural decisions, implementation, production deployment, and ongoing operation.
- Create and defend technical proposals and RFCs, contributing to engineering-wide technical discussions and architectural decisions.
- Collaborate closely with engineering leadership and cross-functional teams while communicating complex technical concepts clearly to stakeholders with different levels of technical expertise.
- Approximately 6+ years of experience in SRE, Platform Engineering, Infrastructure Engineering, or a closely related discipline, with greater emphasis placed on demonstrated impact and technical depth than on a specific number of years.
- Deep, hands-on AWS expertise and strong production experience with Terraform.
- Solid experience designing, deploying, and operating Kubernetes in production environments.
- Proven experience designing distributed systems rather than simply operating existing infrastructure.
- Strong understanding of high availability, multi-region architectures, event-driven systems, and reliability engineering, with the ability to explain and defend architectural trade-offs.
- Experience with infrastructure-as-code and tools such as Terraform, Terragrunt, and Helm.
- Strong understanding of cloud storage, ideally including large-scale object storage and the economics of storage and data egress.
- Experience with event-streaming technologies such as Kafka is highly valuable.
- Strong builder mentality, with the ability to take ambiguous technical problems from discovery and scoping through solution design, implementation, production deployment, and ongoing improvement.
- Ability to work autonomously and take ownership of infrastructure projects from end to end.
- Strong communication skills, including the ability to adapt technical explanations to engineers, product stakeholders, and other audiences.
- A strong security mindset, with practical awareness of infrastructure hardening, vulnerability management, patching, and secure operational practices.
- Comfortable participating in an on-call rotation and taking ownership of production incidents.
- Experience establishing platform foundations, engineering processes, or internal infrastructure practices in an early-stage or rapidly scaling environment is a strong advantage.
- Experience with media, video, storage infrastructure, ingestion pipelines, transcoding, live streaming, or high-volume egress is a plus.
- Familiarity with technologies such as PostgreSQL, OpenSearch, Temporal, or Go is beneficial.
- Comfortable working in a remote-first, international environment and collaborating across distributed teams.
- €90,000–€110,000 annual compensation.
- €150,000 in equity, providing an opportunity to participate in the long-term value created by the business.
- Fully remote position within Europe.
- Remote-first international working environment with regular opportunities for team gatherings.
- Significant autonomy and ownership over technical architecture, infrastructure strategy, and project delivery.
- Opportunity to work on infrastructure operating at substantial scale, including 15+ petabytes of media storage and 145M+ minutes of video ingested monthly.
- Direct influence over platform architecture and infrastructure decisions that have both technical and commercial impact.
- Opportunity to solve challenging problems across AWS, Kubernetes, distributed systems, event-driven architectures, cloud economics, and developer platforms.
- Shared on-call model designed to distribute operational responsibility across the Platform team.
- Strong focus on learning from incidents through structured postmortems and continuous reliability improvements.
- Opportunity to work alongside experienced engineers in an international team spanning Europe and the US.
- Exposure to high-profile sports organizations and mission-critical workloads where infrastructure reliability directly affects live-event operations.
- Engineering culture centered on autonomy, technical ownership, RFC-driven decisions, collaboration, and continuous improvement.