Senior Site Reliability Engineer (Data Platform)
Summary
Senior SRE joining a team running a massive telemetry data pipeline for a global cloud and content delivery leader, collecting data from hundreds of thousands of servers. The work centers on reliability of large-scale distributed data systems: NoSQL databases (Cassandra/MongoDB preferred), observability, automation in Python/Go/Java, and Linux troubleshooting.
On behalf of our client, a global leader in cloud and content delivery, we are seeking a highly skilled Site Reliability Engineer to join a critical team at the heart of their global infrastructure.
This team is responsible for a massive data pipeline that collects telemetry from hundreds of thousands of servers and delivers this vital data to customers for analytics and reporting. This is a role for a true systems engineer who understands that reliability at this scale is fundamentally a data problem.
We are not looking for a typical DevOps engineer. We need a deep-thinking SRE who is passionate about the reliability of large-scale data systems, observability, and solving complex problems at the intersection of software, network, and infrastructure.
Who We're Looking For (Your Profile):
- You have deep, hands-on experience engineering and operating large-scale, distributed data systems. You are comfortable working with data at scale, analyzing metrics, and troubleshooting complex data integrity issues.
- Practical experience with modern NoSQL databases is essential. We have a strong preference for candidates who have worked with Cassandra or MongoDB, but we are open to experts in other similar technologies.
- You are a seasoned SRE or Systems/Infrastructure Engineer with at least 5 years of experience managing mission-critical distributed systems.
- You are a proficient programmer with experience building automation and tooling in languages like Python, Go, or Java.
- You have a strong foundation in Linux/Unix administration and low-level system troubleshooting.
- You have a relentless drive to find the root cause of complex problems and deliver robust, production-grade solutions.