Senior Site Reliability Engineer (Data Platform) @ Link Group
Summary
Senior SRE for a global cloud/content-delivery client's data platform: keeps a massive telemetry pipeline collecting data from hundreds of thousands of servers reliable, working with large-scale NoSQL databases (Cassandra, MongoDB), building observability with Prometheus/Grafana, and writing automation in Python, Go, or Java.
On behalf of our client, a global leader in cloud and content delivery, we are seeking a highly skilled Site Reliability Engineer to join a critical team at the heart of their global infrastructure.
This team is responsible for a massive data pipeline that collects telemetry from hundreds of thousands of servers and delivers this vital data to customers for analytics and reporting. This is a role for a true systems engineer who understands that reliability at this scale is fundamentally a data problem.
We are not looking for a typical DevOps engineer. We need a deep-thinking SRE who is passionate about the reliability of large-scale data systems, observability, and solving complex problems at the intersection of software, network, and infrastructure.
Who We're Looking For (Your Profile):
- You have deep, hands-on experience engineering and operating large-scale, distributed data systems. You are comfortable working with data at scale, analyzing metrics, and troubleshooting complex data integrity issues.
- Practical experience with modern NoSQL databases is essential. We have a strong preference for candidates who have worked with Cassandra or MongoDB, but we are open to experts in other similar technologies.
- You are a seasoned SRE or Systems/Infrastructure Engineer with at least 5 years of experience managing mission-critical distributed systems.
- You are a proficient programmer with experience building automation and tooling in languages like Python, Go, or Java.
- You have a strong foundation in Linux/Unix administration and low-level system troubleshooting.
- You have a relentless drive to find the root cause of complex problems and deliver robust, production-grade solutions.
On behalf of our client, a global leader in cloud and content delivery, we are seeking a highly skilled Site Reliability Engineer to join a critical team at the heart of their global infrastructure.
This team is responsible for a massive data pipeline that collects telemetry from hundreds of thousands of servers and delivers this vital data to customers for analytics and reporting. This is a role for a true systems engineer who understands that reliability at this scale is fundamentally a data problem.
We are not looking for a typical DevOps engineer. We need a deep-thinking SRE who is passionate about the reliability of large-scale data systems, observability, and solving complex problems at the intersection of software, network, and infrastructure.
,(Engineer World-Class Data Systems: Your core focus will be the reliability and performance of massive data platforms. You will work extensively with large-scale distributed databases to ensure data integrity, availability, and low-latency performance. While the environment heavily utilizes Cassandra and MongoDB, your expertise with other modern NoSQL or distributed data systems will be highly valued., Become the Ultimate Troubleshooter: You will be the highest technical escalation point for the most complex reliability and performance issues, leading investigations that span the entire global stack., Build Insightful Observability: You will design and build the observability fabric that allows us to understand the health of the platform in real-time, using tools like Prometheus and Grafana to create actionable insights, not just noise., Solve Problems with Code: You will write high-quality automation and internal tooling using Python, Go, or Java. Your code will streamline operations, automate diagnostics (increasingly with AI assistance), and empower other teams with self-service workflows., Drive Long-Term Reliability: You will partner closely with Engineering, Product, and Network teams to influence architecture, identify systemic weaknesses, and deliver scalable, long-term solutions that prevent future incidents.) Requirements: NoSQL, Cassandra, MongoDB, SRE, Python, Go, Linux, Unix Additionally: Private healthcare, Sport subscription, Foreign languages classes, Life Insurance, Cafeteria system.