Site Reliability / Gitops Engineer
NewBe an early applicantThis position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability / Gitops Engineer based in Germany.
Join a global Information Systems team responsible for operating and evolving critical production services at significant scale.
In this role, you’ll use automation, Infrastructure as Code, and GitOps practices to make cloud operations more reliable, consistent, and efficient.
You’ll work across private and public cloud environments, strengthening infrastructure resilience, scalability, observability, and performance.
Your expertise will also influence the evolution of open-source infrastructure technologies through hands-on feedback, bug reporting, and collaboration.
You’ll troubleshoot complex systems, improve operational processes, and help eliminate repetitive manual work through thoughtful automation.
Working with a distributed team of experienced SREs, you’ll have opportunities to share knowledge, mentor colleagues, and contribute to major engineering initiatives.
This is an ideal opportunity for an automation-first technologist who is passionate about Linux, open source, and building robust systems at scale.
Accountabilities:
- Apply Infrastructure as Code expertise to continuously improve automation practices, processes, and operational consistency.
- Automate software operations across private and public clouds while accounting for the complexities of distributed systems.
- Develop new capabilities and improve the resilience, scalability, and reliability of cloud and container infrastructure.
- Maintain operational responsibility for core services, networks, and infrastructure, ensuring reliable day-to-day performance.
- Troubleshoot complex systems, perform capacity planning and performance investigations, and develop strong operational expertise.
- Set up, maintain, and use observability and monitoring solutions such as Prometheus, Grafana, and Elasticsearch.
- Design and maintain monitoring and alerting for critical systems and services.
- Collaborate with development teams on service architecture, documentation, playbooks, policies, and operational procedures.
- Work closely with globally distributed engineering, operations, and support teams to resolve issues and improve services.
- Dedicate focused development time to larger engineering projects and the automation of repetitive manual processes.
- Share technical knowledge and best practices through design sessions, mentoring, and collaborative problem-solving.
- Take final responsibility for resolving time-critical operational escalations.
- Deep experience defining IT operations through code, using version control, peer review, and CI/CD to deploy application and infrastructure changes.
- Strong modern software engineering practices, including peer review, unit testing, source control management, CI/CD, and Agile methodologies.
- Significant Python development experience, including work on large or complex projects.
- Practical knowledge of Linux networking, routing, firewalls, and related infrastructure concepts.
- Familiarity with Linux storage technologies, ranging from Ceph to database systems.
- Hands-on experience administering enterprise Linux servers.
- Extensive understanding of cloud computing concepts, architectures, and technologies.
- Bachelor’s degree or higher, preferably in computer science, software engineering, or a related technical discipline.
- Strong English communication skills across written and spoken channels, including email, chat, video calls, voice calls, and in-person collaboration.
- Strong troubleshooting abilities, with the curiosity and persistence to investigate issues from the Linux kernel through to the web layer.
- Ability to collaborate effectively while knowing when to seek input from teammates and subject-matter experts.
- Adaptability, willingness to learn quickly, and comfort working in fast-changing technical environments.
- Ability to thrive within globally distributed teams and collaborate across different locations and time zones.
- Strong interest in open-source technologies, particularly Ubuntu or Debian.
- Opportunity to work on production infrastructure supporting large-scale global services.
- Exposure to private and public cloud environments, Infrastructure as Code, GitOps, CI/CD, observability, and open-source technologies.
- Dedicated development time for impactful automation and larger engineering projects.
- Collaboration with a highly experienced, globally distributed SRE and engineering community.
- Opportunities for mentoring, knowledge sharing, and cross-functional technical collaboration.
- Remote work flexibility, with the role available across time zones.
- Opportunities to meet colleagues in person 2–4 times per year at internal events, typically lasting 1–2 weeks.
- International exposure through collaboration with distributed teams and participation in global company events.
- Compensation and benefits are determined according to the role, location, experience, and applicable company policies.
Requirements:
Benefits:
As published by lever
Resume/CV, Full name, Email, Phone, Current location, Current company