Senior Site Reliability Engineer (SRE)

Open 26d posting dated 2 weeks ago

Senior Site Reliability Engineer (SRE)

You will be responsible for the operations and reliability of our SaaS application and the underlying infrastructure on GCP, with a focus on automation and continuous improvement. You will tackle the most complex technical challenges, mentor other engineers, and set the standard for operational excellence. This is not a traditional system administration role; you are expected to write code and build systems to eliminate manual work.

Key Responsibilities:

  • Design and implement scalable, self-healing infrastructure on GCP using Terraform and Kubernetes (GKE).

  • Develop automation tooling to reduce operational toil and improve system efficiency.

  • Provide technical guidance and support to the application team on application reliability, performance, observability, and operational best practices.

  • Lead the technical response during major incidents, performing deep-dive analysis to identify root causes.

  • Architect and implement comprehensive monitoring, logging, and alerting solutions to ensure proactive issue detection.

  • Participate in the 24/7 on-call rotation.

  • Drive post-mortem analysis and ensure follow-up actions are implemented.

Required Qualifications:

  • 5+ years of experience in a Site Reliability, DevOps, or Software Engineering role with a focus on infrastructure.

  • Strong, hands-on experience with GCP, specifically with GKE, VPC, Cloud Load Balancing, IAM, and Cloud Monitoring.

  • Proficient in writing production-quality code for automation (Python or Go preferred).

  • Understanding of Java application operations.

  • Expertise in Terraform for managing complex infrastructure.

  • In-depth knowledge of Kubernetes architecture and operations in a production environment.

  • Experience with CI/CD systems (e.g., GitLab CI) and integrating operational controls into pipelines.

  • A systematic problem-solving approach, coupled with strong communication skills and a sense of ownership.

What we offer:

  • Participation in a strategic Cloud & SaaS transformation in a global company.

  • Long-term career growth opportunities within an international organization.

  • The chance to design and develop PSI’s proprietary products.

  • A team of highly experienced specialists eager to share knowledge.

  • Comfortable office environment (small rooms, chillout space, parking).

  • Stability and security of employment in a company with a 50-year tradition.

  • Flexible working hours and a friendly atmosphere with no artificial hierarchy.

  • Well-defined career development paths.

  • Access to technology conferences, training, and language courses (English and German).

  • Benefits package: private medical care, group insurance, benefits platform, Multisport card