Site Reliability Engineer

Summary

Maintains 24/7 production systems, optimizes performance and security, and resolves incidents using SRE practices and tools like Prometheus and Grafana.

Scope of responsibilities:

  • Maintaining high availability (24/7) of production services.

  • Optimizing system performance, security, and availability in line with SRE practices.

  • Resolving incidents, performing Root Cause Analysis (RCA), and conducting post-mortem reviews.

  • Active participation in software architecture design.

  • Defining SLIs/SLOs, monitoring (observability), and implementing remediation plans.

  • Participating in on-call rotation (one weekend every 6 weeks).

  • Planning and executing migrations, DR (Disaster Recovery) tests, and system upgrades.

Requirements:

  • Experience in Production Support or SRE.

  • Automation and monitoring tools (Ansible, Jenkins, Prometheus, Grafana).

  • Good programming and database knowledge (Java, Python, NodeJS + SQL).

  • Practical knowledge of the Software Development Life Cycle (SDLC).

  • Experience in maintaining large instances of Atlassian Jira and Confluence Data Center.

  • Onsite work in Krakow for any 6 days per month.