Senior Site Reliability Engineer (Python/Kubernetes)
Summary
Luxoft Poland is hiring a Senior Site Reliability Engineer to build and operate a cloud observability platform for SAP, a large German enterprise software company. Day-to-day work centers on L3 incident resolution, Python automation and AI agents for L1/L2 ops, and 24x7 shift rotation, using Kubernetes, Grafana/Prometheus, and CI/CD tooling.
We're building a Cloud Observability platform for market leading large Germany based company with 440,000 customers worldwide.
Originally known for leadership in enterprise resource planning (ERP) software, the company has evolved to become a market leader in end-to-end enterprise application software, database, analytics, intelligent technologies, and experience management. A top cloud company with 200 million users worldwide, the company helps businesses of all sizes and in all industries to operate profitably, adapt continuously, and achieve their purpose.
Responsibilities
- Independently resolve L3 operational issues using engineering expertise, AI tooling, and purpose-built automations
- Build and maintain AI agents for L1 & L2 automated operations
- Develop automation solutions that reduce manual operational effort
- Participate in development tasks for run-operation
- Monitor platform availability and manage incident response/resolution by severity (P1/P2/P3)
- Perform root cause analysis and determine if issues require code-level fixes
- Participate in 24x7 shift rotation during weekdays and on-call on weekends
Mandatory Skills
- Python
- Security Monitoring & Observability
- Telemetry
Mandatory Skills Description
- Strong Python development skills for automation and AI agent development
- Kubernetes administration and troubleshooting
- Experience with observability and monitoring platforms (Grafana, Prometheus, Jaeger, Splunk)
- Incident management experience with severity-based response processes (P1/P2/P3)
- CI/CD pipeline experience (GitHub Actions, ArgoCD)
- Experience with AI/ML concepts for building intelligent automation
- Strong root cause analysis skills
- Willingness to work in 24x7 shift rotation
Nice-to-Have Skills Description
- Experience with Generative AI / LLM for operational automation
- Knowledge of OpenTelemetry and Kafka
- Experience with OpenSearch / Elasticsearch