Senior Service Reliability Engineer

Summary

Senior engineer responsible for ensuring high availability, security, and performance of production applications in cloud environments, using SRE practices, automation, and observability tools.

Job Title

Senior Service Reliability Engineer
Job Description

Key Responsibilities

Application Reliability, Availability & Security

  • Own end-to-end reliability of applications in production environments

  • Ensure adherence to SLA, SLO, and SLI targets

  • Continuously improve availability, latency, and performance metrics

  • Ensure applications comply with security, privacy, and governance standards

  • Support patching, vulnerability remediation, and compliance requirements

Incident Management & Problem Resolution

  • Lead incident response, triage, and resolution across application layers

  • Perform root cause analysis (RCA) and drive permanent fixes

  • Reduce Mean Time to Recovery (MTTR) through automation and process improvements

  • Act as escalation point for critical production issues

Automation & Reliability Engineering

  • Develop automation for deployment, monitoring, and recovery processes

  • Drive “Reliability as Code” and infrastructure automation

  • Build self-healing mechanisms and reduce manual operational effort

  • Design and maintain CI/CD pipelines for application delivery

  • Ensure reliable and consistent deployments using automated pipelines

  • Support application release cycles with zero/low downtime strategies

Observability & Capacity Planning

  • Implement monitoring, logging, and alerting systems

  • Define meaningful alerts and reduce noise/false positives

  • Create dashboards and metrics for real-time health visibility

  • Conduct performance testing and tuning

  • Forecast capacity and ensure scalability of applications

  • Optimize cost vs performance in cloud environments

Required Qualifications

Education

  • Bachelor’s/Master’s degree in Computer Science or related field

Experience

  • 8+ Years in SRE / DevOps / Production Engineering roles

  • Hands-on experience supporting production-grade applications in cloud environments

Technical Skills

Core SRE Skills

  • Experience with JAVA Based Applications

  • Incident management & on-call operations

  • Monitoring & observability (Prometheus, Grafana, ELK, Dynatrace, etc.)

  • Knowledge of SLA/SLO/SLI frameworks

Cloud & Infrastructure

  • AWS / Azure / GCP

  • Kubernetes, Docker (containerization)

  • Infrastructure as Code (Terraform, ARM, etc.)

Automation & CI/CD

  • Jenkins / Azure DevOps / GitHub Actions

  • Scripting: Python / Bash / PowerShell

Application Troubleshooting

  • Strong debugging skills across:

    • Application layer (Java, .NET, Node.js)

    • Middleware (Tomcat, IIS, containers)

    • Database & APIs

Diversity & Inclusion

Amadeus aspires to be a leader in Diversity and Inclusion in the tech industry, enabling every employee to reach their full potential by fostering a culture of belonging and fair treatment, attracting the best talent from all backgrounds, and as a role model for an inclusive employee experience.

Amadeus is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to gender, race, ethnicity, sexual orientation, age, beliefs, disability or any other characteristics protected by law.

Be aware of recruitment scams


Amadeus Group never charges fees, requests payment, or asks for financial information during recruitment. All legitimate opportunities are communicated solely through official Amadeus channels, including our careers website. Any payment request or outreach via unofficial platforms (e.g., WhatsApp, Telegram) should be treated as fraudulent.