Lead/Senior SRE Engineer - Dynatrace
We are in search of an experienced Lead/Senior SRE Engineer with strong Dynatrace expertise to join our team.
In this role, you will drive the implementation of DevOps and SRE practices, shape the technology roadmap, and ensure the reliability, performance, and observability of our production systems. You will collaborate closely with application teams and product owners to foster a culture of operational excellence and continuous improvement.
Responsibilities
- Support implementation of DevOps & SRE practices across teams
- Lead discussions and planning for the SRE technology roadmap
- Establish and maintain SLIs and SLOs, and track MTTR, Lead time for change, Deployment Frequency, and Change Failure Rate
- Create, enhance, and operate monitoring, alerting, operability, and observability for applications using Dynatrace, Splunk, and Grafana
- Evaluate system performance through assessments and monitoring, and propose optimization actions
- Hold application teams accountable for meeting performance and availability SLAs
- Work with product owners to manage error budget, order the toil backlog, and validate progress using team, application, and incident metrics
- Participate in on-call rotation for production events or outages
- Advance continuous improvement of continuous integration & continuous deployment through the CI/CD Pipeline
- Execute troubleshooting techniques, incident management processes, and root cause analysis
- Drive automation wherever feasible to reduce manual work
- Apply cybersecurity measures by continuously performing vulnerability assessment and risk management
- Deliver periodic reporting on progress to management and the customer
- Coordinate with application teams to ease platform adoption, and align communication within the team and with customers
- Assess the current system and develop improvement and enhancement plans
Requirements
- Bachelor's degree in Computer Science or a related discipline, or equivalent work experience
- 5+ years of general IT experience, including 5+ years in DevOps or SRE teams
- Practical experience supporting production infrastructure
- Strong knowledge of CI/CD practices
- Clear understanding of observability fundamentals (monitoring, logging, and tracing)
- Advanced proficiency with Dynatrace and Splunk
- Experience with a leading cloud provider (AWS, Azure, or GCP)
- Capability to run high-availability, fault-tolerant, scalable, distributed software in production environments
- Ability to operate autonomously and in a team, supported by strong organizational and interpersonal skills and experience driving operational maturity
- Excellent analytical thinking and problem-solving skills, with strategic judgment and calm troubleshooting under pressure
- Willingness to adapt quickly to new technologies
- English level B2 (Upper-Intermediate) or higher
Benefits
Opportunity to work on technical challenges that may impact across geographies
Vast opportunities for self-development: online university, knowledge sharing opportunities globally, learning opportunities through external certifications
Opportunity to share your ideas on international platforms
Sponsored Tech Talks & Hackathons
Unlimited access to LinkedIn learning solutions
Possibility to relocate to any EPAM office for short and long-term projects
Focused individual development
Benefit package:
- Health benefits
- Retirement benefits
- Paid time off
- Flexible benefits
Forums to explore beyond work passion (CSR, photography, painting, sports, etc.)