Lead/Senior SRE Engineer - Dynatrace
We are looking for an experienced Lead/Senior SRE Engineer with strong Dynatrace expertise to join our team.
In this role, you will drive the implementation of DevOps and SRE practices, shape the technology roadmap, and ensure the reliability, performance, and observability of our production systems. You will collaborate closely with application teams and product owners to foster a culture of operational excellence and continuous improvement.
Responsibilities
- Implement DevOps & SRE practices across the organization
- Drive conversations and decisions on the SRE team technology roadmap
- Define, refine, and maintain SLIs and SLOs, plus MTTR, Lead time for change, Deployment Frequency, and Change Failure Rate metrics
- Design, build, and run monitoring, alerting, operability, and observability for applications using Dynatrace, Splunk, and Grafana
- Conduct performance assessments and monitoring, then recommend performance improvements
- Ensure application teams meet performance and availability SLAs
- Partner with product owners to manage error budget, prioritize the toil backlog, and validate outcomes against team, application, and incident metrics
- Join an on-call rotation to respond to production events or outages
- Improve continuous integration & continuous deployment via the CI/CD Pipeline
- Apply troubleshooting methods, incident management, and root cause analysis
- Promote and create automated processes wherever possible
- Implement cybersecurity measures through ongoing vulnerability assessment and risk management
- Prepare periodic progress reporting for management and the customer
- Collaborate with application teams to simplify adoption of the platform, and coordinate communication within the team and with customers
- Review the current system and produce plans for enhancements and improvements
Requirements
- Bachelor’s degree in Computer Science or related field, or equivalent practical experience
- 5+ years of general IT experience, including 5+ years within DevOps or SRE teams
- Solid background in supporting production infrastructure
- Hands-on experience with CI/CD
- Deep understanding of observability (monitoring, logging, and tracing)
- Proven expertise in Dynatrace and Splunk
- Familiarity with a major cloud provider (AWS, Azure, or GCP)
- Strong ability to operate high-availability, fault-tolerant, scalable, distributed software in production
- Demonstrated ability to work independently and collaboratively, with strong organization, interpersonal skills, and experience building operational maturity
- Strong analytical and problem-solving mindset, including strategic thinking and the ability to troubleshoot under pressure
- Adaptability to pick up new technologies quickly
- English proficiency at B2 (Upper-Intermediate) or higher
Benefits
Opportunity to work on technical challenges that may impact across geographies
Vast opportunities for self-development: online university, knowledge sharing opportunities globally, learning opportunities through external certifications
Opportunity to share your ideas on international platforms
Sponsored Tech Talks & Hackathons
Unlimited access to LinkedIn learning solutions
Possibility to relocate to any EPAM office for short and long-term projects
Focused individual development
Benefit package:
- Health benefits
- Retirement benefits
- Paid time off
- Flexible benefits
Forums to explore beyond work passion (CSR, photography, painting, sports, etc.)