Senior SRE Engineer
We are seeking a Senior SRE Engineer to join our team on a full-time, in-office basis, responsible for ensuring the reliability, scalability, and performance of critical cloud infrastructure and services.
Responsibilities
- Design and maintain high-availability and disaster recovery strategies for critical workloads
- Build and manage cloud infrastructure using Infrastructure-as-Code tools
- Implement and optimize CI/CD pipelines for automated deployments
- Monitor system performance and reliability through observability platforms
- Develop scripts and automation tools to streamline operations
- Ensure robust networking, security, and identity/access management practices
- Lead cross-team collaboration to resolve complex reliability challenges
- Drive continuous improvement initiatives across infrastructure and platform automation
Requirements
- 4-8 years of overall experience in IT
- 4+ years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles
- Expertise in AWS services such as EC2, S3, RDS, IAM, VPC, and Lambda
- Knowledge of Infrastructure-as-Code using Terraform, AWS CDK, or CloudFormation
- Background in CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI
- Proficiency in containerization and orchestration technologies including Docker, Kubernetes, and ECS/EKS
- Competency in monitoring and observability tools such as Datadog, New Relic, Prometheus, Grafana, ELK, and CloudWatch
- Skills in scripting or programming languages such as Python, Bash, or Go
- Understanding of networking, security, and identity/access management in cloud environments
- Excellent communication, problem-solving, and leadership skills with the ability to influence across teams
Nice to have
- AWS or other Cloud Certification such as Solutions Architect or DevOps Engineer
- Familiarity with AIOps, Serverless Architectures, and event-driven systems
- Understanding of FinOps practices and cost optimization frameworks, along with SaaS monitoring tools such as Sumo Logic and PagerDuty
- Exposure to Atlassian tools including Jira, Confluence, and Bitbucket, plus experience with SQL/NoSQL databases
- Showcase of leading cross-functional reliability initiatives or platform-wide automation projects
Benefits
Opportunity to work on technical challenges that may impact across geographies
Vast opportunities for self-development: online university, knowledge sharing opportunities globally, learning opportunities through external certifications
Opportunity to share your ideas on international platforms
Sponsored Tech Talks & Hackathons
Unlimited access to LinkedIn learning solutions
Possibility to relocate to any EPAM office for short and long-term projects
Focused individual development
Benefit package:
- Health benefits
- Retirement benefits
- Paid time off
- Flexible benefits
Forums to explore beyond work passion (CSR, photography, painting, sports, etc.)
Skills
- Atlassian
- Automation
- AWS
- Bash
- Bitbucket
- CDK
- CI/CD
- Cloud
- CloudFormation
- CloudWatch
- Confluence
- Containerization
- Datadog
- DevOps
- Docker
- EC2
- ECS
- EKS
- ELK
- Event Driven Architecture
- FinOps
- GitHub
- GitHub Actions
- GitLab
- Grafana
- IAM
- Infrastructure as Code
- Jenkins
- Jira
- Kubernetes
- Lambda
- Networking
- New Relic
- NoSQL
- Observability
- PagerDuty
- Prometheus
- Python
- RDS
- SaaS
- Serverless
- SQL
- Terraform
- VPC