Senior DevOps/SRE Engineer
- Setup and maintenance of IT cloud infrastructure for multiple products
- Setup and maintenance of the automated build/deployment jobs
- Maintenance of various support tools and utilities such as logs, monitoring, gathering user feedback
- Monitoring application performance and resource planning
- Coordination with development and user teams to assess risks, goals, and needs and ensure adequate addressing
- Coordination with the client's infrastructure team to align on prerequisites and expectations
- Development of CI/CD pipeline scripts and templates
- Deep understanding of complete software development life cycle and experience in software development
- Strong understanding and hands-on experience with Azure ecosystem
- Strong understanding and hands-on experience with Azure DevOps (Azure Pipelines)
- Strong understanding and hands-on experience with Kubernetes (AKS)
- Strong understanding and hands-on experience with Docker
- Strong understanding and hands-on experience with networking concepts
- Experience with Terraform and Powershell
- In-depth knowledge of SCM tools (preferably Git)
- Proactive problem-solving and ownership mindset
- Strong communication and collaboration skills
- Deep, practical understanding of SRE principles with demonstrated ownership of production reliability
- Proven experience defining, enforcing, and evolving SLIs, SLOs, and error budgets for complex systems
- Hands-on responsibility for high-availability, mission-critical platforms operating at scale
- Expert-level experience with observability systems (metrics, logs, traces) and building signal-driven alerting that minimizes noise
- Ability to diagnose and resolve systemic issues across Kubernetes, cloud infrastructure, networking, and application layers
- Demonstrated experience in capacity planning, performance engineering, and cost-aware scaling
- Expertise in designing and enforcing resilience and failure-tolerant architectures such as self-healing, redundancy, graceful degradation
- Strong production change discipline including safe rollouts, rollback strategies, and risk management