Senior Site Reliability Engineer H 8423
Summary
Maintains and automates the lifecycle of global edge server infrastructure, ensuring reliability, security, and performance for a large-scale enterprise platform.
Join Akvelon — build products used by millions!
Akvelon is an IT company with 20+ years of experience and 1,200+ engineers across 15+ locations worldwide.
We work with both well-known global tech companies, including Microsoft, Facebook, Airbnb, Dropbox, and Pinterest, and with growing startups.
Our teams are involved in different types of engineering projects, from cloud solutions and AI/ML systems to big data, web, and mobile applications.
Since we are remote-first, our engineers work in distributed teams with flexible hours. We value ownership, clear communication, and the ability to take responsibility for your part of the work.
About the role
The client is a multinational technology corporation recognized for its innovation and leadership in the software, hardware, and cloud computing industries. Renowned for developing and delivering cutting-edge products and services that empower individuals and businesses worldwide.
The project focuses on building and maintaining services that manage the lifecycle and operational health of edge server infrastructure. It includes automating workflows, improving reliability, and ensuring efficient performance across global systems, providing a platform that supports diverse workloads and critical enterprise operations.
Requirements
- Proven experience in Site Reliability Engineering (SRE), infrastructure operations, or similar roles, with a strong background in Azure
- Hands-on experience with Azure Kubernetes Service (AKS) and container orchestration
- Strong understanding of templating, configuration generation, and automation
- Solid networking knowledge and troubleshooting skills, including BGP, HTTP, DNS, and TCP/IP (with focus on hybrid and cloud networking in Azure)
- Ability to troubleshoot production issues across infrastructure and application layers, including basic code debugging and fixes
- Experience connecting production operations with software engineering workflows
- Knowledge of compliance and security practices, including patching, package management, dependency updates, and secure authentication
- Understanding of distributed systems, datacenter operations, and infrastructure best practices
- Experience with PowerShell or another scripting language for automation tasks
- Experience managing Dell edge servers or compute infrastructure
- Familiarity with Dell iDRAC or similar remote server management tools
- Experience working with large-scale, multi-regional datacenter environments
- Experience using AI-assisted engineering tools such as GitHub Copilot or Azure AI
Responsibilities
- Monitor infrastructure health, respond to incidents, and perform root cause analysis
- Troubleshoot issues across infrastructure, networking, and application layers, including code-level fixes when needed
- Maintain secure and reliable operations through compliance with SLOs and security requirements
- Support datacenter expansion and coordinate hardware troubleshooting
- Improve automation, alerting, and self-healing capabilities
- Collaborate with engineering teams to improve reliability, scalability, and operational efficiency
- Overlap time requirements until 11:00 AM PST (8:00 PM CET)
- Required to participate in a scheduled DRI on-call rotation, which may include coverage during both business and non-business hours, depending on the team schedule