Middle Site Reliability Engineer H 8423
Summary
Middle Site Reliability Engineer building and maintaining services for edge server infrastructure, focusing on automation, monitoring, and operational health of enterprise workloads worldwide. Core technologies include Site Reliability Engineering, infrastructure operations, networking fundamentals, and scripting languages like PowerShell.
Join Akvelon — build products used by millions!
Akvelon is an IT company with 20+ years of experience and 1,200+ engineers across 15+ locations worldwide.
We work with both well-known global tech companies, including Microsoft, Facebook, Airbnb, Dropbox, and Pinterest, and with growing startups.
Our teams are involved in different types of engineering projects, from cloud solutions and AI/ML systems to big data, web, and mobile applications.
Since we are remote-first, our engineers work in distributed teams with flexible hours. We value ownership, clear communication, and the ability to take responsibility for your part of the work.
About the role
The client is a multinational technology corporation recognized for its innovation and leadership in software, hardware, and cloud computing.
The project focuses on building and maintaining services that support the lifecycle and operational health of edge server infrastructure. You'll help automate operational workflows, improve system reliability, and support large-scale infrastructure used by enterprise workloads worldwide.
Requirements
- Proven experience in SRE role, infrastructure operations, or similar roles, with a strong background in Azure
- Hands-on experience with Azure Kubernetes Service (AKS) and container orchestration
- Strong understanding of templating, configuration generation, and automation
- Solid networking knowledge and troubleshooting skills, including BGP, HTTP, DNS, and TCP/IP (with focus on hybrid and cloud networking in Azure)
- Ability to troubleshoot production issues across infrastructure and application layers, including basic code debugging and fixes
- Experience connecting production operations with software engineering workflows
- Knowledge of compliance and security practices, including patching, package management, dependency updates, and secure authentication
- Understanding of distributed systems, datacenter operations, and infrastructure best practices
- Experience with PowerShell or another scripting language for automation tasks
- Experience managing Dell edge servers or compute infrastructure
- Familiarity with Dell iDRAC or similar remote management tools
- Experience with large-scale, multi-regional datacenter environments
- Experience using AI-assisted engineering tools such as GitHub Copilot or Azure AI
Responsibilities
- Monitor infrastructure health and participate in incident response and root cause analysis
- Troubleshoot infrastructure, networking, and application issues, including basic code-level fixes
- Support secure and reliable production operations by following compliance and security best practices
- Assist with datacenter operations, hardware troubleshooting, and infrastructure maintenance
- Contribute to automation, monitoring, alerting, and operational improvements
- Collaborate with engineering teams to enhance platform reliability and operational efficiency
- Overlap time requirements until 11:00 AM PST (8:00 PM CET)
- Required to participate in a scheduled DRI on-call rotation, which may include coverage during both business and non-business hours, depending on the team schedule