Senior Site Reliability Engineer H 8423 US
Summary
Senior Site Reliability Engineer responsible for maintaining and optimizing cloud infrastructure and services to ensure high availability and performance.
Join Akvelon — build products used by millions!
Akvelon is an IT company with 20+ years of experience and 1,200+ engineers across 15+ locations worldwide.
We work with both well-known global tech companies, including Microsoft, Facebook, Airbnb, Dropbox, and Pinterest, and with growing startups.
Our teams are involved in different types of engineering projects, from cloud solutions and AI/ML systems to big data, web, and mobile applications.
Since we are remote-first, our engineers work in distributed teams with flexible hours. We value ownership, clear communication, and the ability to take responsibility for your part of the work.
About the project
The project focuses on building and operating reliable cloud infrastructure and services in Azure. The team works across Kubernetes-based environments, hybrid and cloud networking, production infrastructure, automation, and distributed systems. The role combines hands-on SRE and infrastructure operations with software engineering practices, troubleshooting complex production issues, and improving the reliability, security, and scalability of large-scale environments.
Requirements
- Proven experience in an SRE, Infrastructure Operations, or similar role, with a strong background in Azure
- Hands-on experience with Azure Kubernetes Service (AKS) and container orchestration
- Strong understanding of templating, configuration generation, and infrastructure automation
- Solid networking knowledge and troubleshooting skills, including BGP, HTTP, DNS, and TCP/IP, with a focus on hybrid and cloud networking in Azure
- Ability to troubleshoot production issues across infrastructure and application layers, including basic code debugging and fixes
- Experience connecting production operations with software engineering workflows
- Knowledge of compliance and security practices, including patching, package management, dependency updates, and secure authentication
- Understanding of distributed systems, datacenter operations, and infrastructure best practices
- Experience with PowerShell or another scripting language for automation tasks
- Experience managing Dell edge servers or compute infrastructure
- Familiarity with Dell iDRAC or similar remote server management tools
- Experience working with large-scale, multi-regional datacenter environments
- Experience using AI-assisted engineering tools, such as GitHub Copilot or Azure AI
Responsibilities
- Operate and improve Azure-based infrastructure and Kubernetes environments
- Troubleshoot complex production issues across infrastructure, networking, and application layers
- Develop and maintain automation for infrastructure provisioning, configuration, and operational workflows
- Support reliable operation of distributed and hybrid cloud environments
- Investigate networking issues involving BGP, HTTP, DNS, and TCP/IP
- Collaborate with software engineering teams to improve production reliability and operational efficiency
- Support security, compliance, patching, and dependency management initiatives
- Contribute to infrastructure and datacenter best practices and continuous improvement
- Participate in incident response and help drive issues through to resolution
- Required to participate in a scheduled DRI on-call rotation
Base compensation
Salary Range: $50 - $55 per hour
Paid vacation and sick leave