Senior DevOps Engineer
Summary
Senior DevOps Engineer with an SRE focus at HCSS, a construction software company. Day to day: managing and optimizing Azure SQL Elastic Pools, building observability with Grafana, automating infrastructure with Terraform and CI/CD pipelines, and leading high availability, disaster recovery, and incident response across Azure/AWS cloud environments.
- 8+ years of experience in DevOps or SRE roles with a strong focus on cloud infrastructure and systems reliability
- 3+ years of hands-on experience with managing Azure SQL Elastic Pools, including performance tuning, scaling, and automation
- 5+ years of expertise in Azure cloud services including networking, compute, databases, and identity
- 3+ years of experience applying SRE principles including SLIs, SLOs, and incident management best practices
- Extensive experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM templates
- Extensive experience building and managing CI/CD pipelines with tools like Azure DevOps or GitHub Actions
- Strong scripting skills using Azure CLI and PowerShell for automation and operational tasks
- Experience with monitoring and observability platforms, ideally Grafana, or a strong foundation in similar tools
- Strong troubleshooting and problem-solving abilities.
- Excellent communication skills and a collaborative mindset to work with cross-functional teams.
- Ability to work independently, manage multiple tasks, and prioritize efficiently.
- A proactive attitude toward continuous improvement and learning.
- Managed 10+ Elastic pools and 100+ databases in Azure
- Advanced level certifications on cloud infrastructure like Az-400 or equivalent
- Take ownership of the design, scaling, and optimization of Azure SQL Elastic Pools
- Monitor and tune pool performance to ensure efficiency and SLA compliance
- Establish observability and alerting for SQL resource consumption, errors, and performance anomalies
- Automate provisioning, scaling, and failover using infrastructure and scripting tools
- Collaborate with database and application teams to align on resource usage strategies
- Design and implement highly available and fault tolerant systems
- Develop and maintain disaster recovery strategies across critical services
- Perform regular failover testing, documentation, and validation of recovery procedures
- Work closely with infrastructure and development teams to ensure business continuity objectives are met
- Implement and manage observability stacks with logs, metrics, traces, and alerting
- Create dashboards and alerts in Grafana or similar platforms to track key system indicators
- Develop automated solutions for provisioning, monitoring, and maintenance tasks
- Continuously improve system visibility and reduce time to detect and resolve issues
- Assist in the automation of performance testing to proactively identify bottlenecks, validate scalability and ensure reliable system behavior
- Establish and refine incident response processes, escalation workflows, and resolution protocols
- Lead root cause analysis, post-incident reviews, and continuous improvement efforts, including following up to ensure identified improvements to the application or process are implemented.
- Develop and maintain runbooks, diagnostic tools, and automated remediation solutions
- Champion a blameless culture of reliability and operational readiness across engineering teams
- Architect and manage scalable and secure cloud infrastructure in Azure or AWS
- Provision and manage services including compute, networking, storage, and containerized workloads
- Continuously monitor performance, latency, and uptime to ensure system health
- Apply cost optimization practices while aligning infrastructure with business goals
- Define and implement infrastructure using tools such as Terraform
- Maintain modular, version-controlled infrastructure code that supports environment consistency
- Apply automation and policy enforcement to reduce drift and improve auditability
- Ensure IaC best practices are embedded in the development lifecycle
- Mentor and support junior DevOps engineers through code reviews, knowledge sharing, and technical guidance
- Partner with cross-functional teams including development, security, and database operations to drive initiatives
- Act as a subject matter expert in SRE practices and reliability-driven engineering
- Lead continuous improvement efforts across infrastructure and operations processes
- Remote Requirements
- Employees will be expected to come into the office on a periodic basis.
- Baseline expectations for roles are as follows but may fluctuate based on manager’s discretion:
- Individual Contributor - Up to 2x per year
- Employees will be expected to attend HCSS sponsored events per manager discretion (ex. UGM)
- Flexibility to work Remotely
- Medical, dental, and vision coverage with company-paid and employee-paid options
- Paid holidays, sick days, and personal time off
- Employee Resource Groups (ERGs) that foster connection and inclusion
- On-site amenities including a covered basketball court, soccer field, track, pickleball/tennis courts, gym, etc.
- Dog-friendly campus and WiFi-accessible courtyards
- 401(k) with a 5% company match
- Coverage for employee professional development and wellness
- And more!