freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform.

Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform

Key Responsibilities

  • Observability, SRE, DevOps roles with expertise in infrastructure and application reliability
  • Dynatrace, ELK, Splunk, PagerDuty
  • SLI/SLO frameworks
  • Azure Kubernetes Service (AKS), Terraform, Azure managed services

What will you do

  • Design and implement observability-as-code solutions using Terraform for monitoring pipelines, dashboards, and alerting across distributed systems
  • Drive observability improvements using Dynatrace, ELK, Splunk, PagerDuty for real-time performance insights and system visibility
  • Instrument applications for end-to-end observability including distributed tracing, metrics collection, and log aggregation across Node.js and .NET microservices and event-driven architectures
  • Troubleshoot complex production incidents across service layers, databases, caches, and APIs using SLI/SLO frameworks
  • Investigate and resolve Azure Kubernetes Service (AKS) infrastructure issues ensuring reliability and scalability of containerized workloads using Terraform and Azure services (SQL MI, Redis, Functions, Event Grid)
  • Translate business requirements into observable, resilient systems aligned to SLIs/SLOs
  • Automate operational tasks using Infrastructure-as-Code and CI/CD to reduce toil and improve resilience
  • Lead incident response and remediation for critical systems, including blameless postmortems and chaos engineering practices
  • Collaborate with development, platform, and business teams to improve availability, scalability, and operational excellence

What do you need to succeed

Must-have

  • 8+ years experience in SRE, DevOps, or Observability roles focused on infrastructure and application reliability
  • Strong expertise in Dynatrace, ELK, Splunk, PagerDuty and observability principles (instrumentation, correlation IDs, SLIs/SLOs)
  • Advanced proficiency in Azure Kubernetes Service (AKS), Terraform, and Azure managed services (SQL MI, Redis, Functions, Event Grid)
  • Hands-on experience with observability instrumentation (distributed tracing, metrics, logs) across Node.js and .NET microservices and event-driven systems
  • Strong troubleshooting skills across distributed systems (services, databases, caches, APIs) in production environments
  • Incident management expertise using PagerDuty and ServiceNow, including high-severity incident resolution and RCA
  • Knowledge of incident, problem, and change management, SRE principles, blameless postmortems, and chaos engineering
  • Strong communication and leadership skills for cross-functional coordination and incident handling


See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available