freehire launches on Product Hunt on 26 August.

Follow →

Senior Site Reliability Engineer, Observability

Open 36d

You'll spend most of your time doing hands-on observability and reliability engineering work, while also coaching and consulting with product teams to help them build operational maturity. You'll join the Technical Operations team and work across Azure and AWS environments supporting predominantly Windows-based infrastructure that handles significant payment volume for enterprise treasury customers. You'll design monitoring, alerting, and dashboards in New Relic, write NRQL queries, define SLOs/SLIs and error budgets, and lead alert noise reduction efforts. You'll develop and maintain Terraform infrastructure as code, establish IaC governance standards, and author Azure DevOps pipelines. You'll administer Incident.IO, build out incident management foundations including on-call rotations and escalation policies, track reliability metrics, and respond to production incidents. You'll also enable engineering teams through workshops and documentation to adopt improved observability and incident management practices.

Responsibilities

  • Design and implement monitoring, alerting, and dashboards in New Relic across Azure and AWS
  • Write NRQL queries for troubleshooting, analysis, and reporting
  • Define and implement SLOs/SLIs and error budgets
  • Coach teams on using SLOs/SLIs to balance feature velocity with reliability
  • Lead alert noise reduction and signal quality engineering
  • Optimize observability costs through log ingestion management and pipeline rules
  • Partner with engineering teams to improve observability maturity
  • Develop and maintain Terraform infrastructure as code for monitoring resources
  • Establish and enforce IaC governance standards for observability infrastructure
  • Author and troubleshoot Azure DevOps pipelines
  • Administer and configure Incident.IO including alert routing and notification workflows
  • Build incident management foundations including PIR/postmortem processes and on-call rotation design
  • Track and report on MTTR, MTTD, and incident frequency
  • Respond to and debrief on production incidents
  • Enable stream-aligned engineering teams through workshops and hands-on guidance
  • Collaborate with the Subsystems Platform Team on self-service observability capabilities
  • Build team competency through documentation and training materials

Requirements

  • 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with a focus on observability and production operations
  • Proven ability to deliver hands-on engineering work while coaching and mentoring teams
  • Experience working in Agile/Scrum environments
  • Expert-level hands-on experience with New Relic (APM, Infrastructure, Logs, Synthetics, Alerts) and strong NRQL proficiency
  • Deep understanding of structured logging, metrics collection (RED/USE methods), and distributed tracing
  • Expertise defining and implementing SLOs/SLIs and error budgets
  • Hands-on experience with incident management platforms (Incident.IO, PagerDuty, OpsGenie, or similar)
  • Experience designing incident response workflows, on-call rotations, and escalation policies
  • Demonstrated ability to troubleshoot complex production issues using observability data
  • Strong Terraform experience for cloud infrastructure and monitoring resources
  • Proficiency with PowerShell scripting
  • Strong experience with Azure cloud and working knowledge of AWS
  • Experience with Azure DevOps for CI/CD pipeline authoring and troubleshooting
  • Experience with Octopus Deploy for deployment management
  • Comfort working across both Windows and Linux server environments
  • Familiarity with Slack for operational workflows and alert routing

Benefits

  • Professional development budget
  • Flexible in-office days
  • Bi-weekly all-company meetings
  • Team offsites, team bonding activities, and happy hours
  • Health, retirement, family forming, and family support benefits
  • Employee giving match
  • Mobile phone stipend
  • R&R days
  • Wellness reimbursement and weekly onsite & virtual programming
  • Generous vacation policy
  • Parental leave and family planning benefits
  • Catered lunches and fully-stocked kitchens

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available