freehire launches on Product Hunt on 26 August.

Follow →

Site Reliability Engineer (SRE) – Dynatrace & AI Observability

Site Reliability Engineer

Location: Toronto, ON
Work Model: Hybrid (2 days per week in-person at the Toronto office preferred)
Required Skills

• Site Reliability Engineering (SRE)
• DevOps
• Dynatrace

Role Summary

• Design, implement, and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability, performance, and reliability across distributed environments.
• Leverage Dynatrace, Davis AI, automation, and cloud technologies to enable proactive monitoring, intelligent automation, and operational excellence.

Role Description

Dynatrace & AI-Driven Observability

• Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.
• Utilize Dynatrace Davis AI for automated root cause analysis, anomaly detection, event correlation, predictive performance insights, and alert noise reduction.
• Configure and manage OneAgent deployments, Smartscape topology mapping, service flow, and distributed tracing.
• Define and monitor SLIs, SLOs, and user experience metrics.
• Build custom dashboards, alerts, and observability pipelines.
• Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.
• Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.
• Enable self-healing automation using Dynatrace event triggers and AI-driven insights.

Automation & Configuration Management

• Design and implement automation solutions using Ansible.
• Automate configuration management, application deployments, and environment provisioning.
• Develop reusable Ansible playbooks and roles for scalable operations.
• Automate operational tasks, patching, compliance processes, and remediation workflows.
• Integrate Ansible with CI/CD pipelines and monitoring systems.

Cloud & DevOps

• Design and manage cloud-native solutions on AWS with exposure to Azure.
• Develop infrastructure using Terraform, CloudFormation, or AWS CDK.
• Build and manage CI/CD pipelines using GitHub Actions, Jenkins, or GitLab CI.
• Develop and deploy serverless solutions using AWS Lambda, API Gateway, and Step Functions.
• Automate DevOps and operational workflows using Python (boto3) and Bash scripting.
• Deploy and maintain production environments through automated pipelines.
• Optimize cloud infrastructure for cost, performance, and scalability.

Monitoring & Reliability Engineering

• Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.
• Design unified observability across multi-cloud environments.
• Implement logging and distributed tracing strategies.
• Work with Docker, Kubernetes, ECS, and AKS environments.
• Design fault-tolerant, highly available, and disaster recovery solutions.
• Support incident response, on-call activities, and root cause analysis (RCA).

Required Qualifications

• Proven experience with Dynatrace APM, Real User Monitoring (RUM), and infrastructure monitoring.
• Strong hands-on experience with Dynatrace Davis AI capabilities.
• Experience with Ansible for automation and configuration management.
• Deep knowledge of AWS services and cloud-native architectures.
• Experience with Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
• Proficiency in Python (boto3) and Bash scripting.
• Experience supporting production-scale environments.
• Business Analyst experience.
• Scrum Master experience.

Nice to Have

• Dynatrace Associate or Professional certification.
• Experience with Dynatrace APIs and automation.
• Experience building self-healing systems using AI-driven triggers.
• Familiarity with Prometheus, Grafana, and the ELK Stack.
• Azure cloud experience and certifications.
• Experience with GitOps and Platform Engineering.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available