Site Reliability Engineer (SRE) – Dynatrace & AI Observability
Site Reliability Engineer
Location: Toronto, ON
Work Model: Hybrid (2 days per week in-person at the Toronto office preferred)
Required Skills
• Site Reliability Engineering (SRE)
• DevOps
• Dynatrace
Role Summary
• Design, implement, and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability, performance, and reliability across distributed environments.
• Leverage Dynatrace, Davis AI, automation, and cloud technologies to enable proactive monitoring, intelligent automation, and operational excellence.
Role Description
Dynatrace & AI-Driven Observability
• Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.
• Utilize Dynatrace Davis AI for automated root cause analysis, anomaly detection, event correlation, predictive performance insights, and alert noise reduction.
• Configure and manage OneAgent deployments, Smartscape topology mapping, service flow, and distributed tracing.
• Define and monitor SLIs, SLOs, and user experience metrics.
• Build custom dashboards, alerts, and observability pipelines.
• Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.
• Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.
• Enable self-healing automation using Dynatrace event triggers and AI-driven insights.
Automation & Configuration Management
• Design and implement automation solutions using Ansible.
• Automate configuration management, application deployments, and environment provisioning.
• Develop reusable Ansible playbooks and roles for scalable operations.
• Automate operational tasks, patching, compliance processes, and remediation workflows.
• Integrate Ansible with CI/CD pipelines and monitoring systems.
Cloud & DevOps
• Design and manage cloud-native solutions on AWS with exposure to Azure.
• Develop infrastructure using Terraform, CloudFormation, or AWS CDK.
• Build and manage CI/CD pipelines using GitHub Actions, Jenkins, or GitLab CI.
• Develop and deploy serverless solutions using AWS Lambda, API Gateway, and Step Functions.
• Automate DevOps and operational workflows using Python (boto3) and Bash scripting.
• Deploy and maintain production environments through automated pipelines.
• Optimize cloud infrastructure for cost, performance, and scalability.
Monitoring & Reliability Engineering
• Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.
• Design unified observability across multi-cloud environments.
• Implement logging and distributed tracing strategies.
• Work with Docker, Kubernetes, ECS, and AKS environments.
• Design fault-tolerant, highly available, and disaster recovery solutions.
• Support incident response, on-call activities, and root cause analysis (RCA).
Required Qualifications
• Proven experience with Dynatrace APM, Real User Monitoring (RUM), and infrastructure monitoring.
• Strong hands-on experience with Dynatrace Davis AI capabilities.
• Experience with Ansible for automation and configuration management.
• Deep knowledge of AWS services and cloud-native architectures.
• Experience with Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
• Proficiency in Python (boto3) and Bash scripting.
• Experience supporting production-scale environments.
• Business Analyst experience.
• Scrum Master experience.
Nice to Have
• Dynatrace Associate or Professional certification.
• Experience with Dynatrace APIs and automation.
• Experience building self-healing systems using AI-driven triggers.
• Familiarity with Prometheus, Grafana, and the ELK Stack.
• Azure cloud experience and certifications.
• Experience with GitOps and Platform Engineering.