Site Reliability Engineer – APM, Dynatrace, Observability
Location: Toronto, ON (Hybrid – 4 Days Onsite)
Duration: 12 Months
Experience: 6–8 Years
Required Skills
- Kubernetes & Containers (Kubernetes 1.24+, CRDs, Operators)
- Prometheus, Grafana, Thanos/Cortex/Mimir, Alertmanager
- Loki/ELK/Splunk, OpenTelemetry, Jaeger/Tempo
- GitOps (ArgoCD/FluxCD), Helm, Terraform
- AWS/Azure/GCP, S3/Object Storage
- PromQL, LogQL/Lucene, SQL
- Python, Bash or Go
- CI/CD (GitHub Actions, GitLab CI, Jenkins)
Key Responsibilities
- Design, deploy, and manage enterprise observability platforms across Kubernetes environments.
- Build and maintain monitoring, logging, tracing, dashboards, and alerting solutions.
- Automate deployments using GitOps and Infrastructure as Code.
- Optimize observability performance, scalability, and cost.
- Integrate observability with CI/CD pipelines, cloud platforms, and incident management tools.
- Collaborate with SRE, DevOps, and application teams to improve platform reliability and reduce MTTR.
Preferred Skills
- AI/ML for Observability, AIOps, LLM integration
- Kubernetes Certifications (CKA/CKAD/CKS)
- Grafana Certification
- Experience with SLIs, SLOs, Error Budgets
- Financial Services or regulated industry experience
- APM tools and Service Mesh observability