Job Description Service Reliability and Monitoring - Monitor availability, performance, and health of Azure-hosted internal applications using tools such as Azure Monitor, Application Insights, Open Telemetry and Log Analytics - Participate in incident response workflows including triage, escalation, and post-incident review with a focus on reducing mean time to recovery (MTTR) - Maintain and refine alerting thresholds and dashboards to surface actionable signals over noise - Contribute to SLO/…