Monitoring And Observability Engineer
Description
The engineer is responsible for monitoring and maintaining the operational health of the private cloud platform. To operate a 24x7 managed monitoring service using Dynatrace and ManageEngine, you will ensure high availability and performance across 100+ applications and 1,000+ infrastructure assets, maintaining strict SLA compliance.
Responsibilities
- Full-stack observability: Manage Dynatrace APM to monitor application performance, transaction tracing, and user experience for 100+ apps and 12 databases.
- Infrastructure oversight: Utilize ManageEngine to monitor 100 network devices and 1,000+ hosts, tracking health metrics like CPU, latency, and packet loss.
- Dependency mapping: Use Dynatrace Smartscape to map digital asset relationships and perform impact analysis during incidents.
- 24x7 NOC operations: Provide continuous alert triage and escalation management within a rotating shift model.
- Performance governance: Generate real-time dashboards and detailed root cause analysis (RCA) reports to drive continuous service improvement.
Requirements
- Bachelor’s degree or diploma in engineering, computer science, or similar (preferred).
- Experience with monitoring tools, specifically Dynatrace and/or ManageEngine monitoring suites.
- Technical depth: strong understanding of application topology, network protocols, and data center infrastructure including servers, storage, and networking.
- Previous experience in 24x7 managed services or NOC environment (preferred).
- Process knowledge: experience with ITIL frameworks, incident management, and meeting strict SLA targets.
- Communication: ability to translate technical telemetry into clear performance reports and dashboards.