freehire launches on Product Hunt on 26 August.

Follow →

Sr. Observability Engineer – Grafana & Prometheus

This role is part of the large-scale integration modernization program, focused on migrating WebMethods-based integrations to AWS-native architecture. The engineer will design and implement a centralized observability platform leveraging AWS Managed Grafana, Prometheus, CloudWatch, and OpenTelemetry across APIs, messaging, MFT, and SAP integrations.
The engineer will contribute to building unified dashboards, SLA monitoring, and AI-driven alerting to provide real-time insights into system health, integration flows, and business KPIs. The role is critical to ensuring reliability, troubleshooting efficiency, and operational intelligence across ~100M monthly transactions and ~1,500 flows.

Observability Platform Engineering

  • Set up and manage Prometheus (metrics collection) and Grafana dashboards (AWS Managed Services)
  • Build and maintain AWS-native observability platform including metrics, logs, and traces using CloudWatch, X-Ray, and OTEL instrumentation
  • Implement instrumentation across APIs, messaging, MFT, and batch integrations for end-to-end visibility

Monitoring & Dashboards

  • Monitor application and infrastructure performance across AWS integration estate
  • Create dashboards for KPIs, system health, SLA tracking, and integration performance
  • Develop master dashboards covering API performance, messaging queues, infra health, and integration status

Alerting & Reliability

  • Configure alerts, thresholds, and proactive monitoring frameworks
  • Enable AI-driven alerting and anomaly detection integrated with ITSM tools
  • Support high availability and reliability through proactive observability practices

Troubleshooting & Operations

  • Support troubleshooting using metrics, logs, traces, and dashboards
  • Provide deep visibility into payload tracking, error handling, and integration health
  • Enable root cause analysis and faster resolution through observability insights

Standardization & Best Practices

  • Establish monitoring standards, observability patterns, and dashboard blueprints
  • Implement structured logging, correlation IDs, and traceability across distributed systems
  • Contribute to observability maturity including automated alerts, runbooks, and analytics

Required Skills & Competencies

  • Strong experience with Grafana and Prometheus
  • Experience setting up monitoring dashboards, alerts, and observability pipelines
  • Basic knowledge of AWS cloud services (CloudWatch, X-Ray, etc.)
  • Solid understanding of monitoring, logging, and alerting concepts
  • Good troubleshooting and analytical skills

Good to Have

  • Knowledge of OpenTelemetry (OTEL) instrumentation
  • Exposure to Kubernetes/EKS monitoring
  • Familiarity with integration monitoring (APIs, messaging, batch flows)
  • Understanding of AI-driven observability and anomaly detection concepts
  • Experience in enterprise integration and cloud modernization programs preferred

Qualifications

  • Bachelor’s degree in Engineering or related field
  • 8–12 years of experience

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available