freehire launches on Product Hunt on 26 August.

Follow →

Observability Engineer – Grafana & Prometheus

This role supports the implementation of the AWS-native observability platform as part of the integration modernization program. The engineer will work on configuring dashboards, metrics collection, and alerting frameworks using Grafana, Prometheus, and AWS monitoring tools.
The role focuses on enabling real-time monitoring and operational visibility for APIs, messaging, and integration flows. The engineer will collaborate with platform and engineering teams to ensure system health, identify issues proactively, and support troubleshooting using observability data. The position plays a key role in maintaining reliability and performance across the integration ecosystem.

Observability Platform Engineering

  • Set up and manage Prometheus (metrics collection) and Grafana dashboards (AWS Managed Services)
  • Create dashboards for system health, performance KPIs, and integration monitoring
  • Maintain and enhance dashboards for APIs, messaging, and infrastructure components

Monitoring & Dashboards

  • Monitor application and infrastructure performance across AWS environments
  • Track key metrics such as latency, error rates, throughput, and availability
  • Support performance tuning through observability insights

Alerting & Reliability

  • Configure alerts, thresholds, and proactive monitoring frameworks
  • Enable AI-driven alerting and anomaly detection integrated with ITSM tools
  • Support high availability and reliability through proactive observability practices

Troubleshooting & Operations

  • Support troubleshooting using metrics, logs, traces, and dashboards
  • Provide deep visibility into payload tracking, error handling, and integration health
  • Enable root cause analysis and faster resolution through observability insights

Standardization & Best Practices

  • Establish monitoring standards, observability patterns, and dashboard blueprints
  • Implement structured logging, correlation IDs, and traceability across distributed systems
  • Contribute to observability maturity including automated alerts, runbooks, and analytics

Required Skills & Competencies

  • Strong experience with Grafana and Prometheus
  • Experience setting up monitoring dashboards, alerts, and observability pipelines
  • Basic knowledge of AWS cloud services (CloudWatch, X-Ray, etc.)
  • Solid understanding of monitoring, logging, and alerting concepts
  • Good troubleshooting and analytical skills

Good to Have

  • Knowledge of OpenTelemetry (OTEL) instrumentation
  • Exposure to Kubernetes/EKS monitoring
  • Familiarity with integration monitoring (APIs, messaging, batch flows)
  • Understanding of AI-driven observability and anomaly detection concepts
  • Experience in enterprise integration and cloud modernization programs preferred
  • Understanding of logging and tracing tools

Qualifications

  • Bachelor’s degree in Engineering or related field
  • 6–10 years of experience

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available