Senior Observability Engineer / Platform Engineer

Summary

This role involves designing and maintaining enterprise-grade observability platforms using tools like Prometheus, Grafana, and OpenTelemetry to ensure system reliability. The engineer will work across cloud-native environments, focusing on metrics, logs, traces, and automation to improve incident detection and troubleshooting.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Observability Engineer / Platform Engineer based in India.

This is a remote opportunity for an experienced engineer to shape and operate enterprise-grade observability platforms across cloud-native environments. You will work across metrics, logs, traces, events, alerting, and service health to strengthen the reliability of critical production systems. The role combines hands-on engineering with platform enablement, automation, and operational excellence. You will support Kubernetes and public cloud environments while helping engineering and SRE teams gain better visibility into system performance and reliability. Your work will directly contribute to proactive incident detection, faster troubleshooting, and improved availability. You will also have the opportunity to contribute to evolving practices around distributed tracing, SLOs, intelligent alerting, and AIOps. This is a permanent remote role designed for professionals who enjoy solving complex platform and reliability challenges.

Accountabilities:

  • Design, implement, and maintain enterprise observability platforms covering metrics, logs, traces, and events.
  • Build and manage observability solutions using platforms such as Prometheus, Grafana, OpenSearch, Splunk, Elastic, Datadog, Dynatrace, New Relic, or equivalent technologies.
  • Develop dashboards, SLOs, SLIs, alerting rules, and service health monitoring frameworks to provide actionable visibility into production environments.
  • Integrate monitoring and observability capabilities across Kubernetes, containerized workloads, and public cloud platforms.
  • Enable effective incident management, root cause analysis, and production troubleshooting through robust observability practices.
  • Automate monitoring configuration, platform onboarding, and operational processes using Infrastructure-as-Code and CI/CD pipelines.
  • Partner with engineering, platform, SRE, DevOps, and operations teams to improve system reliability, performance, scalability, and availability.
  • Contribute to observability maturity initiatives, including distributed tracing, OpenTelemetry, AIOps, intelligent alerting, and automated remediation.
  • Requirements:

    • 6+ years of experience in Platform Engineering, SRE, DevOps, Cloud Operations, Observability Engineering, or a closely related discipline.
    • Strong hands-on expertise with observability technologies such as Prometheus, Grafana, Splunk, OpenSearch, Elastic, or equivalent platforms.
    • Experience with distributed tracing solutions such as Jaeger, Tempo, and OpenTelemetry, with a solid understanding of modern observability practices.
    • Strong knowledge of Kubernetes, Docker, and containerized workloads, ideally gained through enterprise-scale production environments.
    • Practical experience working with AWS, Azure, or GCP cloud environments.
    • Demonstrated ability to manage incidents, tune alerts, troubleshoot complex production issues, and improve operational reliability.
    • Strong scripting and automation capabilities using Python, Shell, or similar languages.
    • Familiarity with CI/CD and Infrastructure-as-Code tools such as Terraform, Jenkins, GitHub Actions, or ArgoCD.
    • Understanding of SRE practices, including SLO/SLI frameworks and error budgets, is preferred.
    • Exposure to AIOps, intelligent incident response, automated remediation, and OpenTelemetry implementations is an advantage.
    • Experience working with enterprise-scale production platforms and an awareness of security, compliance, and governance considerations in cloud-native environments.
    • Strong collaboration, communication, analytical, and problem-solving skills, with the ability to work effectively across engineering and operations teams.
    • Benefits:

      • Permanent remote working model, with the role based in India.
      • Flexible working arrangements designed to accommodate employee, customer, and business needs.
      • Opportunities to work on enterprise-scale cloud, platform engineering, and observability initiatives.
      • Exposure to modern technologies and practices across observability, Kubernetes, cloud platforms, SRE, automation, and AIOps.
      • An inclusive and diverse working environment that values different perspectives, skills, and experiences.
      • Flexibility around working hours and arrangements where business and customer requirements allow.
      • Supportive initiatives for professionals returning to work after an extended career break due to health or family circumstances.
      • Opportunities for professional development, meaningful ownership, and long-term career growth.
      • Salary: Competitive compensation aligned with experience and market standards.
      • Healthcare and additional perks: Details to be confirmed during the hiring process.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available