Associate - SRE - Platform Engineering
Associate Platform Reliability Engineer (SRE)
Location: Mumbai / Pune
Role Overview
We are seeking a highly motivated Associate Platform Reliability Engineer (SRE) to join our global Platform Reliability Engineering team. This is a hands-on engineering role focused on the reliability, scalability, and operational excellence of critical front-to-back platforms supporting post-trade processing.
The ideal candidate will have a strong software engineering foundation, production support experience, and a passion for automation, observability, and reliability engineering. You will work closely with development, infrastructure, and business teams to improve system resilience, enhance operational visibility, reduce manual intervention, and deliver highly available services.
Key Responsibilities
- Proactively monitor platform health and drive improvements in reliability, performance, availability, and operational efficiency.
- Perform incident triage, troubleshooting, communication, and post-incident reviews to minimize business impact and prevent recurrence.
- Collaborate with engineering, infrastructure, and business stakeholders to design and implement scalable and resilient solutions.
- Build and enhance deployment and observability capabilities, including dashboards, alerts, and service health monitoring using Grafana, Prometheus, and OpenTelemetry.
- Analyze logs, metrics, and distributed traces to proactively identify system bottlenecks and reliability issues.
- Support enterprise messaging and event-driven architectures, including Kafka-based platforms and integrations.
- Apply capacity planning and availability management best practices to improve platform resilience.
- Develop automation to reduce operational toil, minimize manual intervention, and improve service efficiency.
- Participate in production support, problem management, release management, and change management activities.
Required Qualifications
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline.
- 3+ years of experience in Site Reliability Engineering (SRE), Platform Reliability Engineering (PRE), DevOps, Production Support, or Application Support.
- Strong programming and scripting experience in one or more languages such as Python, Go, or Java.
- Solid understanding of software engineering principles, data structures, algorithms, and system design.
- Strong working knowledge of Linux/Unix and Windows Server environments.
- Experience with modern monitoring and observability practices and tooling.
- Strong understanding of observability and reliability concepts, including metrics, logs, traces, SLIs, SLOs, and alerting.
- Good understanding of event-driven architectures and enterprise messaging platforms such as Kafka and MQ.
- Experience troubleshooting distributed production systems, including APIs, middleware components, and message flows.
- Understanding incident management, problem management, root cause analysis, and operational support processes.
- Familiarity with source control, CI/CD pipelines, Infrastructure as Code (IaC), and DevOps practices.
- Strong verbal and written communication skills with the ability to engage both technical and business stakeholders.
- Self-motivated, detail-oriented, and capable of working independently in a fast-paced environment.
Preferred Qualifications
Observability & Monitoring: Grafana, Prometheus, OpenTelemetry and Loki
DevOps & Automation: Git, Ansible and CI/CD Frameworks
Container & Platform Technologies: Docker, Kubernetes
Data & Messaging Platforms: Kafka, Redis, MQ
Cloud Technologies: AWS