Point your AI agent at freehire and let it find you a job.

Get the CLI →

Kotak Mahindra Bank

NewBe an early applicant

Tech Ops Engineering II-SUPPORT SERVICES-CTO - In House Engineering

Posted Updated
Discussion

Job Title: Production Site Reliability Engineer (SRE) – Digital Payments

Role Overview

We are seeking a highly technical and driven Production SRE Engineer to manage and monitor mission-critical payment platforms including UPI, IMPS, and Payment Hub systems & various Payment Applications.

The role focuses on ensuring high availability, low latency, and seamless transaction experience for customers. The incumbent will collaborate with cross-functional teams (Engineering, Business, Compliance) and external regulators (RBI, NPCI) to maintain resilient and scalable payment infrastructure.

Key Responsibilities

Production Support & Incident Management

  • Provide L2/L3 production support for UPI, IMPS, and Payment Hub platforms & various Payment Applications.
  • Diagnose, triage, and resolve transaction failures, timeouts, and API disruptions.
  • Lead and participate in Major Incident Management (MIM) calls and ensure timely stakeholder communication.
  • Manage incidents, service requests, and problem tickets via Jira, ServiceNow.
  • Provide regular updates to internal stakeholders and regulatory bodies (NPCI/RBI) during critical issues.

Reliability Engineering & RCA

  • Perform deep-dive Root Cause Analysis (RCA) for recurring payment and system issues.
  • Implement preventive and corrective measures to improve system stability.
  • Drive SRE best practices including error budgets, SLIs/SLOs, and system resilience.
  • Experience in managing DR Drills & Documentations.
  • Reviewing the SOPs & its relative documentations.

Monitoring, Observability & System Engineering

  • Monitor key performance indicators:
    • Transaction success rates
    • Latency and response times
    • Failure trends and retries
  • Build and maintain dashboards using:
    • ELK Stack, Grafana, Kibana, Splunk, Datadog, Prometheus
  • Establish proactive alerting and anomaly detection mechanisms.
  • Work closely with engineering teams to design and optimize:
    • High-throughput payment switches
    • Routing logic
    • Settlement and reconciliation systems
  • Understand and support UPI architecture, IMPS rails, and payment orchestration layers & various Payment Applications.
  • Trace end-to-end transaction lifecycle across distributed systems.

External Partner & Regulatory Coordination

  • Coordinate with NPCI, partner banks, and TPAPs during outages, reconciliation issues, or network disruptions.
  • Lead integrations and ensure seamless onboarding of ecosystem participants.
  • Ensure compliance with:
    • RBI guidelines and data localization mandates
    • NPCI operational and technical standards

Technical Skills & Expertise

Payments Domain Knowledge

  • Strong expertise in:
    • UPI architecture and flows
    • IMPS rails
    • Payment gateway / switch systems
    • Payment Hub orchestration & various Payment Applications.

Core Technical Skills

  • Advanced SQL proficiency (joins, aggregations, stored procedures)
  • Strong hands-on experience in:
    • Linux/UNIX systems administration
    • Shell scripting
  • Ability to:
    • Read , Write and interpret All types documentation (SOPs, workflows, etc.)
    • Understand database schemas
    • Analyse system architecture and latency

Monitoring & Observability Tools

  • Hands-on expertise with:
    • ELK Stack (Elasticsearch, Logstash, Kibana)
    • Grafana, Prometheus
    • Splunk, Datadog

DevOps & Cloud

  • Experience with:
    • CI/CD pipelines, Containerization (Docker, Kubernetes)
  • Cloud platforms:
    • AWS / GCP / Azure

Key Competencies

  • Strong problem-solving and analytical skills
  • High ownership in production environments
  • Ability to work under pressure in real-time systems
  • Strong stakeholder communication and coordination
  • Focus on reliability, scalability, and performance
  • Must have can do, takes initiative, Drives end to end deliverables.
  • Proactive, solution-oriented mindset with ownership to resolve production issues under pressure.
  • Ability to clearly articulate incidents, updates, and RCA to stakeholders, leadership, and regulators.
  • Works effectively with cross-functional teams (engineering, product, partners, regulators).
  • Structured thinking to diagnose complex system failures and drive long-term fixes.
  • Ability to stay calm and effective during high-severity incidents and critical outages.
  • Knowledge of PCI-DSS compliance, Financial data governance & security best practices
  • Quickly adapts to changing technologies, incidents, and regulatory requirements in a fast-evolving payments ecosystem.
  • Precision in analysing logs, transactions, and system behaviour to avoid critical errors in production.
  • Effectively manage multiple incidents, tasks, and escalations in a high-pressure environment.
  • Ability to handle expectations and coordinate with internal teams, partners, and regulators efficiently.
  • Takes ownership to make quick, informed decisions during outages or critical production incidents & communications to various Stake holders including Regulatory.

Skills

See also

Support jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available