freehire launches on Product Hunt on 26 August.

Follow →

Senior Site Reliability Engineer

Summary

The Senior Site Reliability Engineer will design and maintain observability tooling, define SLOs, and lead incident responses for debit card services. The role involves automating operational tasks and partnering with engineering teams to ensure system reliability and scalability using technologies like Kubernetes, Prometheus, and Go.

We are looking for a Senior Site Reliability Engineer to join our team. You'll work closely with backend engineers, product managers, and other SREs to identify risk before it becomes an incident — and when incidents do happen, you'll help lead the response and the follow-up.

Responsibilities

  • Design, build, and maintain monitoring, alerting, and observability tooling, including metrics, logs, and traces, for debit card services
  • Build and maintain dashboards that surface availability, latency, error rates, and other key reliability signals to engineering and leadership
  • Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical debit card flows, such as authorization, settlement, card issuance, and disputes
  • Monitor system availability and reliability on an ongoing basis, proactively identifying degradation trends before they cause customer-facing impact
  • Participate in on-call rotations, leading or supporting incident response, root cause analysis, and blameless postmortems
  • Partner with product and backend engineering teams to review architecture for reliability, scalability, and fault tolerance
  • Automate manual operational work through tooling and scripting to reduce repetitive tasks and operational toil
  • Conduct capacity planning and performance testing to ensure systems scale effectively with transaction volume
  • Improve deployment safety through canary releases, rollback automation, and progressive delivery practices
  • Contribute to and enforce reliability best practices, runbooks, and operational documentation across the team

Requirements

  • A minimum of 3 years of relevant experience as a Site Reliability Engineer, DevOps Engineer, or Production/Infrastructure Engineer
  • Hands-on experience with monitoring and observability tools such as Grafana, Prometheus, Datadog, or M3/Uber's internal metrics stack
  • Strong understanding of SLOs, SLIs, error budgets, and other reliability engineering principles
  • Proficiency in at least one programming language commonly used for tooling and automation, such as Go, Python, or Java
  • Experience with distributed systems and a solid understanding of failure modes in high-throughput, low-latency environments
  • Familiarity with container orchestration and infrastructure, such as Kubernetes and Docker, along with cloud or on-prem infrastructure at scale
  • Experience with incident management processes, including on-call response, root cause analysis, and postmortems
  • Strong scripting and automation skills for reducing operational toil, using tools such as Bash, Python, or similar languages
  • Working knowledge of CI/CD pipelines and safe deployment practices, including canary, blue-green, and rollback strategies
  • Excellent communication skills, with the ability to translate system health data into clear insights for both engineers and non-technical stakeholders
  • Excellent English communication skills (B2 level or higher)

Nice to have

  • Experience in payments, fintech, or other high-compliance, transaction-critical environments
  • Familiarity with PCI-DSS or other financial services compliance and security requirements
  • Experience with chaos engineering or fault-injection testing
  • Background in database reliability, including query performance, replication, and failover, for transactional systems
  • Experience building or maintaining internal tooling and platforms for observability at scale

Benefits

  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available