freehire launches on Product Hunt on 26 August.

Follow →

Sr. Site Reliability Engineer

Summary

Senior Site Reliability Engineer to design and maintain resilient, scalable cloud infrastructure for a fintech SaaS platform, ensuring high availability and observability using Python, Terraform, and Azure/AWS.

About the Role

We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services. This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen.

Key Responsibilities

  • Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistently

  • Lead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behavior

  • Build and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortems

  • Architect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code practices and ensure high availability, disaster recovery, and business continuity

  • Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems; implement intelligent alerting to minimize false positives

  • Drive reliability improvements through load testing, chaos engineering, and failure scenario analysis; identify and eliminate single points of failure

  • Partner with application and backend teams to design reliable systems from inception; conduct architecture reviews and reliability assessments

  • Write production-grade Python tooling for automation, metrics collection, alert management, and operational workflows

  • Champion security and compliance in infrastructure; implement defense-in-depth principles for a regulated fintech environment

Required Qualifications

  • 7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systems

  • Expert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scale

  • Demonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog, New Relic, or similar)

  • Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritization

  • Proven experience designing and troubleshooting highly available, resilient, and scalable systems; deep understanding of distributed systems concepts and failure modes

  • Proficiency in Python, PowerShell, bash, etc. scripting languages for production automation, tooling, and systems programming; ability to write clean, maintainable code for operational workflows

  • Hands-on experience with AIOps practices: event correlation, intelligent alerting, predictive analytics, and automated remediation; familiarity with AIOps platforms is a plus

  • Experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Ansible); version control and CI/CD pipeline design

  • Track record of incident management and on-call ownership; comfort with incident response and the ability to remain calm under pressure

  • Excellent communication skills; ability to work cross-functionally and influence without authority; comfort mentoring junior engineers

Preferred Qualifications

  • Experience in the fintech, payments, banking, or other regulated industries; understanding of compliance requirements (SOC 2, PCI-DSS, etc.)

  • Experience with Kubernetes and container orchestration; deep knowledge of containerized application deployment and management

  • Proficiency with observability as code; experience building custom metrics, dashboards, and alerts programmatically

  • Background in chaos engineering or reliability testing; experience using tools like Gremlin or similar platforms

  • Contribution to open-source observability or infrastructure projects

  • Expertise in network security, application security, or infrastructure hardening

  • Experience with database optimization, query performance tuning, and backup/recovery strategies

What this application asks

ashby

Name, Email, Resume

  • Are you eighteen years of age or older? yes / no
  • May MeridianLink contact your CURRENT or MOST RECENT employer? yes / no
  • May MeridianLink Contact your PAST employers? yes / no
  • Have you ever been fired or asked to resign to avoid being fired from a job? yes / no
  • Can you perform the essential functions of this job with or without reasonable accomodation? yes / no
  • Will you now or in the future require sponsorship (i.e. H-1B visa ect.) to legally work in the United States? yes / no
  • Location
  • I understand that as permitted by law, MeridianLink will conduct an investigative consumer report. I understand the scope of this consumer report may include verification of my references, employment history, credit and indebtedness, criminal conviction history, educational and training background and other matters related to my suitability for employment. Upon timely written request to the personnel department of MeridianLink, the nature and scope of the report will be disclosed to me. If I am offered employment with MeridianLink, I understand that the offer will be conditioned on passing a pre-employment drug screen, timely submitting valid documentation that confirms my identity and authorization to work in the United States, and reading and agreeing to comply with the MeridianLink policies as stated in its Employee Handbook. I certify that the answers given by me in this employment application are true, correct and complete. I agree that the company shall not be liable, in any respect, if my employment is terminated because of misstatements or pertinent omissions made by me in this application. Moreover, I understand that all offers of employment are contingent upon passing the company's prescribed credit check and background screen. A copy of this form may be used as the original. The use of results from this form and/or tests will be used for prudent employment decisions. It is agreed and understood that completion of this application does not mean a job opening exists and in no way obligates the company to employ me. In the event of employment, I will comply with all company rules and regulations as established from time to time. I am willing to work all assigned overtime or other special work assignments as requested by the company. I also understand that MeridianLink, Inc. retains the right to amend, modify, add or delete any or all policies or procedures at its sole and absolute discretion. I understand that nothing contained herein is intended to create a contract between the company and me for either employment or the provision of any compensation or benefits. I hereby understand and acknowledge that any employment relationship with this Company is of an “At-Will” nature, which means that the Employee may resign at any time and the Employer may discharge Employee at any time, with or without notice, with or without cause. It is further understood that this ‘At-Will’ employment relationship may not be changed by any written document or by verbal agreement unless such change is specifically acknowledged in writing by an authorized Executive of this Company. During my employment with MeridianLink, Inc. and after my employment ends, I agree not to disclose any confidential or proprietary information regarding operating and trade secrets. CERTIFICATE OF APPLICANT: I certify that all statements made in this application and attachments are true, and I agree and understand that misstatements or omissions of any material fact may be cause for disqualification or dismissal from employment with MeridianLink choose one
  • Please let us know where you heard about this job opening? choose one
  • If above question is "Other" please add the source below.  written answer · optional

See also