freehire launches on Product Hunt on 26 August.

Follow →

SRE Engineer

Job Description & Requirements

About Bounteous

Bounteous is a leading global digital engineering and technology solutions provider, trusted by top-tier organizations across Banking, Financial Services & Insurance (BFSI). With a combined strength of more than 5,000 engineers worldwide, we specialize in delivering high-impact solutions across Capital Markets, Core Banking, Payments, Digital Transformation, Data & AI, Cloud Engineering, and modern enterprise platforms.

In Singapore, we partner closely with major regional and global financial institutions, including banks, asset managers, and market infrastructure providers, to drive advanced engineering, modernization, and AI-led transformation. Our teams bring deep expertise across Site Reliability Engineering (SRE), Cloud-Native Platforms, DevSecOps, Infrastructure Automation, Observability, AI Engineering, and Enterprise Operations, supported by a strong delivery presence across APAC, India, Europe, and North America.

About the Role

We are seeking an experienced Site Reliability Engineer (SRE) to support the implementation, operation, and continuous improvement of an enterprise-grade Agentic AI Security Control Platform (MSAgentShield).

The successful candidate will play a key role in ensuring platform reliability, availability, observability, monitoring, automation, incident management, and operational excellence across a high-availability production environment. This role requires strong expertise in monitoring, alerting, log analytics, production support, and operational engineering practices.

Key Responsibilities

  • Design, implement, and maintain monitoring, observability, and alerting frameworks for mission-critical production systems.
  • Develop and optimize dashboards, metrics, logging, and alerting solutions using tools such as Grafana, Splunk, Datadog, or equivalent platforms.
  • Monitor system health, platform availability, and performance to proactively identify and resolve issues.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational KPIs.
  • Participate in production incident management, root cause analysis (RCA), and post-incident reviews.
  • Support and enhance enterprise log analytics and monitoring solutions.
  • Develop automation scripts and tooling to reduce operational overhead and improve platform reliability.
  • Maintain and improve CI/CD pipelines and deployment processes.
  • Collaborate with engineering, security, infrastructure, and cloud teams to ensure operational readiness.
  • Create and maintain runbooks, standard operating procedures, troubleshooting guides, and escalation processes.
  • Support platform upgrades, maintenance activities, disaster recovery testing, and operational resilience initiatives.
  • Contribute to continuous improvement initiatives focused on reliability, scalability, and service quality.

Required Skills & Experience

  • Strong experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Infrastructure Engineering, or Production Support Engineering.
  • Solid hands-on experience with Linux administration, shell scripting, Git, and CI/CD pipelines.
  • Working knowledge of Docker and Kubernetes in production environments.
  • Strong expertise in observability, monitoring, dashboarding, and alerting.
  • Hands-on experience with Grafana and enterprise monitoring platforms.
  • Experience with Splunk, Datadog, ELK Stack, Prometheus, OpenTelemetry, or similar observability solutions.
  • Experience defining and managing SLOs, SLIs, alert thresholds, and operational metrics.
  • Strong incident response, troubleshooting, and root cause analysis experience.
  • Basic to intermediate Python scripting for automation and operational tooling.
  • Experience supporting high-availability production systems.
  • Strong analytical, problem-solving, and communication skills.
  • Experience working in Agile delivery environments.

Preferred Qualifications

  • Bachelor's Degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • Experience supporting large-scale cloud-native environments on AWS, Azure, or GCP.
  • Experience with log analytics and enterprise monitoring ecosystems.
  • Exposure to AI/ML platforms, agentic AI solutions, or security platforms.
  • Experience implementing Infrastructure as Code (Terraform, Ansible, etc.).
  • Experience with ITIL-based operational processes.
  • Familiarity with security monitoring, SIEM platforms, and DevSecOps practices.
  • Experience working in global enterprise environments.

What's on Offer

You will join a high-caliber engineering environment where reliability, automation, observability, and operational excellence are critical to success. This role offers the opportunity to work on a next-generation Agentic AI Security Platform supporting enterprise-scale production environments.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available