freehire launches on Product Hunt on 26 August.

Follow →

Production Support / Site Reliability Engineer- Cloud Native Data Analytics Platform (JD#11280)

Summary

This role involves providing production support and site reliability engineering for a data analytics and reporting platform within the financial sector. The engineer will utilize Google Cloud technologies, observability tools, and AIOps practices to ensure platform stability and performance during trading hours.

Job Summary

We are looking for a proactive and technically strong Production Support / Site Reliability Engineer to join the team supporting the Data Analytics and Reporting Platform. This is a business-critical role combining production support, cloud technology, data analytics, observability, automation, and financial markets knowledge.

Mandatory Skill-set

  • Bachelor's or Master's degree in Computer Science, Mathematics, Finance, Engineering, or a related discipline;
  • Must have 3-4 years of experience in supporting Google Cloud technologies such as Terraform, BigQuery, Cloud Composer, Dataflow and cloud-native data analytics platforms;
  • Understanding of financial markets products, particularly, Rates & Credits, Fixed Income, Repos, Foreign Exchange (FX), Money Markets;
  • Experience with SRE principles, service reliability engineering, observability, and operational resilience;
  • Experience applying AI or machine learning to IT operations, such as AIOps, intelligent alerting, incident summarization, anomaly detection, or automated root-cause analysis
  • Experience with production support, incident management, problem management, and operational processes;
  • Hands-on experience with monitoring and observability tools such as Grafana, Prometheus, ELK/Kibana, or equivalent technologies;
  • Understanding of CI/CD, automation, DevOps/SRE practices, and iterative software delivery;
  • Understanding of ITIL processes and their practical application in a production environment;
  • Knowledge of IT risk and security concepts, including SOx controls, vulnerability management, certificates, authentication, and secure communication protocols;
  • Strong troubleshooting and analytical skills, with the ability to understand complex systems and identify root causes;
  • Experience supporting business-critical platforms during market/trading hours;
  • Strong communication and stakeholder-management skills, with the ability to communicate effectively with both technical and non-technical audiences.

Desired Skill-set

  • Experience applying AI or machine learning to IT operations, such as AIOps, intelligent alerting, incident summarization, anomaly detection, or automated root-cause analysis;
  • Experience developing operational automation using scripting, APIs, or cloud-native services;
  • ITIL certification.

Responsibilities

  • Provide operational support for the Data Analytics and Reporting platform, ensuring reliable service delivery during Asia trading hours;
  • Take ownership of production incidents, coordinating resolution activities and stakeholder communication to minimize business impact;
  • Drive effective problem management by identifying recurring issues, implementing preventive measures, and collaborating with the Product Owner to prioritize improvements through the squad backlog;
  • Contribute to the continuous enhancement of the platform with a strong focus on reliability, performance, security, and operational excellence;
  • Drive automation and AI-assisted operational improvements to enhance platform reliability, efficiency, and observability;
  • Support testing, release, and deployment activities to ensure safe and stable delivery of changes into production;
  • Maintain and continuously improve operational runbooks, knowledge articles, and user guidance documentation to support platform onboarding;
  • You understand the entire stack’s technology on which the application runs and how it fits in the overall chain;
  • Collaborate closely with Trading, Risk, Infrastructure teams and other squads to resolve complex issues and continuously improve platform stability and user experience;
  • Participate in a shared global standby rotation (typically one week per month), contributing to the reliability and continuity of business-critical services;
  • As part of a cross-border Squad work in an Agile/Scrum way, on the backlog prioritized by a Product Owner, and demonstrate your features/stories to other colleagues and the stakeholders

Should you be interested in this career opportunity, please send in your updated resume to apply@sciente.com at the earliest.

When you apply, you voluntarily consent to the disclosure, collection and use of your personal data for employment/recruitment and related purposes in accordance with the SCIENTE Group Privacy Policy, a copy of which is published at SCIENTE’s website (https://www.sciente.com/privacy-policy).

Confidentiality is assured, and only shortlisted candidates will be notified for interviews.

EA Licence No. 07C5639

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available