Site Reliability Engineer/SRE/DevOps/L2 Support (Large-scale Consumer Platform, Engineering-dri[...]
DADACONSULTANTS PTE. LTD. Site Reliability Engineer/SRE/DevOps/L2 Support (Large-scale Consumer Platform, Engineering-dri[...]
Summary
Ensure high availability and reliability for large-scale consumer platforms by managing incidents, capacity planning, and automation using Python, Go, and cloud infrastructure.
My client:
My client is a well-established technology company building large-scale, consumer-facing digital platforms used by millions of users globally. The team is highly engineering-driven, with strong emphasis on system reliability, automation, and operational excellence. This role offers the opportunity to work on production-critical systems, collaborate closely with backend and platform engineers, and take real ownership of reliability and infrastructure in a fast-moving, high-impact environment.
Job Responsibilities:
- Ensure high availability and reliability of production systems supporting large-scale online services.
- Own core operational functions including resource management, incident response, capacity planning, monitoring, and reliability improvements.
- Review system and infrastructure designs, proactively identifying potential reliability and scalability risks.
- Analyze system bottlenecks and structural weaknesses, and lead optimization initiatives to improve stability and cost efficiency.
- Participate in 24/7 on-call rotations, responding quickly to production incidents and driving root cause analysis.
- Continuously improve operational processes through automation, reducing manual intervention and operational overhead.
Job Requirements:
- Bachelor's degree in Computer Science or a related technical field.
- Proficiency in Python, Go, or Shell scripting, with the ability to independently build tools or platforms.
- Solid experience with cloud infrastructure; exposure to multi-cloud or hybrid environments is a plus.
- Strong fundamentals in Linux systems, networking, load balancing, and high-availability / disaster recovery design.
- Hands-on experience operating or supporting large-scale web services in production environments.
- Self-motivated, responsible, and comfortable working in a fast-paced, reliability-critical environment.
What They Offer
- A large-scale, production-grade technical environment supporting high-traffic, consumer-facing platforms.
- A stable yet fast-evolving engineering environment with long-term technical depth and learning potential.
- Close collaboration with experienced engineers in a flat, open, and execution-focused team culture.
About Us
Dada Consultants was established in 2017, with the commitment of providing the best recruitment services in Singapore. We are comprised of a dynamic head-hunting team dedicated to sourcing for highly competent professionals in IT industry. We provide enterprises with customized talent solutions, and bring talents to career advancement.