Senior Site Reliability Engineer
Summary
Designs and maintains cloud-based platform services for Autodesk’s APIs, focusing on reliability, observability, and automation using AWS, IaC, and SRE principles.
Job Description
Position Overview
We are looking for a passionate skilled Site Reliability Engineer (SRE) to join our platform team in Singapore. Autodesk Platform Services (APS) is a cloud service platform that powers custom and pre-built applications, integrations, and innovative solutions. It offers APIs and web services to unlock the values of our customers' Design and Make data, and connects custom and end-to-end workflows. It is an opportunity to work on the APIs and services that directly impact the millions of users of Autodesk products.
In this hybrid role, you will contribute to the development of the foundational platform APIs, that are the building blocks for next-generation design apps. You will play a pivotal role in designing, developing, and optimizing our cloud-based platform, with a strong focus on Infrastructure as Code (IaC), Application Performance Monitoring (APM), Observability, Continuous Deployment orchestration, incident response and AI-driven automation. You will work in a global organization and collaborate with local and remote colleagues from various disciplines like business, engineering, operations and support. This is an exciting opportunity to unequivocally help our customers build better and custom business solution, within and across disciplines and industries.
Responsibilities
Develop and maintain secure, high-performance cloud services
Collaborate with architects, designers, engineers and key stakeholders to translate requirements into product features and capabilities
Enhance system design and architecture with your cloud expertise throughout the development lifecycle
Improve team processes to meet business needs efficiently
Review services, assess implementations, and recommend improvements
Develop AI-based solutions to boost reliability, efficiency and productivity
Participate in on-call rotations for production support
Minimum Qualifications
BS or MS in Computer Science or related technical field or relevant experience
6 to 10 years of hands-on experience with cloud services and applications
Demonstrated ability to solve problems and work with complex systems
Understanding of data structures, algorithms, and programming
Practical experience with AWS or other major Cloud providers
Proficiency in Infrastructure as Code (IaC) tools such as Terraform or CloudFormation
Knowledge of SRE principles and incident metrics.
Experience designing and maintaining scalable, production-grade platforms and services
Preferred Qualifications
Experience with Cloud Platforms like AWS, Azure or GCP
Familiarity with databases such as MySQL, Redis, DynamoDB
Experience with APM, observability, logging and alerting tool (e.g. Dynatrace, NewRelic, DataDog, Splunk)
Proficient in Python or similar scripting languages
Skilled in building scalable, secure, observable and production-grade cloud services
Experience with AIOps and AI-driven tools for system reliability, automated incident response, and predictive maintenance
Experience working in a Scrum team and Agile setup