Data Center Engineering Operations Facility Manager
Data Center Engineering Operations Facility Manager
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we're the people who keep the cloud running. We support all AWS data centers and all the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we're looking for talented people who want to help.
You'll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You'll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you'll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.
Key job responsibilities
- Oversee the day-to-day operations and maintenance of all mechanical, electrical, fire protection, and control systems, ensuring safety, security, availability, and performance across the data center.
- Lead and develop teams of 24x7 Engineering Operations Technicians and Chief Engineer, covering all aspects of people management, performance, career growth, and technical management.
- Manage and coordinate planned and reactive maintenance activities, ensuring minimal disruption to data center operations.
- Develop and review operation documentation including MOP, SOP, and EOP where applicable.
- Engage in expansion and improvement projects from conception to completion, coordinating across internal support teams and external stakeholders.
- Oversee third-party vendor relationships, setting high performance standards, driving accountability, and ensuring adherence to contracted SLAs through regular performance reviews.
- Manage CapEx and OpEx effectively to suit operational needs.
- Serve as the primary point of contact for the data center, manage Large Scale Events (LSEs) or outages, and maintain readiness through frequent and effective drills.
- Lead incident investigations for unplanned events, ensuring root cause analysis and corrective actions are documented, tracked, and closed.
- Support audits, inspections, and compliance reviews conducted by internal or external parties.
About the team
About AWS
AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.
Why AWS?
Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.
Inclusive Team Culture
Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empowers us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (diversity) conferences, inspire us to never stop embracing our uniqueness.
Mentorship & Career Growth
We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.
Work/Life Balance
We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.
Bachelor's degree in Electrical Engineering, Mechanical Engineering, or a related field
Knowledge of the electrical and mechanical systems involved in critical data center operations including systems such as feeders, transformers, generators, switchgear, UPS systems, ATS units, PDU units, chillers, pumps, air handling units, and CRAC units
8+ years of critical facility engineering and operations experience, 5+ years in people management experience with direct reports as their performance manager
A valid Registered Electrical Worker (REW) Certificate (Grade B0 or above)
In-depth experience in change, problem and incident management
Demonstrated communication skills in presentation, open discussion, team coaching
Fluent verbal and written proficiency in Chinese and English
Financial understanding in annual budgeting, Capex and Opex concepts
Preferred qualifications
A holder of Grade C0 Registered Electrical Worker (REW) Certificate
Experience in major facilities incident management
Experience in data center Safety and Security measures
Knowledge of Data Center capacity planning and KPI reporting
Experience of critical maintenance works, including LV WR2, HV WR2, Pull the Plug test
Knowledge of IT services such as servers, network, cabling and platform
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.