Data Center IT management Engineer
Responsibilities
Team Insight: The DataCenter Service / Datacenter Cloud System (DCS) team sits within TikTok’s global technology structure and supports the company’s fast growth by building and operating hyper-scale datacenters, managing the life cycle of server fleet, providing cloud solutions, and developing various infrastructure services, making sure they are scalable and are reliable.
Role Insight: We are seeking an experienced Data Center Operations Engineers to apply technical expertise in a dynamic, fast-paced environment. This role requires strong knowledge of server hardware and a foundational understanding of mechanical and electrical infrastructure in large-scale data centers. You will be responsible for diagnosing and resolving server and infrastructure issues, collaborating with remote teams, and supporting the full server rack lifecycle, including buildout of compute and storage environments. In addition, you will lead and guide DC Technicians, ensuring high-quality execution of operational tasks and adherence to best practices. Candidates should have hands‑on experience in at least one of the following areas: Networking, Scripting, or Hardware Repair. Success in this role requires strong communication skills, the ability to work independently and within a team, and adaptability in a rapidly changing environment.
Responsibilities
- Own day-to-day IT infrastructure operations across assigned data centre sites, ensuring uptime and SLA compliance for all break-fix activity.
- Lead daily operations by guiding Data Center Technicians and overseeing task allocation, hardware troubleshooting, maintenance, and diagnostics.
- Manage the full incident and change lifecycle, resolving incidents to SLA, executing changes under formal change control, and leading root cause analysis and corrective actions to closure.
- Track and report IT operational performance metrics, including availability, MTTR, ticket ageing, first-time fix, break-fix volumes, and rack capacity utilisation, using trend analysis to drive continuous improvement.
- Manage data centre IT on‑site stability, identifying risk events and driving improvement initiatives to enhance stability and reduce operational errors.
- Support data centre projects and operational improvements, including capacity expansions, retrofits, infrastructure upgrades and implementation of new tools and processes; support new site builds and operational handover readiness.
- Maintain and improve site documentation (SOPs, MOPs, EOPs, escalation matrices), build strong cross‑functional relationships, act as an escalation point, and identify recurring issues to drive vendor escalation and continuous improvement.
- Willingness to participate in on-call rotation and travel to other sites as required.
Qualifications
Minimum Qualifications
- 5+ years of experience in data centre operations, IT infrastructure, hardware support, and structured cabling in a live data centre environment.
- Strong hands‑on knowledge of server, storage, and network hardware, including component‑level diagnostics and replacement.
- Capable of server disassembly, hardware replacement, rack deployment, mandatory knowledge etc.
- Basic Linux OS administration and BMC/BIOS configuration skills.
- Basic understanding of critical facility infrastructure (UPS, generators, PDUs, CRAC/CRAH, cooling topology) to work safely alongside facilities teams.
- Competent with DCIM, ticketing, and monitoring platforms, with the ability to analyse ticket data to identify recurring faults and support IT infrastructure performance.
- Clear written and verbal communication skills, with the ability to report to both technical and non‑technical stakeholders.
Preferred Qualifications
- Relevant certification such as CDCP, CDCS, ITIL Foundation, CompTIA Server+, or CCNA.
- Experience in large‑scale data centre environments.
- Strong working knowledge of data centre power infrastructure and cooling systems, with proven competence in identifying server overheating risks and responding to power emergencies.
- Hands‑on experience operating and maintaining GPU servers and integrated rack systems, including hardware troubleshooting on GB300, B200, or comparable GPU platforms (e.g., GPU error analysis, NVLink fault diagnosis).
- Experience coordinating vendors, outsourced technicians, or smart‑hand resources on site.
- Experience supporting projects, incident management, or small team leadership.
- Työsuhde: Työsuhde
- Palkka: Muu Sopimuksen mukaan
- Työskentelyaika: Päivätyö, Arkisin
- Paikkoja: 1
- Lähde: Työmarkkinatori
Osaamisalue: IT & tietoliikenne, Tutkimus, kehitys & tuotekehitys