Infrastructure Engineer
You will diagnose and repair rack-level GPU hardware, investigate faults using system and kernel logs and BMC Redfish APIs, and work with engineering specialists when additional diagnostic data is needed. You will automate diagnostics, provisioning, and repair workflows; coordinate hardware remediation and fleet upgrades; document repeatable processes; validate new and repaired servers; and participate in a follow-the-sun on-call rotation.
Responsibilities
- Investigate and troubleshoot GPU platform problems and hardware faults using system logs, kernel logs, and BMC Redfish APIs
- Coordinate with data center operations, hardware engineering, and capacity planning to repair failed hardware, deliver new hardware, and roll out fleet upgrades
- Automate routine processes and build hardware diagnostics, provisioning, and repair tooling
- Build processes, documentation, and tooling for recurring problems
- Test and validate new and repaired AI hardware and servers
- Participate in an on-call rotation providing follow-the-sun coverage
Requirements
- Strong analytical, troubleshooting, and problem-solving skills
- Linux experience and understanding of Linux internals
- Exposure to server-class hardware and provisioning
- Knowledge of hardware and networking fundamentals
- Excellent communication and collaboration skills
- Bachelor's degree in computer science or a related field, or self-education in computer science fundamentals
Benefits
- Pension contributions
- Private health insurance
- Dental insurance
- Income protection
- Life assurance