HPC Operations Engineer
You provide frontline operational support for 24/7 Linux HPC compute, storage, and interconnects. You resolve research community problems, respond to alerts, perform coordinated maintenance, develop diagnostic and automation code, support monitoring systems and production tools, document systems, manage vendor relationships, participate in on-call rotations, and work on global infrastructure projects.
Responsibilities
- Provide frontline operational support for Linux HPC compute, storage, and interconnects
- Resolve research community problem reports and questions
- Respond to alerts
- Participate in coordinated maintenance operations
- Work on global infrastructure projects
- Write code for diagnosis, resolution, triage, and automation
- Develop code and testing infrastructure across codebases
- Manage outside vendor relationships
- Implement and support performance and fault monitoring systems
- Develop systems and user documentation
- Develop and monitor production computing tools
- Participate in on-call rotations
- Work from the company office an average of five days a week
Requirements
- At least 2+ years of professional Linux systems administration experience
- HPC experience with parallel filesystems, batch systems, and high-performance network interconnects is a plus
- High proficiency in at least one programming or scripting language such as Go, Python, or C
- Ability to perform root cause analysis
- Strong verbal and written communication skills
- Strong collaboration skills
- Ability to independently manage complex projects and multiple workstreams
- Willingness to perform operational maintenance during evenings and weekends
- Reliable and predictable availability
Benefits
- Private medical insurance
- Vision insurance
- Dental insurance
- Travel medical insurance
- Group pension scheme
- Group life assurance
- Income protection schemes
- Paid parental leave
- Parking benefits
- Commuter benefits