HPC Operations Engineer
You provide frontline operational support for 24/7 Linux HPC compute, storage, and interconnects. You resolve research community problems, respond to alerts, perform maintenance, develop diagnostic and automation code, support monitoring systems and production tools, document systems, manage vendor relationships, participate in on-call rotations, and work on global infrastructure projects.
Responsibilities
- Provide frontline operational support for Linux HPC compute, storage, and interconnects
- Resolve research community problem reports and questions
- Respond to alerts
- Participate in coordinated maintenance operations
- Work on global infrastructure projects
- Write code for diagnosis, resolution, triage, and automation
- Develop code and testing infrastructure across codebases
- Manage outside vendor relationships
- Implement and support performance and fault monitoring systems
- Develop systems and user documentation
- Develop and monitor production computing tools
- Participate in on-call rotations
- Work from the company office an average of five days a week
- Work a rotating Friday evening or Saturday morning maintenance window
Requirements
- At least 2+ years of professional Linux systems experience
- HPC experience with parallel filesystems, batch systems, and high-performance network interconnects is a plus
- High proficiency in at least one programming or scripting language such as Go, Python, or C
- Ability to perform root cause analysis
- Strong verbal and written communication skills
- Strong collaboration skills
- Ability to independently manage complex projects and multiple workstreams
- Willingness to perform operational maintenance during evenings and weekends
- Reliable and predictable availability