freehire launches on Product Hunt on 26 August.

Follow →

HPC Data Center Production Engineer

You will build and own automation and tooling for HPC data center operations in Chicago or New York. You will automate hardware onboarding and lifecycle management, develop capacity planning and outage simulation tools, integrate monitoring and metrics, maintain reliable production systems, document workflows, and use AI tools daily to accelerate development and improve operations.

Responsibilities

  • Automate onboarding of data center hardware
  • Build end-to-end hardware provisioning workflows
  • Adapt onboarding automation for new hardware platforms
  • Develop power and cooling capacity planning tools
  • Build outage simulation tooling
  • Maintain data center lifecycle, inventory, change management, and diagnostic tools
  • Integrate infrastructure telemetry into observability platforms
  • Integrate colocation and data center provider metrics
  • Implement monitoring and alerting strategies
  • Partner with HPC engineering on provisioning integrations
  • Automate manual operational processes
  • Own system reliability and lifecycle
  • Respond to system issues and improve tools
  • Maintain tooling and workflow documentation
  • Participate in evening and weekend maintenance operations
  • Use AI tools for coding, analysis, debugging, documentation, and development
  • Apply AI to anomaly detection, capacity planning, and alerting

Requirements

  • 5+ years of experience in production engineering, infrastructure automation, or site reliability engineering
  • Experience in HPC or large-scale data center environments preferred
  • Track record of building and shipping reliable production automation and tooling
  • Experience automating hardware provisioning and lifecycle management
  • Strong understanding of data center power, cooling, environmental monitoring, and structured cabling
  • Experience with IPMI, BMC, Redfish, SNMP, and vendor APIs
  • High proficiency in Golang and at least one additional language such as Python
  • Strong Linux systems knowledge
  • Experience with Grafana and observability platforms such as Prometheus or InfluxDB
  • Experience with SaltStack, Ansible, Terraform, or similar tools
  • Understanding of L2/L3 protocols, VLANs, BGP, SNMP, and network device configuration
  • Experience with APIs and data integration
  • Experience with ClickHouse and MySQL
  • Experience with GitHub, code review, and CI/CD workflows
  • Professional daily use of AI tools
  • Excellent written and verbal communication skills
  • Reliable and predictable availability
  • Bachelor's degree preferred

Benefits

  • Discretionary bonus eligibility
  • Medical, dental, and vision insurance
  • HSA, FSA, and Dependent Care options
  • Employer Paid Group Term Life and AD&D Insurance
  • Voluntary Life & AD&D insurance
  • Paid vacation plus paid holidays
  • Retirement plan with employer match
  • Paid parental leave
  • Wellness Programs

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available