freehire launches on Product Hunt on 26 August.

Follow →

AI/ML Infrastructure Engineer

Summary

Build and maintain the infrastructure that powers AI/ML workloads, including GPU clusters, Kubernetes, and observability stacks for large-scale systems.

Delivery Head | Canada Recruitment | Talent Acquisition

Location: Montreal, Quebec, Canada

Seniority level: Mid-Senior level

Employment type: Full-time

Job function: Information Technology

Skills Required

  • Production experience in SRE / Infrastructure / ops for large-scale systems
  • Strong programming/scripting skills (Python, Go, Java, or equivalent)
  • Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
  • Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures
  • Experience with monitoring / observability / logging / alerting tools (Prometheus, Grafana, ELK / EFK, Datadog, etc.)
  • Networking & systems engineering knowledge (TCP/IP, DNS, routing, load balancing, distributed storage)
  • Solid experience in capacity planning, performance tuning, scaling, and incident response
  • Demonstrated ability to lead RCAs, deploy fixes, and drive reliability improvements
  • Experience in regulated environments (financial services, compliance, audit, security) is a strong plus
  • Excellent communication, documentation, and cross-team collaboration skills
  • Proven track record of reducing operational toil via automation

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available