Point your AI agent at freehire and let it find you a job.

Get the CLI →

THE SUPREME HR ADVISORY PTE. LTD.

NewBe an early applicant

6723 - AI Systems Infrastructure Engineer [Up to $7K | Kaki Bukit | Hand On Exp in Infrastructure engineering]

Discussion

AI Infrastructure Engineer

  • 5 days, Mon - Fri 8.30am to 5.30pm

  • Salary: $5,000 to $7,000

  • Location: Kaki Bukit

  • Operating Systems: Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).

  • Accelerated Compute: Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.

  • Orchestration & Workload Scheduling: Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).

  • High-Speed Networking: Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations.

  • Storage Systems: Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.

  • Automation: Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).

  • Bachelor’s Degree in Computer Science, Information Technology, Computer Engineering, or equivalent practical experience.

  • 3–6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.

  • Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified Associate/Professional, AWS/Azure/GCP Solutions Architect).

Job scopes:

Compute & Cluster Management

  • Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures).

  • Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.

  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.

High-Performance Networking & Storage

  • Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).

  • Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines.

Automation & Infrastructure as Code (IaC)

  • Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.

  • Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.

Operations, Observability & Performance

  • Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).

  • Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.

  • Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.


📲Interested candidates, please WhatsApp your resume to: +65 9642 0989 (Han) 📧 Or email your resume to: supreme.cc.han@gmail.com

👤Chaw Chiaw Han | R22106723🏢 The Supreme HR Advisory Pte Ltd | EA 14C7279

Skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available