Point your AI agent at freehire and let it find you a job.

Get the CLI →

ASUS GLOBAL PTE. LTD.

NewBe an early applicant

ASUS AICS SG - Machine Learning Engineer

Posted Updated
Discussion

About ASUS

AICS is part of ASUS, a multinational company known for the world’s best motherboards, PCs, monitors, graphics cards and routers. Along with an expanding range of superior gaming, content-creation and AIoT solutions, ASUS leads the industry through cutting-edge design and innovations made to create the most ubiquitous, intelligent, heartfelt and joyful smart life for everyone. With a global workforce that includes more than 5,000 R&D professionals, ASUS is driven to become the world’s most admired innovative leading technology enterprise.

About AICS

The mission of ASUS Intelligent Cloud Services (AICS) is to build revolutionary healthcare solutions with natural language processing, computer vision, and big data analytics. We provide Software as a Service (SaaS) applications to accelerate the effective use of medical data and improve the efficiency of hospital operations, unleashing the power of data for precision healthcare and bringing transformative impact to the industry.

Job Overview

We are looking for an experienced Machine Learning Engineer (ML Ops Engineer) to build and operate the infrastructure that enables reliable, scalable, and efficient AI model deployment.

You will work closely with ML Engineers, AI Researchers, Software Engineers, and Product Teams to manage the production lifecycle of AI models, including model evaluation, release, deployment, monitoring, and updates.

The role focuses on GPU infrastructure, model serving, model lifecycle management, and AI platform reliability.

Responsibilities

GPU & Compute Resource Management

  • Design and operate infrastructure for efficient GPU resource allocation and utilization across AI workloads.
  • Manage GPU workloads in Kubernetes and containerized environments.
  • Implement resource scheduling, quotas, priorities, and workload isolation for multiple AI workloads.
  • Monitor GPU utilization, capacity, performance, and resource consumption.
  • Optimize GPU utilization and inference efficiency as workloads scale.
  • Troubleshoot GPU, container, networking, and infrastructure issues in production.

Model Lifecycle & Evaluation

  • Build and maintain processes for model versioning, evaluation, release and rollback.
  • Establish automated workflows to evaluate new model versions against defined quality, performance, and reliability criteria.
  • Design model release gates to ensure new models meet predefined requirements before production deployment.
  • Compare model versions across metrics such as model quality, latency, throughput, and resource consumption.
  • Support controlled model rollout, including canary deployment, A/B testing, and rollback.
  • Maintain model metadata, evaluation results, deployment history, and release status for traceability.

Model Serving &Deployment

  • Build and operate reliable infrastructure for self-hosted AI model inference and serving.
  • Deploy and optimize LLM inference services using vLLM or similar inference engines.
  • Optimize model serving for latency, throughput, GPU utilization, and reliability.
  • Automate model deployment and configuration changes across environments.
  • Implement deployment strategies that minimize service disruption during model updates.
  • Monitor and troubleshoot production inference workloads.

Model Gateway & AI Platform

  • Build and maintain model gateway / model routing infrastructure that provides a unified interface to multiple AI models and inference backends.
  • Support model routing, traffic management, authentication, rate limiting, and observability.
  • Enable applications and AI agents to consume models through a consistent and reliable interface.
  • Integrate different model providers and self-hosted inference services into a unified platform.

Platform Reliability &Observability

  • Build monitoring and observability for AI workloads, including:
  1. GPU utilization and health
  2. Inference latency and throughput
  3. Model quality metrics
  4. Service availability
  5. Model version and deployment status
  6. Resource consumption
  • Establish logging, metrics, tracing, alerting, and operational dashboards.
  • Investigate production incidents and perform root-cause analysis.
  • Continuously improve system reliability, scalability, and operational efficiency.

Requirements

  • 4+ years of experience in software engineering, MLOps, ML infrastructure, DevOps, or a related field.
  • Strong programming skills in Python and/or Go.
  • Hands-on experience operating production AI/ML infrastructure.
  • Strong experience with Docker and Kubernetes.
  • Solid understanding of GPU infrastructure and resource management.
  • Experience with GPU scheduling, resource allocation, monitoring, or capacity planning.
  • Experience operating model serving or inference infrastructure in production.
  • Familiarity with LLM inference and serving, preferably with hands-on experience using vLLM.
  • Experience designing or operating model evaluation and release processes.
  • Experience with CI/CD and infrastructure automation.
  • Strong understanding of Linux, networking, distributed systems, and cloud-native infrastructure.
  • Experience with monitoring and observability.
  • Strong troubleshooting and problem-solving skills.
  • Ability to collaborate effectively with ML researchers, ML engineers, and software engineers.

Nice to Have

  • Experience building or operating a Model Gateway / AI Gateway.
  • Experience with model routing, traffic management, rate limiting, and multi-model serving.
  • Experience with vLLM internals and performance tuning, such as batching, KV cache, GPU memory utilization, and concurrency.
  • Experience with Kubernetes GPU scheduling and NVIDIA GPU infrastructure.
  • Experience operating on-premises GPU clusters.
  • Experience with LLM / Generative AI / Agentic AI infrastructure.
  • Experience implementing canary deployment, A/B testing, or automated model rollback.
  • Experience building internal AI / ML platforms used by multiple teams.
  • Experience with Infrastructure as Code such as Terraform.
  • Experience working in healthcare or other regulated environments is a plus.

Skills

See also

ML / AI jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available