Machine Learning / AI Engineer
Responsibilities
- Build and operate end-to-end ML and AI systems covering data preparation, model training, evaluation, inference, deployment, and production iteration.
- Develop production-grade AI infrastructure and backend services, including inference pipelines, orchestration layers, model-serving services, APIs, caching, batching, and streaming.
- Design and implement agentic AI workflows supporting multi-step planning, tool use, memory, failure handling, recovery, and integration with LLMs and external systems.
- Monitor, debug, and optimize AI systems in production using logging, metrics, tracing, alerting, and production signals to improve latency, throughput, cost, reliability, and safety.
- Develop and integrate AI-powered product features across frontend, backend, ML, and AI systems, using production performance and usage data to drive continuous improvements.
Requirements
- Strong production software engineering fundamentals, with experience building and operating high-throughput, low-latency backend services.
- Hands-on experience developing and deploying ML or AI systems, including model training, evaluation, inference, and production monitoring; experience with PyTorch.
- Experience with LLM-based systems and AI inference patterns, including OpenAI, Anthropic, or open-source LLMs, embeddings, tool calling, agent workflows, memory, or multimodal AI.
- Proficiency in Python and Node.js, with experience using SQL and/or NoSQL databases, Docker, and Kubernetes.
- Experience debugging and optimizing distributed AI systems in production, including inference latency, throughput, caching, batching, streaming, observability, reliability, fallback mechanisms, and resource/cost optimization.
GMP Recruitment Services (S) Pte Ltd | EA Licence: 09C3051 | VO UYEN AI LINH | Registration No: R22109232