AI Platform Engineer / MLOps Engineer
Summary
Build and scale GPU-backed Kubernetes infrastructure to host and operate production LLM services, RAG pipelines, and agentic workflows for a quantitative investment firm.
We are seeking a Senior AI Platform / MLOps Engineer to join a premier global quantitative investment firm.
To be completely transparent about what this role is—and what it isn't: This is a Systems, Infrastructure, and MLOps role at its core. We are not looking for a pure application software engineer or a research-focused ML scientist. We need an infrastructure expert who has grown into the AI/LLM space, someone who knows how to host, operate, and scale GPU-backed systems in production Unix environments.
You will own and scale the internal AI infrastructure that powers high-performance LLM services, RAG pipelines, and agentic workflows across the entire enterprise.
Key Responsibilities
Infrastructure & Platform Operations: Build, scale, and maintain GPU-context Kubernetes environments and Unix infrastructure hosting production LLM workloads.
AI Service Architecture: Integrate model serving endpoints into core application services, maintaining production-grade APIs.
RAG & Vector Infrastructure: Manage the performance, scalability, data freshness, and retrieval quality of vector databases and ingestion pipelines.
Reliability & Observability: Implement end-to-end monitoring, tracing, prompt versioning, and fallback mechanisms for critical AI endpoints.
-
Service Lifecycle: Own service rollouts, deprecation, latency optimization, and operational incident response for high-throughput AI services.
What We Are Looking For (Must-Haves)
Heavy Systems & Infra Focus: Strong background as a Platform, MLOps, or Infrastructure Engineer with deep Unix and Kubernetes expertise (specifically handling containerized environments and GPU workloads).
Production LLM Experience: Proven hands-on experience hosting, operating, and scaling LLM systems and production RAG pipelines in live environments (simply calling an API or building a basic chatbot wrapper won't fit this role).
Cloud Infrastructure: Strong hands-on knowledge of AWS (networking fundamentals, IAM, cloud security, and orchestration).
Targeted Coding Ability: Solid Python skills sufficient to write custom infrastructure tasks, build production APIs, and handle service integration.
-
Vector DB & Retrieval: Experience tuning vector databases and understanding context/token constraints, embeddings, and output reliability.
Nice to Have
Hands-on experience with agentic AI systems and workflow orchestration frameworks.
Familiarity with inference optimization, model serving platforms, and LLM evaluation/quality measurement frameworks.
-
AWS or Kubernetes (CKA/CKAD) certifications.
Why Apply?
Work with massive computational resources and cutting-edge tech in a data-driven, highly analytical environment.
Direct ownership over next-generation AI platform architecture without bureaucratic drag.
-
Top-tier competitive compensation package and flexible working model.
Interested in leading the infrastructure side of AI? Apply now or send over your CV to start the conversation!