Senior AI Platform Engineer – LLM Inference Backend
Summary
Build and optimize backend services for LLM inference, focusing on routing, batching, scheduling, and GPU utilization to improve performance and reliability of AI systems.
JPMorgan Chase & Co. is hiring a Software Engineer III in London to build backend services for LLM inference and scalable production systems.
You will work on routing, batching, scheduling, streaming responses, and quota management while improving APIs, observability, and reliability across the platform. You will explore model architectures, tokenization costs, and GPU utilization, collaborating with product teams and using enterprise AI tooling to boost performance and security of critical