Hybrid ML Engineer: Python, PyTorch, Distributed Training
Summary
Build and deploy production-grade ML models in San Jose, scaling distributed training across multi-GPU clusters and optimizing performance for low-latency serving.
Enigma is seeking an experienced Machine Learning Engineer in San Jose to drive the end-to-end lifecycle of production ML models. You will scale training across multi-GPU clusters, optimize performance and cost, and build robust serving systems with attention to latency and reliability.
You will work across Research, Platform/Infra, Data, and Product teams, applying distributed training techniques (DDP/FSDP/ZeRO) and modern tooling (Triton, vLLM, ONNX, TensorRT) to deliver production-ready