Tech jobs
Job listings
Senior Platform & Infrastructure Engineer MLBots
Build and maintain the distributed compute and orchestration platforms powering large-scale ML training and simulation for AI game agents.
Sr. Software Development Engineer
Develops and debugs AI/ML software (Python, PyTorch, JAX, TensorFlow) for AMD’s semiconductor operations, focusing on distributed training, transformer architectures, and LLM/Diffusion model fine-tuning to advance AI hardware and cloud platforms.
Machine Learning Engineer, II - 3D Perception
Build and improve 3D perception models (LiDAR, camera, BEV) for autonomous trucks using PyTorch and Python, integrating them into Torc’s autonomy stack.
Machine Learning Engineer, II - 3D Perception
Build and improve 3D perception models (LiDAR, cameras) for autonomous trucks using PyTorch and Python, integrating ML into Torc’s autonomy stack.
Helix AI Engineer, Training Performance
Optimize distributed AI training for 100B+ parameter models across 100k+ GPUs, writing CUDA/Triton kernels and co-designing hardware-efficient training recipes.
Founding Engineer - ML Research
Build and scale the AI research backbone for a Series A startup, designing and training LLMs/diffusion models, optimizing training pipelines, and shipping research-driven features.
Machine Learning Engineer – Feed Recommendation Singapore
Build and optimize large-scale feed recommendation systems using ML to improve user engagement and retention for a social media platform.
Member of Technical Staff — Pretraining Infra (Experienced)
Build and scale distributed training infrastructure for large-scale AI avatar models, optimizing GPU clusters, parallelism, and multimodal data pipelines for real-time, full-duplex training.
Senior ML Engineer _TT
Own the design, training, and deployment of a novel foundation model, including custom CUDA kernels, distributed training pipelines, and cloud-based scaling.
LLM Pre-training & Distributed Engineer (AI Infrastructure)
Engineer large-scale LLM pre-training pipelines and distributed GPU clusters using PyTorch, DeepSpeed, or Megatron-LM, optimizing networking, memory, and fault tolerance for month-long runs.
Technical Operational Program Manager, AI Infrastructure
Lead cross-functional programs to scale AI server infrastructure from prototype to volume production, coordinating hardware design, ODM partners, and supply chains to deliver production-ready systems for large-scale AI workloads.
Software Technical Program Manager
Technical Program Manager at an AI startup leading cross-functional programs to build and deploy custom AI servers, coordinating hardware, software, and supplier teams.
Software Technical Program Manager
Lead cross-functional programs to build and deploy custom AI servers, coordinating hardware, software, and ODM partners while unblocking complex technical bottlenecks.
Senior Deep Learning Algorithm Engineer
Senior engineer building and optimizing open-source AI frameworks (Megatron Core, NeMo) for large language and multimodal model training, fine-tuning, and deployment on NVIDIA GPUs.
Machine Learning Engineer
Build and deploy AI models for defense/intelligence missions using Python and PyTorch, from research to production deployment in Linux environments.
Senior Deep Learning Compiler Engineer - XLA
Build and optimize compiler algorithms for deep learning workloads, focusing on JAX and OpenXLA to accelerate AI training and inference on NVIDIA GPUs.
AI Infrastructure Engineer
Designs and maintains GPU clusters, Kubernetes, and high-speed networking/storage for AI/ML workloads, using Terraform, Ansible, and observability tools.