Tech jobs
Job listings
Member of Technical Staff — Inference-Core Engine
Builds and optimizes large-scale AI inference systems for frontier models, focusing on performance, latency, and cost across thousands of GPUs.
Principal Architect, Simulation Platform
Principal Architect leads the design of a GPU-powered simulation platform for AI data centers, enabling real-time power orchestration and validation of energy-aware compute systems.
Member of Technical Staff — Performance
Performance Engineer at RadixArk in Palo Alto optimizes LLM inference and training systems for latency, throughput, and cost efficiency across production workloads using SGLang, Miles, and GPU/TPU infrastructure.
Member of Technical Staff — Developer Technology
Optimize and accelerate LLM inference and training systems like SGLang and Miles by profiling GPU performance, writing custom kernels, and enabling new models on modern hardware.
Staff Machine Learning Engineer
Build and ship production-grade AI systems—LLMs, agents, retrieval, and evals—from prototype to scalable deployment for high-stakes decision-making.
MLOps Engineer (JAX, PyTorch, Pallas/Triton)
Build and evaluate MLOps pipelines for cutting-edge GenAI models using JAX, PyTorch, and GPU kernels (Pallas/Triton) to improve training data quality.
Ingénieur(e) IA Générative - Hébergement De Modèles LLM
Design and build a cloud-native SaaS platform for data management and MLOps that supports 5G core networks, using Python, Go, Kubernetes, and hyperscale cloud services.
Machine Learning Engineer, Foundation Model Services
Build and optimize production-grade inference services for Apple’s large language, vision, and speech models, ensuring low-latency performance across products like Siri and Photos.
Senior Lead AI Application Engineer (Global Security)
Lead the design and delivery of secure, cloud-native AI agent systems and full-stack microservices that automate risk/security controls and integrate with RBC’s pipelines and regulated environments.
Software Engineer, Cloud Infrastructure (Multiple Seniority Levels)
Build and run AWS-based cloud and ML infrastructure for an AI aviation platform, including LLM endpoints, RAG pipelines, and IoT deployments using CDK/Terraform and LangChain.
Senior MLOps Engineer
Senior MLOps Engineer builds and runs scalable LLM inference and training pipelines, deploys large open-weight models, and maintains a cost-aware multi-provider gateway for production ML services.
Machine Learning Performance Engineer (Inference)
Optimize and deploy ultra-low-latency ML inference pipelines for quantitative trading, tuning GPU/FPGA kernels and memory hierarchies to microsecond-level performance.
Machine Learning Engineer
Build and deploy India’s open healthcare foundation model: train a 30B-parameter MoE on 500B+ Indic-language tokens, then ship fast inference and robust evals for doctors and government use.
AI Infrastructure Engineer
Build and optimize distributed AI training systems, profiling bottlenecks in PyTorch stacks and low-level GPU kernels to speed up model convergence.
Machine Learning Compiler and Systems Engineer
Build and optimize a compiler stack for robotics simulation and AI training, using LLVM, JIT, and GPU codegen to maximize performance.
AI System Architect
Design AI accelerators and chips for ML training/inference, model performance, and hardware-software integration in a cutting-edge Oxfordshire startup.
AI System Architect
Design the hardware and software stack that powers Lumai’s optical AI accelerator, ensuring it meets the compute and memory demands of modern transformer models and other AI workloads.
AI / Machine Learning Engineer
Build, evaluate, and deploy ML systems (PyTorch/TensorFlow) and LLMs in production using FastAPI/Triton/Ray, with MLOps and applied research responsibilities.