Tech jobs
Job listings
Software Engineer — Distributed LLM Inference Systems
Build and optimize distributed systems that run large language models efficiently across Intel hardware, focusing on inference performance, parallelism, and communication.
Sr. Product Manager, Unstructured Data Workloads for AI & Analytics Solutions
Owns product strategy for unstructured data workloads powering AI/analytics solutions, partnering with engineering to launch features and drive adoption across markets.
Expert Software Engineer - Edge AI
Build and maintain a Rust-based platform that deploys and orchestrates AI models on edge devices, optimizing for constrained hardware and integrating with cloud services.
Senior Computer Vision Algorithm/Software Engineer
Design and implement advanced image-processing and computer-vision algorithms (CNNs/SNNs) for next-gen EO/IR camera systems, integrating them into embedded hardware from concept to production.
GenAI / Agentic Architect
Lead enterprise GenAI strategy, architect scalable AI systems, and advise C-suite on tech stacks, trade-offs, and ROI for large client engagements.
AI Platform Engineer
Design and operate scalable AI inference platforms for production ML workloads, optimizing GPU utilization, LLM serving, and cloud-native infrastructure.
AI Engineer
Build and deploy AI models for edge devices: port vision/speech models to NPU-enabled SoCs, extend AI tooling for hardware-aware development, and create reference designs for smart-home and surveillance use cases.
Staff Machine Learning Engineer - LLM Quantization & Deployment
Build and deploy quantized large language models for in-vehicle AI, focusing on PTQ, QAT, and low-bit inference to ensure numerical consistency and performance on XPENG’s Turing AI chip.
Senior Machine Learning Engineer - LLM Quantization & Deployment
Build and deploy quantized large language models for in-vehicle AI, focusing on PTQ, QAT, and low-bit inference to optimize performance on XPENG’s Turing AI chip.
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Builds and optimizes inference stacks for large-scale Apple foundation models, enabling AI features across services like Siri and Photos with low latency.
System Performance Engineer - AI/ML, Platform Architecture
Optimize AI/ML performance across Apple devices by benchmarking workloads, profiling hardware/software, and driving insights to enhance customer experiences.
Machine Learning Engineer
Build and deploy deep-learning models for vision-based AI systems, focusing on pose, facial, and action recognition using CNNs, transformers, and embedding techniques.
Machine Learning Engineer
About Sunset At its core, Sunset was founded to help founders. We started by supporting startups through shutting down, but we have since expanded into unlocking a new revenue stream for all types of businesses. In…
Senior Software Engineer I, Inference
Senior engineer builds and optimizes CoreWeave’s Kubernetes-native AI inference platform, improving latency, throughput, and reliability to meet strict P99 SLAs while mentoring peers.
Senior Software Engineer, Inference
Senior engineer building and optimizing CoreWeave's Kubernetes-native AI inference platform to meet strict latency and reliability SLAs, using Python/Go, CUDA, and distributed systems.
Senior Software Engineer - Perf and Benchmarking
Senior engineer builds and runs Kubernetes-native benchmarking services to measure latency, throughput, and reliability across CoreWeave’s global AI cloud, publishing results like MLPerf.
Software Engineer, Inference AI/ML
Develops and optimizes AI model-serving systems on GPU infrastructure, focusing on latency, reliability, and cost while working with tools like Triton, vLLM, and Kubernetes.
Staff Software Engineer, Inference
Leads architecture and performance for CoreWeave’s Kubernetes-native AI inference platform, optimizing low-latency, high-throughput systems and GPU resource management.
Senior Software Engineer - GPU Kernel Authoring & Optimization
Write, profile, and optimize CUDA kernels for LLM inference to maximize throughput and minimize latency on NVIDIA GPUs, using DSLs like Triton or Mojo and benchmarking with MLPerf.
Applied AI Engineer, Inference
Applied AI Engineer, Inference at CoreWeave. Improve real-world performance of AI models via benchmarking, profiling, and optimization.