Tech jobs
Job listings
Data Scientist / ML Engineer (Computer Vision), Junior+ / Middle
Build and ship computer-vision models: train, evaluate, wrap in services, and demo via simple web UIs. Core stack: Python, PyTorch, OpenCV, FastAPI/Flask, Docker.
Member of Technical Staff — Model Optimization and Inference (New Grad)
Optimize and deploy real-time AI avatars by accelerating LLM, audio, and diffusion model inference to sub-500ms latency using quantization, KV cache tuning, and custom kernels.
Member of Technical Staff — Model Optimization and Inference (Experienced)
Optimize and deploy real-time AI avatar models for sub-500ms latency, focusing on KV cache, quantization, and kernel-level acceleration across LLMs, audio, and diffusion components.
Software Engineer — Distributed LLM Inference Systems
Build and optimize distributed systems that run large language models efficiently across Intel hardware, focusing on inference performance, parallelism, and communication.
Sr. Product Manager, Unstructured Data Workloads for AI & Analytics Solutions
Owns product strategy for unstructured data workloads powering AI/analytics solutions, partnering with engineering to launch features and drive adoption across markets.
Expert Software Engineer - Edge AI
Build and maintain a Rust-based platform that deploys and orchestrates AI models on edge devices, optimizing for constrained hardware and integrating with cloud services.
Senior Computer Vision Algorithm/Software Engineer
Design and implement advanced image-processing and computer-vision algorithms (CNNs/SNNs) for next-gen EO/IR camera systems, integrating them into embedded hardware from concept to production.
GenAI / Agentic Architect
Lead enterprise GenAI strategy, architect scalable AI systems, and advise C-suite on tech stacks, trade-offs, and ROI for large client engagements.
AI Platform Engineer
Design and operate scalable AI inference platforms for production ML workloads, optimizing GPU utilization, LLM serving, and cloud-native infrastructure.
AI Engineer
Build and deploy AI models for edge devices: port vision/speech models to NPU-enabled SoCs, extend AI tooling for hardware-aware development, and create reference designs for smart-home and surveillance use cases.
Staff Machine Learning Engineer - LLM Quantization & Deployment
Build and deploy quantized large language models for in-vehicle AI, focusing on PTQ, QAT, and low-bit inference to ensure numerical consistency and performance on XPENG’s Turing AI chip.
Senior Machine Learning Engineer - LLM Quantization & Deployment
Build and deploy quantized large language models for in-vehicle AI, focusing on PTQ, QAT, and low-bit inference to optimize performance on XPENG’s Turing AI chip.
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Builds and optimizes inference stacks for large-scale Apple foundation models, enabling AI features across services like Siri and Photos with low latency.
System Performance Engineer - AI/ML, Platform Architecture
Optimize AI/ML performance across Apple devices by benchmarking workloads, profiling hardware/software, and driving insights to enhance customer experiences.
Machine Learning Engineer
Build and deploy deep-learning models for vision-based AI systems, focusing on pose, facial, and action recognition using CNNs, transformers, and embedding techniques.
Machine Learning Engineer
About Sunset At its core, Sunset was founded to help founders. We started by supporting startups through shutting down, but we have since expanded into unlocking a new revenue stream for all types of businesses. In…
Senior Software Engineer I, Inference
Senior engineer builds and optimizes CoreWeave’s Kubernetes-native AI inference platform, improving latency, throughput, and reliability to meet strict P99 SLAs while mentoring peers.
Senior Software Engineer, Inference
Senior engineer building and optimizing CoreWeave's Kubernetes-native AI inference platform to meet strict latency and reliability SLAs, using Python/Go, CUDA, and distributed systems.
Senior Software Engineer - Perf and Benchmarking
Senior engineer builds and runs Kubernetes-native benchmarking services to measure latency, throughput, and reliability across CoreWeave’s global AI cloud, publishing results like MLPerf.
Software Engineer, Inference AI/ML
Develops and optimizes AI model-serving systems on GPU infrastructure, focusing on latency, reliability, and cost while working with tools like Triton, vLLM, and Kubernetes.