Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Own and optimize the CI/CD infrastructure for SGLang, an open-source LLM inference engine, ensuring fast, reliable, and secure test pipelines across multiple GPU hardware pools.
Senior HPC systems administrator managing a 600-node cluster, storage, InfiniBand/Ethernet networks, and Slurm workloads to support AI/ML and computational research at a Saudi research university.
Build and maintain ML models to analyze 4G/5G network performance, troubleshoot issues, and create dashboards for stakeholders using Python, Spark, and cloud tools.
Build and deploy AI/ML models using Python, FastAPI, and Generative AI tools like LangChain and vector databases to create scalable solutions for clients in media, e-commerce, and other tech sectors.
Develop GPU-accelerated open-source software for scalable tomographic reconstruction pipelines using C/C++ and Python, optimizing data flows across CPUs, GPUs, and multi-node systems.
Build and maintain scalable cloud infrastructure for AI/ML video systems using Kubernetes, GPU workloads, and MLOps pipelines.
Build and scale ML infrastructure for brain-computer interface R&D, including distributed training pipelines and large-scale data platforms to support neuroscientific modeling and neural decoding.
Designs and prototypes hardware-software systems for defense, automotive, and aerospace, translating algorithms into deployable solutions on GPUs, FPGAs, and ASICs.
Develop and optimize computer vision algorithms (object detection, OCR, 3D processing) from prototypes into production-ready C++/Python code for embedded and mobile hardware.
Develop and optimize computer vision algorithms into production-ready code for embedded systems, focusing on performance and hardware constraints using C/C++/Python.
Support and architect AI platforms for DDN’s Hyperpod, diagnosing issues across NVIDIA AI Enterprise, vector databases, GPUs, Kubernetes, and high-performance storage/networking.
Technical Program Manager at RadixArk coordinates large-scale AI infrastructure programs, including inference engines, training frameworks, and GPU integration, to deliver systems serving billions of tokens daily.
Builds and optimizes large-scale AI inference systems for frontier models, focusing on performance, latency, and cost across thousands of GPUs.
Build and optimize GPU/CPU-accelerated image and data processing libraries (e.g., nvComp, NPP) for AI, computer vision, and scientific workflows using C/C++ and CUDA.
Drive adoption of NVIDIA’s AI and data platform tools with strategic partners by building integrations, crafting joint solutions, and aligning technical and business goals.
Design, develop, and test embedded software for undersea payload systems using C++, Python, and Linux, ensuring compliance with requirements and supporting agile teams.
Performance Engineer at RadixArk in Palo Alto optimizes LLM inference and training systems for latency, throughput, and cost efficiency across production workloads using SGLang, Miles, and GPU/TPU infrastructure.
Builds and engages the technical community around SGLang and Miles by creating content, speaking at events, and collaborating with developers to optimize AI infrastructure.
Optimize and accelerate LLM inference and training systems like SGLang and Miles by profiling GPU performance, writing custom kernels, and enabling new models on modern hardware.
Principal-level role shaping MinIO’s AI ecosystem strategy, designing AI Factory architectures with partners like NVIDIA and Databricks, and driving technical thought leadership across GSIs.
We couldn't check your fit for this role — add a CV to your profile to see it next time.