Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Build and deploy large-scale machine learning models and AI agents to optimize Micron’s semiconductor manufacturing workflows using distributed training and GPU optimization techniques.
Support and optimize high-performance computing (HPC) systems for scientific and AI workloads, troubleshooting applications and collaborating with engineering to improve platform usability.
Build and integrate AI/ML features into enterprise Kubernetes environments using Python/Golang, vLLM, PyTorch, and OpenShift, while contributing to open source projects.
Build AI-driven tools to automate gameplay testing and validate NVIDIA GPUs using deep learning, reinforcement learning, and generative models.
Build and optimize NVIDIA’s deep-learning inference stack (TensorRT, TensorRT-LLM) by profiling models, writing GPU kernels, and integrating OSS frameworks to maximize GenAI performance across datacenter and edge GPUs.
Build and optimize CUDA features for AI frameworks like PyTorch and TRT-LLM, improving multi-GPU performance and distributed runtime for training and inference workloads.
Senior engineer building and optimizing NVIDIA’s Dynamo inference platform, integrating open-source frameworks like vLLM and TensorRT-LLM to maximize AI throughput and latency.
Build and optimize GPU-powered inference systems for large language models, improving serving efficiency and performance through kernel tuning and novel algorithms.
Build and maintain scalable CI/CD infrastructure for NVIDIA’s TensorRT Edge-LLM, automating builds, tests, and deployments across embedded and cloud platforms using tools like GitLab, Kubernetes, and Docker.
Design and optimize ML pipelines, LLM serving, and GPU architectures for generative AI solutions using frameworks like PyTorch, Hugging Face, and vLLM.
Build and run a secure Kubernetes-based platform that hosts AI engineering tools, GitOps workflows, and observability stacks for AI-assisted software development in sovereignty-sensitive environments.
Build and maintain secure, scalable AWS and Kubernetes platforms for AI workloads, including LLM inference, while automating CI/CD and ensuring SOC 2/HITRUST compliance.
Build and integrate a GenAI platform using large language models, RAG pipelines, and self-hosted frontier models for a Polish tech firm.
Build autonomous AI agents that plan, collaborate, and act using frameworks like LangChain and DSPy, integrating LLMs with tools and APIs for real-world execution.
Forward Deployed Engineer at VESSL AI: a hands-on role bridging customers and GPU cloud tech, designing PoCs, solving AI workload issues, and driving adoption of VESSL’s GPUaaS platform for training and inference.
Lead post-training and alignment of large language models using DPO, GRPO, and RLAIF; design reward models and optimize inference with vLLM/SGLang for a leading crypto exchange.
Build and improve Avito’s in-house large language models by designing experiments, training pipelines, and optimizing inference for production use.
Architect and own LSports’ GCP-based cloud platform and its agentic AI layer, designing scalable, secure systems for real-time sports data and LLM-powered automation.
Build and run large-scale, cloud-native systems that power Adobe’s AI features, including ML inference infrastructure, Kubernetes clusters, and automated patch management across AWS, Azure, and GCP.
Orange Business is here! About us Join us at Orange Business! We are a network and digital integrator that understands the entire value chain of the digital world, freeing our customers to focus on the strategic…
We couldn't check your fit for this role — add a CV to your profile to see it next time.