Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Lead ROCm software validation for AMD Instinct GPUs, defining test architecture, CI/CD pipelines, and release gates for AI/ML and HPC workloads across multi-node systems.
Develop and optimize deep learning frameworks (PyTorch, TensorFlow, SGLang) for AMD GPUs, improving kernel performance and scaling AI workloads across multi-GPU and multi-node systems.
Architect GPU performance for next-gen graphics and AI workloads, analyzing bottlenecks and proposing hardware improvements using C/C++, Python, and GPU APIs like Vulkan/CUDA.
Build and optimize the Triton compiler and runtime for AMD GPUs to enable scalable, distributed AI workloads across multi-GPU systems.
Builds high-performance computing and quantitative software for finance and other performance-critical systems, solving complex compute-intensive problems.
Build and optimize LM Studio’s inference runtime for on-device and cloud AI, integrating new engines and models while improving performance across CPU/GPU targets.
Build and optimize inference systems for AI models powering trading decisions, focusing on GPU kernels, FPGAs/ASICs, and data streaming to maximize performance in a high-impact, research-driven environment.
Build and optimize large-scale AI model training systems (kernels, data loading, parallelism) for HRT’s trading-focused foundation models, working closely with researchers to improve performance and impact.
The Role: Build the software that lives next to—or directly inside—the robot. Munari’s edge runtime must capture and synchronize high-rate multimodal data, run local models and detection logic, manage policy execution,…
The role: Own the loop from raw robot experience to a model running on real hardware. Munari observes what robots see, sense, decide, and do. We want to use that data to detect abnormal behavior, understand failures,…
Evangelize Modular’s MAX AI inference and serving platform through technical content, benchmarks, and community engagement to help developers deploy models efficiently.
Champion Mojo, a new systems language for AI workloads, by creating tutorials, blog posts, and conference talks to help developers adopt it.
The Role Lucid builds verifiable compute infrastructure: systems that produce cryptographic evidence about what AI hardware is doing, where it's doing it, and whether the computation matched what was claimed. Where…
Write and optimize low-level compute kernels (matmul, attention, quantization) in C++23 for a custom RISC-V chip, using SIMD intrinsics and memory-hierarchy tuning to accelerate LLM inference/training.
Build high-fidelity 3D surgical simulations and synthetic data pipelines to train perception and policy models for autonomous robotic surgery.
Develop embedded Linux software for real-time, high-definition stereo video processing and illumination control in the da Vinci surgical system.
Builds and maintains embedded Linux software for real-time, high-definition stereo video processing and illumination control in Intuitive’s da Vinci surgical system.
Build and maintain the Kubernetes-based infrastructure and MLOps tooling that trains, deploys, and monitors AI models for robotic-assisted surgery at scale.
Build and deploy large-scale machine learning models and AI agents to optimize Micron’s semiconductor manufacturing workflows using distributed training and GPU optimization techniques.
Staff HPC Applications Engineer Location: Austin, TX (US) Description NextSilicon is reimagining high-performance computing. Our accelerated compute solutions leverage intelligent adaptive algorithms to vastly…
We couldn't check your fit for this role — add a CV to your profile to see it next time.