freehire launches on Product Hunt on 26 August.

Follow →

Ingénieur(e) IA Générative - Hébergement De Modèles LLM

Summary

Build and operate a secure, scalable platform to host large language models, integrating GPU clusters, MLOps pipelines, and observability while optimizing inference performance and tenant isolation.

- Administer secure model hosting and marketplace - Apply SRE SecOps incident management - Apply quantization and kernel tuning - Build Ray based platform architecture - Deploy GPU clusters for training and inference - Design GPU cluster infrastructure - Develop agentic platform runtimes - Develop model inference services and APIs - Implement autoscaling and capacity planning - Implement observability with metrics traces and logs - Implement tenant isolation and access control - Integrate MLOps pipelines CI CD - Integrate NVIDIA CUDA Triton and NCCL - Lead technical mentorship and design reviews - Maintain runbooks and monitoring alerts - Operate multi tenant ML workloads - Optimize batching and latency - Perform chaos testing and capacity prediction - Provide CLI and SDK developer tools

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available