Ingénieur(e) IA Générative - Hébergement De Modèles LLM
Summary
Build and maintain LLM hosting infrastructure, deploy GPU clusters, and create developer tooling for AI model serving and orchestration.
- Apply SRE and SecOps practices for reliability and security
- Build and maintain model serving services and APIs
- Create developer experience tooling with CLI SDK and templates
- Design deploy and operate GPU cluster infrastructure
- Develop agent platform runtimes and multi step workflow orchestration
- Implement secure auditable model marketplace hosting with tenant isolation
- Integrate platform with MLOps pipelines and model governance
- Lead platform architecture and implementation with Ray and NVIDIA integrations
- Optimize batching and latency for inference workloads
- Provide technical mentorship and lead design reviews