Ingénieur(e) IA Générative - Hébergement De Modèles LLM
Summary
Build and operate a secure, scalable platform to host large language models, integrating GPU clusters, MLOps pipelines, and observability while optimizing inference performance and tenant isolation.
- Administer secure model hosting and marketplace
- Apply SRE SecOps incident management
- Apply quantization and kernel tuning
- Build Ray based platform architecture
- Deploy GPU clusters for training and inference
- Design GPU cluster infrastructure
- Develop agentic platform runtimes
- Develop model inference services and APIs
- Implement autoscaling and capacity planning
- Implement observability with metrics traces and logs
- Implement tenant isolation and access control
- Integrate MLOps pipelines CI CD
- Integrate NVIDIA CUDA Triton and NCCL
- Lead technical mentorship and design reviews
- Maintain runbooks and monitoring alerts
- Operate multi tenant ML workloads
- Optimize batching and latency
- Perform chaos testing and capacity prediction
- Provide CLI and SDK developer tools