多模态理解算法实习生
- Analyze multimodal model capability limits
- Build high quality training data for foundation models
- Build video understanding datasets for human behavior event recognition video caption temporal grounding
- Construct multimodal data pipelines
- Develop embodied understanding models for physical world
- Improve model perception and cognition in visual localization
- Label chain of thought reasoning and build evaluation sets
- Support long horizon behavior and task decomposition
- Synthesize clean annotate format and evaluate text image video and action data