Research Engineer - LLM Training Infrastructure - Seed Infra
Summary
Research Engineer focused on optimizing and scaling infrastructure for large language model (LLM) training, addressing performance bottlenecks and designing distributed training strategies for exascale systems.
- Analyze performance bottlenecks in exascale training systems
- Conduct research and development on large scale LLM training infrastructure
- Design and optimize distributed training strategies for LLMs
- Investigate system reliability and resilience
- Research and optimize network scheduling and GPU memory management
- Translate research ideas into scalable AI infrastructure solutions