Machine Learning Research Engineer, Pre-training (LLMs)
Summary
Build and optimize India’s open healthcare foundation model (30B-parameter MoE) by designing data mixes, curricula, and ablation ladders at scale using Megatron/NeMo-class trainers.
ML Research Engineer - Pre-training (LLMs)
Location: Bengaluru, India
Department: Data Science.
Experience: 2 - 4
ML Research Engineer; Pre-training (LLMs)
Bengaluru · Full-time · Experience: 2–4 yrs
About EkaCare and the mission
EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.
Now, under the IndiaAI Mission, EkaCare has been selected to build India's open healthcare foundation model: a ~30B-parameter MoE trained on ~500B+ tokens across 12+ Indic languages. Weights and the India-specific clinical eval suite ship in open domain.
The role
The recipe is the game. You'll work at the heart of the 30B CPT: what goes into the ~500B-token mix, in what order, at what scale, proven cheaply at proxy scale, then spent confidently on the big run.
What you'll do
- Design and run CPT/mid-training ablation ladders at proxy scale, the experimental backbone that decides the real run.
- Own data-mix and curriculum empirics: medical vs general, Indic vs English, replay ratios (~70% design point), annealing schedules.
- Debug training at scale: loss spikes, precision issues, dataloader stalls, checkpoint pathologies.
- Build per-stage eval hooks so every CPT phase has a scoreboard, not a vibe.
- Run tokeniser, long-context and MoE-health experiments (routing balance, expert utilisation).
What we look for
- 2–4 years in ML with pretraining or CPT you personally ran at ≥1B scale (ideally ≥7B, 100B+ tokens) — the recipe was yours to break and fix.
- Fluency with Megatron/NeMo/TorchTitan-class trainers and distributed fundamentals (TP/PP/DP, mixed precision).
- Empirical rigour: you design ablations that answer questions.
- You read papers fast and implement faster.
Bonus
- MoE training exposure; scaling-laws mindset.
- Indic-language or domain-specific (medical/legal/code) pretraining.
- Kernels curiosity — you've opened a profiler and enjoyed it.
Why this is a rare gig
- Open source, with your name on it: weights and technical reports ship publicly.
- India-scale mission: models for a billion people in their own languages.
- Compute that’s rare to fine: dedicated multi-node H200 training under a national grant.
- Small senior team: you work with the people who own the recipe.
- A live deployment path: Government institutes, EkaCare's doctors and patients use what you ship.