Senior Generative Video ML Engineer, Post-Training & Fine-Tuning
Eros is building sovereign cultural AI: systems designed to make advanced intelligence culturally relevant, rights-aware, governable, and commercially useful. We are bringing together a major film and media catalog, a new AI platform, and a founding technical team, with every training and evaluation asset subject to rights, consent, provenance, territory, and permitted-use controls.
We are hiring a senior, hands-on ML engineer as a founding member of a new generative-video team. The work centers on open-weight video models, reproducible inference, controlled post-training or adaptation, and measurable improvement in character, scene, and temporal consistency.
This is not a prompt-engineering role and not a general AI-app build. We need someone who has personally trained, adapted, evaluated, and debugged modern video-generation models and can show what changed, why it changed, and how the result was measured.
What you will own
- Stand up a reproducible, versioned inference and experimentation environment for one or more open-weight video-generation models.
- Establish frozen baseline results before adaptation begins.
- Inspect data readiness, define train, validation, and sealed holdout splits, and prevent identity or scene leakage.
- Design and run license-permitted post-training experiments using the lightest justified method, such as LoRA, adapter training, supervised tuning, preference optimization, or conditioning changes.
- Build or integrate structured character and scene conditioning while preserving immutable run manifests.
- Measure identity consistency, appearance continuity, scene adherence, action fidelity, temporal stability, latency, and cost.
- Diagnose failures and compare controlled remediation strategies such as repair, retry, reroute, or rejection.
- Package code, configs, run logs, checkpoints or adapters, evaluation results, architecture notes, and complete technical handover materials.
Environment and constraints
- The work draws on a major film and media catalog, but every training and evaluation asset must pass documented rights, consent, provenance, territory, and permitted-use controls.
- All proprietary data and resulting artifacts remain inside a controlled private environment.
- No data may be copied to personal storage or external inference services.
- Training and model use must follow checkpoint-specific commercial-license, data-rights, consent, provenance, security, and territory approvals.
- Work will begin with bounded evidence phases before production-scale training runs.
- Results will be judged against frozen baselines and held-out cases. We do not accept self-selected demos as proof.
Required experience
- Deep hands-on experience with diffusion or flow-based video generation, transformer-based video models, or closely related multimodal generation systems.
- Personally executed post-training or adaptation of an open-weight image or video model using PyTorch.
- Strong GPU systems knowledge, including distributed training or inference, memory optimization, mixed precision, checkpointing, and experiment tracking.
- Built evaluation pipelines for identity, visual consistency, temporal quality, prompt or scene adherence, or production usability.
- Comfortable operating under strict data custody, reproducibility, and evidence requirements.
Strong pluses
- Direct experience with model families such as Wan, MiniMax, HunyuanVideo, CogVideoX, LTX Video, Mochi, or comparable systems.
- Character-consistency techniques, reference conditioning, identity embeddings, temporal adapters, video inpainting, or localized repair.
- Experience optimizing inference on H100 or H200-class infrastructure.
- Experience designing blinded human review alongside automated evaluation.
To apply, please address these questions
- Describe the most relevant video-generation model you personally trained or adapted. Name the base model, method, data scale, compute, your exact contribution, and the measured before-and-after result.
- How did you evaluate identity consistency and temporal quality on held-out examples? Include metrics, human review, and one failure your evaluation caught.
- What is the largest multi-GPU training or inference job you personally operated, and what failed during the run?
- Share one sanitized artifact you can walk through live, such as code, config, an experiment report, an evaluation dashboard, or an architecture document.
- Are you able to work full time and provide reliable overlap with a distributed team?