Senior AI Engineer
About the role
We are building AI systems that generate production-quality motion from natural language. The system combines frontier language models, an AI generation harness, and a text-native Motion DSL designed for structured, editable animation.
This role spans two connected areas: improving the production generation system used today, and developing specialized models that can generate the Motion DSL directly with higher quality, lower latency, and better cost efficiency.
You will work at the intersection of LLM systems, post-training, code generation, compilers, evaluation, data engineering, and motion design. This is a hands-on engineering role with end-to-end ownership and measurable product impact.
Key Responsibilities
Build and improve production generative systems
- Design and ship improvements across prompt interpretation, model orchestration, routing, retrieval, tool use, structured generation, validation, repair, and visual verification.
- Diagnose recurring failure modes and turn them into durable improvements in prompts, data, system logic, constraints, or evaluation.
- Build compiler-backed feedback loops and deterministic quality gates that prevent invalid or low-quality outputs from reaching users.
- Develop experiments and fixed evaluation batteries that show whether a change genuinely improves output quality.
Train specialized generative models
- Design supervised fine-tuning datasets, training recipes, and post-training experiments for direct Motion DSL generation.
- Explore distillation, preference optimization, synthetic-data generation, reinforcement-learning approaches, and constrained generation where they are the right tools.
- Select checkpoints using robust evaluations across correctness, visual quality, reliability, latency, and cost - not training loss alone.
- Determine whether a model failure is best addressed through data, training, inference, evaluation, or the underlying language/runtime.
Build the data and evaluation foundation
- Turn production generations into high-quality training and evaluation datasets using filtering, provenance, versioning, deduplication, and contamination controls.
- Design train, validation, and evaluation splits that minimize leakage and preserve meaningful generalization tests.
- Create failure taxonomies, hard negatives, regression suites, and representative prompt batteries.
- Combine deterministic checks, model-based judges, render evidence, and human review into a reliable evaluation system.
What we're looking for
- Strong ML and software engineering
You have built and operated production AI or machine-learning systems, not only prototypes. You are comfortable moving across model behavior, data pipelines, APIs, infrastructure, evaluation, and product code.
- Hands-on LLM training experience
You have practical experience with supervised fine-tuning and modern post-training workflows. You understand how dataset construction affects model behavior and can explain how you prevent leakage, contamination, and misleading evaluation results.
- Strong evaluation instincts
You know that generative systems improve only when they can be measured. You can design experiments, regression suites, automated graders, and evaluation datasets that distinguish real gains from noise.
- Experience with structured or code generation
Experience with code-generation models, DSLs, grammars, parsers, compilers, structured outputs, constrained decoding, or program synthesis is especially relevant. The generated output is executable structured code, so syntactic and semantic correctness both matter.
- Production engineering judgment
You treat observability, reliability, latency, inference cost, caching, failure recovery, and maintainability as part of the ML system itself.
- Product and visual judgment
You can distinguish technically valid output from work that feels polished. Experience with animation, motion design, graphics, creative tooling, or multimodal systems is valuable, but not required.
Nice to have
- Experience fine-tuning or evaluating code-generation models.
- Experience with multimodal or vision-language models.
- Experience building model-based, human-in-the-loop, or rubric-driven evaluation systems.
- Experience with compilers, interpreters, language tooling, or program analysis.
- Experience with Rust, PyTorch, or distributed training infrastructure.
- Experience with preference optimization, reinforcement learning, or synthetic-data pipelines.
- Experience with animation, graphics, rendering, or creative software.
LottieFiles Perks
- Fully Remote Working Environment
- Flexible Work Hours
- A welcome gift and LottieFiles swag pack
- Bonus to set up your workstation at home
- Unlimited Leave Days
- Medical Insurance
- Generous learning budget
- Gym membership
- Co-working space membership
Please note: To proceed with your application, you must confirm your acknowledgment of this Privacy Policy by ticking the checkbox on the next page.
Read our Privacy Policy here: LottieFiles: Privacy Policy

