Member of Technical Staff, Video Data & Model Evaluation
Summary
Join Cantina's Singapore research team to own the data and evaluation strategy for its video generation models: defining quality standards, building and annotating datasets, and running model evaluations that feed back into training. Core work spans generative media datasets, annotation/captioning systems, and tools like Python, SQL, and vision-language models.
About Cantina:
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.
About the Role:
Cantina is looking for a Member of Technical Staff to join our growing Singapore research team. In this role, you will shape the data and evaluation strategies behind our video generation models.
You will own the full data lifecycle—from defining data requirements and annotation standards to building high-quality datasets, evaluating model outputs, and translating findings into actionable recommendations for model improvement. You will work closely with research scientists, machine learning engineers, product teams, and external data partners to strengthen the data–model–evaluation feedback loop.
What You’ll Do
Define quality standards for video generation across aesthetics, realism, composition, motion, temporal consistency, instruction following, subject consistency, editing preservation, and text rendering.
Evaluate model iterations, analyze failure patterns, and translate findings into actionable data, training, and prompting strategies.
Own end-to-end training data projects, including sourcing, construction, annotation, captioning, quality control, and delivery.
Develop labeling systems that support dataset segmentation, coverage analysis, quality filtering, training-data mixtures, and model performance attribution.
Design annotation and captioning standards for video generation, editing, and multi-reference tasks.
Create and evaluate synthetic-data workflows based on their stability, usability, quality, and training value.
Develop automated annotation and filtering workflows using vision-language models, system prompts, and quality-scoring models.
Build structured model evaluation processes covering test-set design, evaluation criteria, failure attribution, and reporting.
Establish feedback loops connecting model evaluation and negative-case analysis with data strategy and model iteration.
Train and calibrate annotation teams and external data partners, conduct quality audits, and improve delivery standards.
Partner with research, engineering, product, and creative teams to execute data initiatives and identify future model capabilities.
What You’ll Bring
Hands-on experience in data strategy, dataset development, annotation, or model evaluation for generative image, video, or multimodal models.
Experience supporting model development across continued training, supervised fine-tuning, post-training, or evaluation.
Strong understanding of generative media tasks, including image and video generation, multimodal generation, and instruction-based editing.
Experience developing labeling taxonomies, annotation guidelines, captioning standards, quality thresholds, and evaluation frameworks.
Strong visual judgment and the ability to translate subjective quality expectations into measurable data and evaluation standards.
Experience managing large-scale data-production or annotation projects and collaborating with research, engineering, creative, and external partner teams.
Familiarity with automated captioning, synthetic-data generation, vision-language models, prompt design, Python, SQL, or other data and workflow automation tools is a plus.
Benefits We Offer:
Competitive salary and generous company equity
Personal time off and paid holidays
Health insurance
Global travel insurance: Covers you when traveling internationally
Monthly spending stipend: $500 (~S$635)
Equipment: All equipment needed for your home office
What they ask for
Required
- Hands-on experience in data strategy, dataset development, annotation, or model evaluation for generative image, video, or multimodal models
- Experience supporting model development across continued training, supervised fine-tuning, post-training, or evaluation
- Strong understanding of generative media tasks, including image and video generation, multimodal generation, and instruction-based editing
- Experience developing labeling taxonomies, annotation guidelines, captioning standards, quality thresholds, and evaluation frameworks
- Strong visual judgment and ability to translate subjective quality expectations into measurable data and evaluation standards
- Experience managing large-scale data-production or annotation projects and collaborating with research, engineering, creative, and external partner teams
Preferred
- Familiarity with automated captioning, synthetic-data generation, vision-language models, prompt design, Python, SQL, or other data and workflow automation tools
