MTS - Research Scientist
Summary
Ironsite is hiring a Research Scientist to develop vision-language models that analyze construction site footage to improve labor productivity and project visibility. The role involves training, benchmarking, and deploying state-of-the-art AI models using proprietary datasets and PyTorch or JAX frameworks.
About Ironsite
The Role
What You'll Do
- Design, train, and iterate on vision-language models fine-tuned for spatial intelligence in construction environments. Your work will directly determine how good our models get.
- Run experiments end to end, from data preparation through training, evaluation, and post-training. Move fast, measure honestly, and share what you learn with the team.
- Contribute to the frontier of what's possible on our data, whether that's establishing new baselines, improving post-training recipes, or exploring long-context architectures.
- Contribute to our Construction Intelligence Benchmark suite across video question answering, temporal reasoning, activity recognition, and site-level analytical reasoning.
- Design evaluation metrics that measure real-world construction task performance, not just standard academic benchmarks.
- Own the evaluation loop for your own work so we always know what's actually improving.
- Take your best models from research code to production deployment, working closely with our infrastructure and hardware teams.
- Apply distillation, quantization, and model routing techniques so state-of-the-art understanding runs affordably across our growing fleet.
- Learn from what happens in the field. Field data and production feedback should shape your next experiment.
- Learn from senior researchers, contribute to the intellectual culture of the team, and start to develop your own point of view on what Ironsite's research should look like at scale.
- Read papers, share ideas, and help set the technical bar for the whole team.
Technical Challenges You'll Solve
- Training models efficiently under real compute budgets while maximizing performance on the problems that matter for our customers.
- Working with novel pre-training and post-training objectives that capture construction-specific knowledge, temporal reasoning, and fine-grained perception.
- Handling the challenges of long-context video data, including temporal reasoning, memory across multi-hour footage, and efficient processing.
- Designing evaluation metrics that predict real-world construction task performance.
- Balancing model capability with deployment constraints for edge, on-prem, and cloud inference.
What We're Looking For
- 2-4 years of hands-on research experience designing and training deep learning models, particularly transformer-based architectures. Industry experience, top-tier PhD program, or equivalent.
- Deep expertise with modern deep learning frameworks (PyTorch, JAX, or similar) and strong proficiency in Python with solid software engineering fundamentals.
- Experience working with large-scale vision or language datasets.
- A track record of shipping meaningful research results, whether at a company, in a lab, or in publications.
- A background in Computer Science, Machine Learning, AI, Robotics, or a related field.
- Experience with vision-language models, video understanding, or multimodal architectures.
- Hands-on experience with post-training techniques for large language or vision-language models (SFT, RL methods, parameter-efficient tuning such as LoRA).
- Familiarity with the challenges of video data, including temporal reasoning and long-context modeling.
- Publications at top-tier AI, ML, or CV conferences.
- Experience optimizing inference at scale (quantization, distillation, sparsity).
- Familiarity with MLOps tools for model training and deployment.
- Interest in vision-language models applied to real-world physical problems, and genuine curiosity about the day-to-day lives of construction workers.
What Success Looks Like
- First 30 days. You know our data, our benchmarks, and our production models cold. You've completed your first end-to-end experiment and shipped at least one meaningful improvement to a model in production.
- First 3 months. You've owned a research initiative from problem definition through deployment, and your work has visibly moved model performance on a problem that matters.
- First 6 months. You're a full contributor to the research team's roadmap and a trusted collaborator across the company. You've grown from executing on research problems to helping shape which ones we take on next.
Location, Compensation, & Perks
- San Francisco Bay Area (on-site)
- Base salary: $175k-$275k per year, commensurate with experience
- Significant early-stage equity
- Full benefits including health, dental, vision, and 401(k) with 6% match
- Access to dedicated GPU compute resources for research and experimentation
- Daily catered breakfast and lunch
- Office in San Francisco, next to Oracle Park and the Caltrain
Why You'll Love Working at Ironsite
- Foundational impact. Solve fundamental AI problems to transform one of the world's largest and least-digitized industries. Your models ship to real jobsites, not just papers.
- Growth opportunity. We're at the stage where every researcher's work materially shapes the company. You'll get scope early, and the growth trajectory is defined by how much impact you're willing to have.
- Dream dataset. Exclusive access to a massive, proprietary, and continuously growing corpus of egocentric jobsite video from hundreds of devices deployed on active construction sites. A moat that enables frontier research.
- World-class team. Learn from and collaborate with a small, elite team of researchers and engineers who have shipped cutting-edge AI products at companies like DeepMind, Etched, Meta, Apple, and NVIDIA.

