Senior Data Engineer

Summary

Lead the design and scaling of high-throughput distributed data pipelines and storage systems for Luma's multimodal foundation models and Physical AI, using Python, Spark/C++/Rust, and PyTorch/Ray.

Intro to the Role and Team

We are expanding our Singapore Infrastructure & Data Systems Hub to power Luma's next-generation multimodal foundation models and Physical AI pipelines. As a Senior Data Engineer, you will lead the design and scaling of the core data engines feeding our AI models, collaborating directly with our global research and infrastructure teams and setting technical direction for high-throughput data systems, distributed storage, and multimodal dataset processing.

What You'll Own

  • High-Throughput Data Infrastructure: Architect, scale, and operate distributed pipelines capable of ingesting and filtering 100M+ multimodal media items per day.
  • Physical AI & Synthetic Data Engines: Design end-to-end hardware and software data pipelines for egocentric video, Physical AI, and robotics datasets, and build domain-specific synthetic data platforms.
  • Reliability & Storage: Own SRE operations, unified access control, and distributed storage systems across multi-cloud and GPU environments (PyTorch, Ray).
  • Systems Performance: Identify and resolve bottlenecks across network, storage, and compute layers to maximize GPU training cluster utilization; drive engineering best practices and mentor junior engineers.

What You'll Bring

Basic Requirements:

  • Education: Bachelor's or Master's degree in Computer Science, Computer Engineering, Physics, Mathematics, or a related quantitative field.
  • Experience: 5+ years building and operating large-scale data systems in production.
  • Core Technical Stack: Strong proficiency in Python plus systems languages or frameworks for high-concurrency workloads (Scala, Spark, C++, or Rust).
  • Domain Expertise: Deep hands-on experience in at least two of the following: distributed data pipelines and vector storage at scale (100M+ items/day); large-scale web data acquisition and network infrastructure; multimodal, Physical AI, or synthetic/robotics dataset engineering; SRE practices and distributed systems management.
  • Leadership: Track record of driving technical direction autonomously and raising the bar for engineering quality.

Nice-to-Haves:

  • Experience building or maintaining data loaders for large-scale GPU training frameworks (PyTorch, Ray).
  • Familiarity with multimodal vector databases or egocentric video dataset processing.
  • Prior experience scaling globally distributed teams in high-growth AI startups.

About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available