Research Scientist, Gemini for Android XR Devices

Summary

Research Scientist at Google in Zürich develops and improves AI models (LLMs, video understanding) for XR devices, focusing on Gemini integration and egocentric applications.

As an organization, Google maintains a portfolio of research projects driven by fundamental research, new product innovation, product contribution and infrastructure goals, while providing individuals and teams the freedom to emphasize specific types of work. As a Research Scientist, you'll setup large-scale tests and deploy promising ideas quickly and broadly, managing deadlines and deliverables while applying the latest theories to develop new and improved products, processes, or technologies. From creating experiments and prototyping implementations to designing new architectures, our research scientists work on real-world problems that span the breadth of computer science, such as machine (and deep) learning, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more.

As a Research Scientist, you'll also actively contribute to the wider research community by sharing and publishing your findings, with ideas inspired by internal projects as well as from collaborations with research programs at partner universities and technical institutes all over the world.

In this role, you will partner with the Gemini modeling team to make Gemini a powerful AI assistant for existing and upcoming XR devices. Your work will cover multiple aspects of model development and improvement. You will collaborate with data collection teams to source, annotate, and format valuable data, and work to improve overall model performance, robustness, and reliability through advanced training, deployment, and customization techniques. You will possess expertise in developing and improving LLMs for video understanding, ideally with an egocentric focus. This expertise can stem from academic research or direct industry experience.

For decades, the computing revolution has reshaped our world driven by
breakthroughs in compute, connectivity, mobile, and now, AI. Google's XR
team is at the forefront of the next major leap – the convergence of AI and XR. This is more than just new devices – it's about reimagining how we interact with the world around us. We're building a future where
lightweight XR devices like smart glasses and headsets pair with helpful AI to augment human intelligence, offering personalized, conversational, and contextually aware experiences.
  • Define the data structure, framework, design, and evaluation metrics for research solution development and implementation under minimal guidance. Identify timelines and obtain resources needed.
  • Work on the integration of Gemini into XR products. Design and implement customization layers to enable the next generation of immersive XR experiences powered by Gemini.
  • Work together with research teams to highlight model limitations and improve performance in key areas critical for XR applications like spatial and temporal understanding.
  • Develop expertise on the full data collection, filtering and pre-processing pipeline to contribute with freshly sourced data to improve the quality of models in egocentric video understanding applications.
  • Establish and maintain relationships with main stakeholders and keep recurring updates on the advancements of the project, making sure that their expectations are aligned to the research and engineering work.

Minimum qualifications:

  • PhD in Computer Science, a related field, or equivalent practical experience.
  • Experience with Computer Vision, Generative AI, or Large Language Models.
  • One of more scientific publication submission(s) for conferences, journals, or public repositories (such as CVPR, ICCV, NeurIPS, ICML, ICLR, etc.).
  • Experience in machine learning, specifically for Generative AI and video applications.

Preferred qualifications:

  • 2 years of coding experience.
  • 1 year of experience owning and initiating research agendas.
  • Experience working on multimodal understanding problems, including image or video understanding tasks.
  • Experience working with egocentric data and egocentric problems.
  • Experience working in an industry or academic research lab, focusing on multiple aspects of the "research to product" pipeline.