Software Engineer III, ML Infrastructure, Google Cloud
Summary
Builds and optimizes Google Cloud’s managed ML training platform, enabling customers to run scalable, distributed AI workloads with high performance and reliability.
Vertex Training offers data scientists and machine learning engineers managed services to develop, train, and tune machine learning models. This flexible platform opens access to hardware and frameworks for developing machine learning models, without requiring users to manage the underlying resources. Vertex Training assists data scientists in running and monitoring their machine learning workloads, while still giving them the flexibility to bring on any type of model or code they need.
The team’s mission is to build a Machine Learning (ML) Training platform that empowers customers to develop scalable and performant ML models.The Google Cloud AI Research team addresses AI challenges motivated by Google Cloud’s mission of bringing AI to tech, healthcare, finance, retail and many other industries. We work on a range of unique problems focused on research topics that maximize scientific and real-world impact, aiming to push the state-of-the-art in AI and share findings with the broader research community. We also collaborate with product teams to bring innovations to real-world impact that benefits our customers.
- Build components to enable large-scale distributed training with focus on usability, performance and resiliency.
- Collaborate with partner teams, dependencies (e.g., Google Compute Engine (GCE), Google Kubernetes Engine (GKE), Cloud Tensor Processing Unit (Cloud TPU), Core Infra, CoreML) to develop pre-training/post-training software to enable customer run reliable and performant workloads.
- Work with Product Team, Account teams and customers to identify key pain points, define scope of the problems, translate them into projects and lead them to execution.
- Contribute to product excellence, and improve overall product quality and user experience.
Minimum qualifications:
- Bachelor’s degree or equivalent practical experience.
- 2 years of experience with software development in one or more programming languages (e.g., Python, C++ or Java).
- Experience in machine learning infrastructure.
- Experience in distributed machine learning.
- Experience with distributed computing and Cloud APIs.
Preferred qualifications:
- Experience with large-scale distributed machine learning, ML performance optimizations, etc.
- Experience with ML frameworks (e.g., PyTorch).
- Experience working on Cloud and ML products.
- Experience in back-end system development and familiarity with common Google Infrastructure.
- Experience with OSS AI/ML Job scheduling software.