Forward Deployed Engineer, GenAI, Google Cloud, Data
Summary
Build and deploy production-grade GenAI data pipelines and evaluation harnesses for enterprise customers using Google Cloud tools like BigQuery, Dataflow, and Vertex AI.
Google will be prioritizing applicants who have a current right to work in Singapore, and do not require Google's sponsorship of a visa. Minimum qualifications:
- Bachelor's degree in Engineering, Computer Science, a related field, or equivalent practical experience.
- 8 years of experience with software development and data engineering with SQL, Python, Java, Scala, or Go.
- Experience with modern Extract, Transform, Load/Extract, Load, Transform (ETL/ELT) frameworks (e.g., dbt, Dataform) and designing enterprise data modeling layers or data marts.
- Master's degree or PhD in Computer Science, Data Science, Artificial Intelligence, or a related technical field.
- Experience integrating semantic metadata formats enterprise taxonomies, or ontologies into large-scale data warehouses and lakes.
- Deep experience designing batch, offline, and online evaluation harnesses and intelligence mining jobs to benchmark LLM capabilities (e.g., Text-to-SQL accuracy, semantic parsing, etc).
- Practical knowledge of configuring and deploying secure code execution harnesses and interpreter sandboxes (e.g., Python/SQL execution environments) for automated data analysis.
- Advanced expertise in synthetic data generation at scale while maintaining multi-table referential integrity using tools like Faker, Snowfakery, or custom constraint engines.
- Serve as a developer for complex AI applications, transitioning from rapid prototypes to production-grade agentic workflows that drive measurable Return on Investment (ROI).
- Co-build with customer engineering teams to instill Google-grade development best practices, ensuring long-term project success and high end-user adoption.
- Design and build high-throughput batch and streaming data pipelines and utilities to curate multi-terabyte evaluation datasets and execute offline/online evaluation generation jobs for model intelligence mining.
- Construct scalable ETL/ELT pipelines using Dataform, dbt, BigQuery, or Dataproc to design enterprise data marts and semantic modeling layers specifically engineered to maximize data quality, schema clarity, and accuracy for Text-to-SQL and natural language analytical interfaces.
- Create mechanisms for large-scale synthetic data generation that maintain strict referential integrity across complex relational schemas, leveraging advanced tools and custom generative utilities for privacy-safe model benchmarking and fine-tuning.