Senior Software Engineer, Search Platforms, LLM Evaluation Infrastructure
At Google, our mission is to organize the world's information and make it universally accessible and useful. The Search Platforms Large Language Model (LLM) Evaluation Infrastructure team is at the forefront of this mission, building the next-generation core infrastructure and benchmarking platforms that power robust, automated, and scalable LLM evaluations across Google Search.
As a Senior Software Engineer, you will lead the architecture, design, and implementation of LLM evaluation platforms and continuous benchmarking pipelines. You will transform complex, bespoke evaluation workflows into a unified, high-throughput platform capable of assessing quality, factual accuracy, safety, and latency across advanced model architectures. You will be responsible for creating and evolving scalable, reproducible back-end evaluation infrastructure for systems operating at a massive global scale.In Google Search, we're reimagining what it means to search for information – any way and anywhere. To do that, we need to solve complex engineering challenges and expand our infrastructure, while maintaining a universally accessible and useful experience that people around the world rely on. In joining the Search team, you'll have an opportunity to make an impact on billions of people globally.
- Architect, build, and maintain high-throughput, low-latency platform services and orchestration pipelines that execute automated evaluation, benchmark suites, and metric computations for Generative Artificial Intelligence (GenAI) Search experiences.
- Design robust evaluation systems to detect metric drift, optimize evaluation throughput, and scale model-assisted rating (e.g., LLM-as-a-judge/Autorater) workflows while maintaining compute efficiency.
- Drive the modernization and consolidation of fragmented, bespoke evaluation pipelines into a unified core platform, identifying capability gaps and creating reusable, extensible frameworks.
- Optimize system throughput, execution turnaround times, and resource efficiency for continuous, high-volume evaluation workloads supporting production-critical Search models.
- Integrate modern AI tooling, automated validation harnesses, and AI-assisted workflows into daily engineering practices to boost developer velocity, evaluation excellence, and system robustness.
Minimum qualifications:
- Bachelor’s degree or equivalent practical experience.
- 5 years of experience with software development in one or more programming languages.
- 3 years of experience working with embedded operating systems.
- 3 years of experience testing, maintaining, or launching software products.
- 1 year of experience with software design and architecture.
Preferred qualifications:
- Master's degree or PhD in Computer Science or a related technical field.
- 5 years of experience with data structures and algorithms.
- 1 year of experience in a technical leadership role.
- Experience developing accessible technologies.