Data Scientist (NLP/ GENAI Evaluation)
Summary
Build and run evaluation frameworks for generative AI and NLP models, analyze performance gaps, and work with engineers to deploy improvements in a public-sector setting.
Responsibilities
- Collaborate with policy and communications officers, product owners, engineers, domain experts and subject-matter specialists to understand user requirements and translate them into well-defined analytical and machine-learning problems.
- Build and maintain robust evaluation frameworks and datasets for natural-language, generative AI and other machine-learning applications.
- Plan and carry out experiments to evaluate and enhance model quality, accuracy, consistency, reliability, latency and cost.
- Assess suitable models, methods and emerging technologies, and recommend solutions based on evidence, user requirements and operational factors.
- Conduct systematic error analysis, identify performance gaps across different use cases and user segments, and work with the team to prioritise areas for improvement.
- Develop appropriate automated and human-evaluation methods, while recognising the limitations and risks associated with individual metrics and AI-assisted evaluation.
- Work closely with engineers to implement validated improvements, establish quality checks and monitor performance in production environments.
- Ensure that data, experiments and model-related decisions are reproducible, properly documented and aligned with responsible AI, privacy and security requirements.
- Present findings, trade-offs and recommendations clearly to both technical and non-technical stakeholders.
Requirements:
- A degree in Computer Science, Data Science, Statistics, Artificial Intelligence, Computational Linguistics or a related quantitative field, or equivalent practical experience.
- Proven experience applying data science or machine learning to real-world problems, preferably in natural language processing, generative AI, search or information retrieval.
- Strong Python programming skills, together with working knowledge of SQL, data processing, version control and software-development practices.
- A solid understanding of statistics, experimental design, evaluation methods, sampling, error analysis and model validation.
- Experience working with unstructured text or other complex data types, as well as evaluating machine-learning or generative AI systems using more than a single aggregate metric.
- Familiarity with current NLP and AI concepts, including embeddings, language models, prompt design and model evaluation.
- The ability to develop maintainable code and collaborate with engineers to deploy data-science solutions into production.
- Strong analytical, problem-solving and communication capabilities, with the ability to explain technical findings and trade-offs clearly to a range of audiences.
Preferred Skills
- Experience in multilingual NLP, translation-quality evaluation, or working alongside linguists and language reviewers.
- Experience with cloud-based AI services, vector search, MLOps, production monitoring or responsible AI practices.
- Experience developing AI-enabled products within government, communications, regulated sectors or other high-assurance environments.
- An understanding of Singapore’s public communications landscape and the requirements of public-sector stakeholders.