freehire launches on Product Hunt on 26 August.

Follow →

Data Scientist (NLP/ GENAI Evaluation)

Summary

Build and run evaluation frameworks for generative AI and NLP models, analyze performance gaps, and work with engineers to deploy improvements in a public-sector setting.

Responsibilities

  • Collaborate with policy and communications officers, product owners, engineers, domain experts and subject-matter specialists to understand user requirements and translate them into well-defined analytical and machine-learning problems.
  • Build and maintain robust evaluation frameworks and datasets for natural-language, generative AI and other machine-learning applications.
  • Plan and carry out experiments to evaluate and enhance model quality, accuracy, consistency, reliability, latency and cost.
  • Assess suitable models, methods and emerging technologies, and recommend solutions based on evidence, user requirements and operational factors.
  • Conduct systematic error analysis, identify performance gaps across different use cases and user segments, and work with the team to prioritise areas for improvement.
  • Develop appropriate automated and human-evaluation methods, while recognising the limitations and risks associated with individual metrics and AI-assisted evaluation.
  • Work closely with engineers to implement validated improvements, establish quality checks and monitor performance in production environments.
  • Ensure that data, experiments and model-related decisions are reproducible, properly documented and aligned with responsible AI, privacy and security requirements.
  • Present findings, trade-offs and recommendations clearly to both technical and non-technical stakeholders.

Requirements:

  • A degree in Computer Science, Data Science, Statistics, Artificial Intelligence, Computational Linguistics or a related quantitative field, or equivalent practical experience.
  • Proven experience applying data science or machine learning to real-world problems, preferably in natural language processing, generative AI, search or information retrieval.
  • Strong Python programming skills, together with working knowledge of SQL, data processing, version control and software-development practices.
  • A solid understanding of statistics, experimental design, evaluation methods, sampling, error analysis and model validation.
  • Experience working with unstructured text or other complex data types, as well as evaluating machine-learning or generative AI systems using more than a single aggregate metric.
  • Familiarity with current NLP and AI concepts, including embeddings, language models, prompt design and model evaluation.
  • The ability to develop maintainable code and collaborate with engineers to deploy data-science solutions into production.
  • Strong analytical, problem-solving and communication capabilities, with the ability to explain technical findings and trade-offs clearly to a range of audiences.

Preferred Skills

  • Experience in multilingual NLP, translation-quality evaluation, or working alongside linguists and language reviewers.
  • Experience with cloud-based AI services, vector search, MLOps, production monitoring or responsible AI practices.
  • Experience developing AI-enabled products within government, communications, regulated sectors or other high-assurance environments.
  • An understanding of Singapore’s public communications landscape and the requirements of public-sector stakeholders.

See also