Senior Applied Data Scientist | NDA
Summary
Senior Applied Data Scientist focused on improving entity resolution at scale by developing and evaluating ML, embedding, and LLM-based matching approaches for complex business records, partnering with engineering to productionize successful models.
GT was founded in 2019 by a former Apple, Nest, and Google executive. GT’s mission is to connect the world’s best talent with product careers offered by high-growth companies in the UK, USA, Canada, Germany, and the Netherlands.
On behalf of our client, GT is looking for a Senior Applied Data Scientist interested in developing and testing new ML, embedding, and LLM-based approaches to solve complex data matching problems at scale.
About the Client
Our client is a leading global management consultancy known for tackling some of the world’s most complex business challenges. With a focus on strategy, transformation, and performance improvement, the firm partners with major organizations across industries to drive lasting impact.
About the Role
We are looking for a Senior Applied Data Scientist to improve how entity resolution is performed at scale.
You will develop and test new ML, embedding, and LLM-based approaches for matching complex business records across multiple data sources.
The work is centered on model quality, experimentation, and evaluation; engineering partners will help productionize successful approaches.
A key part of the role is exploring how newer foundation-model techniques can improve matching quality while remaining practical and scalable for very large datasets.
Responsibilities:
Develop better ways to match company records
Build new ML, embedding, and LLM-based approaches for matching entities
Improve how the system handles messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies.
Develop scoring and ranking approaches to distinguish accurate matches from duplicates, similar-looking records, and unrelated entities.
Evaluate and implement AI and machine learning techniques to improve matching quality while considering accuracy, scalability, and cost.
Design approaches that can operate efficiently at scale, taking model usage and computational cost into consideration.
Improve evaluation, experimentation, and match quality
Define and improve methods for evaluating match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review effort.
Assist in building trusted benchmark sets that allow us to compare new models against the current matching engine before production rollout.
Explore LLM-assisted review and validation to assess matching performance and benchmark more scalable approaches.
Turn ambiguous matching problems into clear hypotheses, experiments, metrics, and recommendations.
Partner with engineering to bring successful ideas into production
Work closely with data engineering and software engineering teams to turn promising prototypes into production-ready matching logic.
Provide engineering partners with clear model specifications, evaluation results, expected behavior, edge cases, and rollout requirements.
Help determine the most appropriate matching techniques based on data characteristics, confidence levels, and cost considerations.
Continuously evaluate matching performance, investigate regressions, and recommend improvements to models and matching logic.
Clearly communicate technical tradeoffs related to matching performance, scalability, cost, latency, explainability, and operational considerations.
Essential knowledge, skills & experience:
5–8 years of relevant experience in Data Science, Applied Data Science, Applied Machine Learning, or a similar role.
Strong applied ML fundamentals, with hands-on experience building and evaluating models on real data.
Excellent Python and SQL skills.
Practical experience with embeddings, semantic similarity, LLMs, or related AI techniques.
Hands-on experience training supervised and unsupervised models, including classification and NLP tasks.
Working knowledge of neural network and transformer architectures.
Proficiency with common ML frameworks such as TensorFlow, PyTorch, and PyCaret.
Experience retraining a taxonomy classifier or maintaining classification models in production.
Experimental judgment: able to define baselines, metrics, test sets, and error analysis that show whether quality improved.
Ability to explain model behavior, tradeoffs, and edge cases clearly to engineering and business partners.
Nice-to-have:
Experience with entity resolution, record linkage, deduplication, or similar matching problems.
Experience with ranking, similarity scoring, retrieval, clustering, or candidate generation.
Experience applying LLMs or embeddings to business problems where cost and scale matter.
Exposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQuery.
Familiarity with company, domain, website, firmographic, or other business-entity data.
Interview Steps:
GT interview with Recruiter
Technical interview
Final interview
As published by ashby
Resume, Full Name, Email, Location
- Where have you heard about this opportunity?
- Please add the link to your LinkedIn page optional
- Cover Letter OR Additional Information written answer · optional