AI Data Scientist
Summary
Build and deploy AI models (including LLMs) for investment research, using Python, ML libraries, and financial data to enhance decision-making and automate analysis.
AI Data Scientist
Location: Hong Kong, Hong Kong
Department: Trading
AI Data ScientistABOUT THE TEAM AND ROLE
The investment team is building an AI-enabled research and decision platform that brings together proprietary knowledge, public information, market and alternative data, analytical tools and modern machine-learning capabilities.
We are seeking an AI Data Scientist to work directly with portfolio managers, investment professionals and engineers. The role will own high-impact projects from problem definition through model development, deployment, evaluation and ongoing improvement. The successful candidate will combine scientific depth with strong engineering judgement and a practical understanding of how data and AI can improve investment research and decision-making.
PRINCIPAL RESPONSIBILITIES
- Translate investment and research questions into well-defined data-science problems, measurable objectives and practical technical solutions.
- Develop and deploy machine-learning, statistical and LLM-enabled models for company research, industry analysis, market monitoring, event detection and knowledge discovery.
- Build robust workflows across the full data lifecycle, including data sourcing, cleaning, transformation, feature engineering, quality checks, modelling and monitoring.
- Develop retrieval, search and knowledge systems using structured and unstructured data, with rigorous source attribution and evaluation.
- Design experiments and evaluation frameworks covering model quality, factual accuracy, robustness, latency, cost and user impact.
- Work with engineers to productionize models and analytical tools through APIs, batch pipelines and monitored applications.
- Partner closely with investment users to understand workflows, communicate trade-offs and iterate based on evidence and feedback.
- Identify promising models, research and open-source technologies, and determine when they are—or are not—appropriate for real investment use cases.
- Improve tooling, documentation and processes to increase reliability, reduce manual work and enable reuse across the team.
QUALIFICATIONS / SKILLS REQUIRED
- PhD in Computer Science, Machine Learning, Artificial Intelligence, Statistics, Applied Mathematics, Engineering, Physics or another highly quantitative field.
- Minimum 2 years of professional, full-time experience in data science, machine learning, applied AI or a closely related role. Doctoral research alone does not replace the professional-experience requirement.
- Strong Python proficiency and experience with core scientific and machine-learning libraries such as Pandas, NumPy, scikit-learn and PyTorch or equivalent frameworks.
- Strong grounding in machine-learning and statistical fundamentals, including problem framing, experimental design, validation, metrics, feature engineering, overfitting and uncertainty.
- Demonstrated experience delivering at least one end-to-end model or data product used by real stakeholders, from initial scoping through deployment and monitoring.
- Practical experience with LLMs and modern NLP, including retrieval-augmented generation, embeddings, vector search, prompt or context design and systematic evaluation.
- Proficiency in SQL and experience working with relational, columnar or document-oriented data systems.
- Ability to work with messy, incomplete and heterogeneous data while maintaining strong standards for data quality, testing, reproducibility and documentation.
- Strong written and verbal communication skills, with professional fluency in English and Mandarin.
PREFERRED QUALIFICATIONS
- Experience with financial, market, regulatory or alternative datasets, or with research-intensive decision environments.
- Experience building data or AI products in cloud environments and deploying APIs, batch jobs, monitoring or feedback loops.
- Familiarity with knowledge graphs, document processing, browser automation, data visualization or time-series and event-driven modelling.
- Evidence of technical depth through publications, patents, open-source contributions or substantial production projects.
- Genuine interest in companies, industries and investing; prior investment experience is valued but not required.
HOW WE WORK
- Ownership: Scope work clearly, set realistic milestones, communicate risks early and follow through on outcomes.
- Scientific rigour: Prefer measurable evidence, reproducible analysis and honest uncertainty over impressive demonstrations.
- Practical judgement: Start with the simplest viable approach, use advanced methods where they add value and understand when not to use AI.
- Collaboration: Work closely with investment and technology colleagues, seek feedback and communicate complex ideas clearly.
- Continuous improvement: Track developments in models, research and open-source tooling, and translate relevant advances into durable capabilities.