Senior Data & AI Engineer
Summary
Senior Data & AI Engineer building a structured research data platform for investment research. Involves data pipelines, LLM applications, backend services, and cloud infrastructure.
About Our Client
Our client is an early-stage AI startup building a next-generation structured research data platform for the investment research industry .
The company works with professional investment teams conducting fundamental equity research and has already onboarded its first group of institutional clients, including established hedge funds.
With a lean, highly technical team, they are now looking for an early core engineer to work closely with the founding team and build the platform from 0 to 1. This is a hands-on role with significant ownership across the entire data and AI stack — from data ingestion and document intelligence to LLM applications and analyst-facing query services.
About the Role
You will be responsible for building the core data infrastructure powering the company's AI-driven investment research platform.
This is an end-to-end engineering role spanning data pipelines, document processing, data architecture, LLM applications, backend services, and cloud infrastructure . You will have the opportunity to influence key technical decisions and help establish the engineering foundations as the company scales.
What You'll Do
Data Ingestion & Pipelines
- Build and maintain robust, production-grade pipelines across multiple external data sources.
- Design reliable ingestion workflows with scheduling, monitoring, and failure recovery.
Document Processing
- Build systems to extract and structure information from PDF, HTML, XBRL, and other real-world documents .
- Handle scanned PDFs, complex tables, and multilingual content in both Chinese and English.
Data Architecture
- Design and maintain structured, schema-based data models.
- Build on modern data warehouses and databases such as Snowflake, BigQuery, or PostgreSQL .
- Ensure data structures remain scalable, reliable, and efficient to query.
ETL/ELT Orchestration & Data Quality
- Build production-grade ETL/ELT workflows using Airflow, Prefect, Dagster , or similar technologies.
- Implement data quality checks, alerting, observability, and lineage tracking.
LLM Application Engineering
- Build production applications using OpenAI, Anthropic, and/or open-source LLMs .
- Develop reliable structured-output and Retrieval-Augmented Generation (RAG) workflows.
- Implement practical LLM evaluation, quality monitoring, and reliability mechanisms.
Backend Engineering
- Build high-performance APIs and query services using FastAPI or similar frameworks.
- Develop backend services powering downstream investment research and analyst-facing products.
Cloud Infrastructure
- Build and operate services within AWS or GCP .
- Leverage managed cloud services to maintain a lean and efficient infrastructure.
What We're Looking For
- 5–10 years of production software engineering experience , with a strong track record of building and deploying complete systems.
- Strong proficiency in Python and SQL .
- Strong understanding of data modeling, testing, system reliability, and performance optimization .
- Hands-on experience building production data pipelines using Airflow, Prefect, Dagster , or similar orchestration frameworks.
- Experience designing schemas and data models using Snowflake, BigQuery, PostgreSQL , or similar technologies.
- Hands-on experience processing messy, real-world documents such as PDFs, HTML, XBRL, scanned documents, or complex tables .
- 2+ years of hands-on experience building production LLM applications , ideally including RAG, structured outputs, or other real-world AI features.
- Experience working with AWS or GCP .
- Strong communication skills with the ability to clearly articulate technical designs and trade-offs.
- Professional working proficiency in English.
- Experience as a Founding Engineer, early-stage engineer, or core member of a startup engineering team is highly valued.
- Experience building end-to-end data platforms or infrastructure from 0 to 1 is a strong plus.
- Experience building data infrastructure specifically for NLP or LLM workloads is a plus.
- Experience working with financial, investment research, or other complex document-heavy datasets is advantageous.
- Experience with dbt , multi-tenant system design, or data access control is a plus.
- Familiarity with graph databases such as Neo4j is advantageous.
- Experience with OCR, document intelligence, table extraction, or financial filings/XBRL is a plus.
AIM Global Talent Pte. Ltd. | EA Licence No. 25C3207