freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer

Summary

Build and own a structured research data platform for investment teams, ingesting and parsing documents (PDF, HTML, XBRL) and deploying LLM-powered query APIs on cloud data warehouses.

We are an early-state AI startup building a structured research data platform for the investment research domain. We serve professional investment teams focused on industry fundamental research. With our first batch of seed clients (mainstream hedge funds) already onboarded, we are a lean and highly efficient team.

We are looking for an early core engineer to work alongside the founding team to build this data platform from 0 to 1: from data ingestion and structured parsing, to query services for analysts. This is an end-to-end role where you will have complete ownership.

Core Responsibilities

1. Data Ingestion

  • Build and maintain multi-source data ingestion and pipelines.
  • Ensure pipelines are robust, stable, and production-scheduled.

2. Document Parsing

  • Extract structured data from formats like PDF, HTML, and XBRL.
  • Handle scanned PDFs, complex tables, and multilingual text (Chinese/English).

3. Data Warehouse Design

  • Design and maintain schema-based tables.
  • Leverage modern cloud data warehouses (Snowflake, BigQuery, or Postgres).
  • Build production-grade ETL/ELT using Airflow, Prefect, or Dagster.
  • Implement data quality checks, alerting, and data lineage tracking.

5. LLM Application Engineering

  • Implement OpenAI/Anthropic APIs and open-source models.
  • Productionize structured outputs and Retrieval-Augmented Generation (RAG).
  • Set up basic LLM evaluation and quality monitoring frameworks.
  • Build high-performance query APIs using FastAPI or similar frameworks.
  • Power downstream investment analyst tools.
  • Utilize cloud ecosystems (GCP or AWS).
  • Use managed services to minimize operational and DevOps overhead.

Key Requirements

  • 5–10 years of production-level software engineering experience.
  • Proven track record of building and deploying complete online systems.

2. Technical Stack

  • Strong proficiency in Python and SQL.
  • Solid experience in data modeling, writing tests, and performance tuning.

3. Document Processing

  • Hands-on experience parsing messy, real-world documents (PDF, HTML, XBRL).

4. AI & LLM Expertise

  • 2–3 years of hands-on experience deploying production LLM applications.
  • Proven experience in launching real-world AI features.
  • Mastery of at least one core orchestration framework (Airflow, Prefect, Dagster).
  • Proficiency in schema design for at least one modern cloud data warehouse.

6. Soft Skills & Communication

  • Strong communication skills to align with research and founding teams.
  • Ability to articulate technical designs clearly in discussions.
  • Professional working proficiency in English.

7. Strong Preferred

  • Previous experience as a "Founding Engineer".
  • Experience building end-to-end data systems at early-stage AI startups.

Preferred Qualifications (Bonus Points)

  • Experience building data infrastructure tailored for NLP/LLM workloads.
  • Hands-on experience with dbt (Data Build Tool).
  • Design experience in multi-tenancy and data access control.
  • Basic familiarity with graph databases (e.g., Neo4j).

What We Offer

  • Direct Impact: Your work forms the bedrock of our platform, directly impacting investment research daily.
  • Competitive Compensation: An attractive salary and equity package.
  • Elite Collaboration: Work directly with seasoned investment professionals (analysts, PMs, and quants).
  • High-Execution Culture: A rigorous, fast-paced, and execution-oriented work environment.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available