freehire launches on Product Hunt on 26 August.

Follow →

Staff/Senior Staff Web3 Big Data Engineer

Summary

Build and own a large-scale Web3 big data platform that ingests on-chain transactions, trading behavior, and user profiles, then layer AI/ML tools for fraud detection, risk modeling, and natural-language data querying.

Who We Are

At the company, we believe that the future will be reshaped by crypto, and ultimately contribute to every individual's freedom. the company is a leading crypto exchange, and the developer of the company Wallet, giving millions access to crypto trading and decentralized crypto applications (dApps). the company is also a trusted brand by hundreds of large institutions seeking access to crypto markets. We are safe and reliable, backed by our Proof of Reserves. Across our multiple offices globally, we are united by our core principles: We Before Me, Do the Right Thing, and Get Things Done. These shared values drive our culture, shape our processes, and foster a friendly, rewarding, and diverse environment for every OK-er. the company is part of OKG, a group that brings the value of Blockchain to users around the world, through our leading products the company, the company Pay, the company Wallet, OKLink and more.

Responsibilities

  • Platform Architecture: Own the architecture, development, and optimization of our big data platform, supporting large-scale collection, processing, and analysis of on-chain data, trading behavior, and user profiles.
  • On-Chain Data Pipelines: Design and build real-time/batch data pipelines for on-chain data — blockchain transactions, smart contract events, wallet address behavior.
  • Data Warehouse & Lake: Build and maintain the data warehouse/lake, define data-layering standards, and ensure data quality, consistency, and timeliness.
  • AI-Driven Data Applications: Combine big data with LLM capabilities to build intelligent applications — on-chain anomaly/fraud detection, smart risk models, user behavior prediction, and natural-language data querying (Text2SQL).
  • AI Infrastructure: Build data infrastructure for AI use cases — vector databases, feature platforms, and Embedding pipelines supporting RAG retrieval and Agent data supply.
  • Cross-Team Collaboration: Partner with Algorithm/AI teams on large-scale data processing and pipelines for model training data and feature engineering; partner with Product, Risk, and Growth teams on data needs for trading analytics, anti-fraud, growth, and operations.
  • Performance & Reliability: Optimize performance and resource efficiency of big data jobs, ensuring stability and SLAs on core pipelines.
  • Technical Direction: Track industry developments in big data, Web3 data infrastructure, and AI, driving technology selection and architecture evolution.
  • Global Collaboration: Communicate and document in English with overseas colleagues and partners (exchanges, public chain teams, etc.).

Requirements

Big Data Engineering

  • Bachelor's degree or above in Computer Science, Software Engineering, or related field; 7+ years of big data development experience.
  • Proficiency with Hadoop, Spark, and Flink, with experience building batch/real-time data warehouses.
  • Familiarity with Hive, Kafka, HBase, ClickHouse, and Doris, with large-scale cluster tuning experience.
  • Strong big data development skills in Java/Scala/Python, with solid SQL tuning ability.
  • Experience with data governance, data quality monitoring, or metadata management is a plus.

AI Capabilities

  • Understanding of LLM fundamentals and application patterns, with hands‑on experience in prompt engineering and RAG.
  • Experience with vector databases (e.g. Milvus, Pinecone, Weaviate, pgvector) or Embedding data processing.
  • Familiarity with emerging AI application architectures such as AI Agents or MCP (Model Context Protocol) is a plus.
  • Experience connecting big data platforms with AI/ML training and inference pipelines (e.g. feature platforms, real-time feature serving).
  • Experience using LLMs to accelerate data engineering (automated data quality checks, intelligent ETL generation, Text2SQL) is a plus.
  • Basic ML/deep learning knowledge and ability to collaborate effectively with algorithm teams is a plus.

Web3 Industry Knowledge

  • Un

See also