Blockchain Data Analyst & Researcher
Summary
Research blockchain protocols, reverse-engineer DeFi smart contracts, and build real-time data pipelines to extract metrics like TVL and volume from on-chain data.
You will research the internals of blockchain networks, reverse-engineer DeFi protocols, and translate complex on-chain interactions into structured, production-ready data. You will architect intuitive data models, design schemas for new blockchains, and build SQL/dbt models that deliver accurate metrics such as TVL, volume, and fees. You will define data quality standards, filter out wash trading, bot activity, and Sybil attacks, and build scalable real-time pipelines using tools like Flink, Spark Streaming, and data lakehouse platforms. You will write advanced SQL and Python to process large, semi-structured datasets, and use APIs and node queries to ingest external data.
Responsibilities
- Research blockchain protocols and their internals
- Reverse-engineer DeFi smart contracts and on-chain interactions
- Architect data models and schemas for blockchain data
- Index new blockchains and expand ecosystem coverage
- Map, cleanse, and normalize raw on-chain data into structured datasets
- Define key metrics such as TVL, volume, and fees
- Filter out wash trading, bot activity, and Sybil attacks
- Build production-grade SQL/dbt models and data pipelines
- Maintain data quality and integrity across analytics and machine learning systems
Requirements
- 3+ years of hands-on experience in Data Science, Data Engineering, or a hybrid role
- Blockchain or crypto analytics background, with familiarity with on-chain data structures, transaction semantics, and asset classification
- Experience building and operating real-time data pipelines with Apache Flink or Spark Streaming
- Experience managing large-scale data lakehouses with Apache Iceberg or Apache Hudi
- Experience with AWS S3 and RDS
- Proficiency in Spark SQL and PySpark
- Familiarity with data grid technologies such as Hazelcast, Apache Ignite, or Redis
- Advanced SQL skills including window functions, query plan analysis, and performance tuning
- Proficiency in Python with Pandas, Polars, and Scikit-learn
- Experience architecting high-throughput ingestion pipelines using REST APIs, WebSocket APIs, and JSON-RPC node queries
- Strong understanding of ETL fundamentals and schema mismatches
- Systematic approach to debugging and root-cause analysis
- Comfort working with large, semi-structured, or undocumented data sources