Big Data Engineer, Web3
Summary
Build and maintain the data infrastructure for a top crypto exchange, including distributed compute/storage and AI-native scheduling agents that auto-tune pipelines and respond to incidents.
- Platform Core: Design and operate large-scale distributed data systems
- Own the big data compute and storage infrastructure (MaxCompute/ODPS, Hologres, Spark)
- Build and maintain multi-site task orchestration that dynamically selects engines and enforces policy
- Drive reliability and performance improvements across batch and real-time pipelines
- AI Integration: Build the AI-native platform layer
- Develop and expose MCP (Model Context Protocol) tool interfaces so AI agents can interact with platform APIs
- Build the scheduling and cost-optimization agents that auto-tune resource allocation and alert severity
- Instrument platform telemetry to feed AI-driven SLA monitoring and anomaly detection
- Design context retrieval pipelines (RAG / vector search) for SQL code and config knowledge bases
- Tooling & DX: Evolve the developer experience
- Own the internal data development platform — IDE integrations, code review automation, deployment tooling
- Build APIs-first tools (backfill, ingestion automation) designed for future MCP integration
- Collaborate with data warehouse and service teams to define platform contracts
- Ops & Governance: Drive operational excellence
- Establish SLA benchmarks, cost metrics, and latency dashboards as AI optimization targets
- Build automated incident response and root-cause analysis pipelines
- Define and enforce infrastructure policies across multi-cloud environments
- Scheduling Agent: auto-configure task dependencies, engine selection, cost/performance trade-offs, and alert tiers
- Operations Agent: detect pipeline latency, performance degradation, and schema drift; trigger remediation
- Incident Response Agent: trace SLA breaches to root cause, assign accountability, generate post-mortems
- MCP Tool Layer: design and maintain the cross-platform tool interfaces that all agents call into
- 5+ years of experience building large-scale data platforms (Hadoop/Spark/Flink or equivalent)
- Deep expertise in distributed storage and compute systems (MaxCompute, Hologres, ClickHouse, Hive)
- Strong software engineering skills in Java, Scala, or Python; experience with API-first design
- Hands-on experience with task scheduling systems (Airflow, DolphinScheduler, or in-house equivalents)
- Solid understanding of multi-cloud architectures and cost governance
- Familiarity with LLM integration patterns: tool calling, RAG pipelines, context management
- Experience with MCP or similar agent-tool frameworks is a strong plus
- Passion for building systems that make other engineers 10x more productive
- Competitive total compensation package
- L&D programs and education subsidy for employees' growth and development
- Various team building programs and company events
- Wellness and meal allowances
- Comprehensive healthcare schemes for employees and dependants
- More that we love to tell you along the process!