Senior Batch Data Engineer (Alibaba Cloud Stack)
Summary
Senior data engineer with end-to-end ownership of OneBullEx's T-1 enterprise batch data warehouse on Alibaba Cloud: designing and optimising ETL/ELT pipelines and warehouse layers (ODS, DWD, DWS, ADS) with DataWorks, MaxCompute and Hologres using SQL and Python, enforcing data quality and monitoring, and collaborating bilingually in English and Chinese.
OneBullEx is looking for a highly accountable, self-driven Senior Batch Data Engineer to take end-to-end ownership of our enterprise batch data warehouse on Alibaba Cloud. In this role, you will design, maintain, and optimise batch pipelines (T-1) supporting executive dashboards, business analytics, and AI/ML services — leveraging modern AI-assisted engineering techniques to rapidly build data layers (ODS, DWD, DWS, ADS) and enforcing strict data quality, automated validation, and proactive monitoring across the platform.
Key Responsibilities
Pipeline Architecture & Data Warehouse Modelling
- Design, build, and maintain scalable ETL/ELT batch pipelines using Alibaba Cloud DataWorks and MaxCompute.
- Rapidly design and build enterprise data warehouse layers (ODS, DWD, DWS, ADS) using modular SQL/Python scripts and AI generation tools.
- Develop and optimise data ingestion, transformation, and batch processing workflows across Lakehouse and Data Warehouse architectures (MaxCompute, Hologres, DataWorks, DTS).
- Ensure data reliability, high query performance, schema stability, and cost-effective cloud resource usage.
AI-Accelerated Engineering & Testing
- Integrate modern AI tools and techniques (e.g. Claude, GitHub Copilot, prompt engineering) into daily workflow to accelerate code generation, SQL refactoring, and data validation.
- Apply new AI testing methodologies to rapidly write unit tests, simulate data scenarios, and audit data accuracy prior to production release.
Data Quality, Monitoring & Alerting
- Implement strict data quality rules, automated reconciliation scripts, and validation checks directly inside DataWorks jobs.
- Build proactive alerting and monitoring workflows (integrated with Lark) to detect data drift, schema breaks, or pipeline failures before business impact occurs.
- Support downstream analytics, BI reporting, and AI/LLM services with clean, trusted datasets.
End-to-End Ownership & Collaboration
- Take full ownership of daily T-1 batch pipeline stability, job scheduling, historical backfills, and legacy issue cleanup.
- Act as a bilingual technical bridge, collaborating fluently in both English and Chinese with cross-functional technical leads, product managers, and developers.
Required Qualifications
- 5+ years of hands-on expertise with Alibaba Cloud data services: DataWorks, MaxCompute, DTS, Hologres, FC Functions, Data Agents, and MaxCompute Lakehouse architectures.
- Advanced proficiency in SQL and Python for data engineering, performance tuning, and database optimisation.
- Proven experience modelling multi-layer enterprise data structures (ODS, DWD, DWS, ADS) and managing batch Lakehouse architectures.
- Fluent in both English and Chinese, spoken and written, for seamless technical collaboration with regional teams.
- Strong sense of ownership, high velocity, proactive communication, and a "deliver fast, iterate continuously" attitude.
Preferred
- Active experience leveraging AI coding assistants (Copilot, LLMs) to speed up ETL development, code reviews, and query optimisation.
- Experience building alerting and monitoring integrations with Lark or a similar collaboration platform.
- Track record applying AI-driven testing methodologies ahead of production release.
Skills
As published by greenhouse · 3 questions
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter
- Preferred First Name optional
- LinkedIn Profile optional
- Website optional