Senior Data Engineer
Summary
Senior Data Engineer builds and scales web-scale data pipelines for AI research, focusing on crawling, normalization, and metadata governance to power sovereign AI models like Ilmu.
About the Role
At YTL AI Labs, we build sovereign AI models that perform on par with the world’s best, while staying grounded in local needs, values, and context. Our flagship model, Ilmu, is designed to be culturally aware, contextually intelligent, and fluent in Bahasa Melayu, delivering cutting‑edge solutions that empower Malaysian businesses with intelligence that truly understands the market and the people they serve.
As pioneers of sovereign AI, we believe every nation should have the power to shape its own intelligence, guided by its people, priorities, and principles.
As the Senior Data Engineer on the Data Team, you will design and implement scalable, high‑fidelity data acquisition systems that power cutting‑edge AI research and deployment. You will architect end‑to‑end data pipelines that span diverse modalities, such as text, code, and multimedia, ensuring efficiency, reliability, and quality at web‑scale. In this role, you will mentor engineers, shape technical direction, and collaborate across research, infrastructure, and product teams to ensure our data ecosystem meets the evolving needs of large‑scale AI systems.
You’ll be responsible for
1. Strategic Web‑Scale Data Acquisition
- Architect and oversee the development of robust, distributed pipelines for crawling and ingesting structured and unstructured data from various sources (e.g., forums).
- Lead the design of scraping systems using Scrapy, Playwright, Selenium, etc., optimized for dynamic content.
- Define and enforce standards for data normalization and harmonization to ensure semantic consistency across datasets.
- Drive initiatives to expand multilingual and multimodal data coverage.
2. Metadata Strategy and Dataset Lineage
- Define metadata schemas and enrichment strategies (e.g., licensing, timestamps, source reliability).
- Build metadata validation and tracing systems to ensure transparency, reproducibility, and ethical dataset use.
3. Scalable, Secure, Resilient Web Interaction
- Develop web interaction strategies.
- Deploy and monitor scraping infrastructure for high‑throughput, fault‑tolerant operations.
4. Data Infrastructure and Governance
- Design modular data lake architectures and metadata‑aware repositories (e.g., Elasticsearch, HBase, Snowflake, MongoDB).
- Standardize dataset versioning and lineage using tools like DVC, Dolt, or HuggingFace Datasets.
- Establish IaC‑based deployment and monitoring using Terraform, Helm, Kubernetes, Airflow, etc.
5. Team Leadership and Cross‑Functional Collaboration
- Partner with AI research teams to align data priorities with model development goals.
- Work with infrastructure teams to optimize scalability, performance, observability, and cost.
- Ensure compliance with legal and responsible data sourcing standards.
What We’re Looking For
- Bachelor’s or Master’s degree in Computer Science, Data Science, Engineering, or related field.
- 2+ years of experience building and scaling complex data pipelines or backend systems, with leadership responsibility.
- Deep expertise in web crawling systems, distributed systems, and data normalization workflows.
- Strong understanding of metadata governance and dataset traceability in AI/ML workflows.
- Proficient in Python, SQL, and distributed compute frameworks (e.g., Spark, Dask); experienced with orchestrators like Airflow or Prefect.
- Experience with dataset versioning tools (DVC, Dolt, HuggingFace Datasets) and data stores (Elasticsearch, MongoDB, Snowflake, object storage).
- Strong communication and mentoring capability.
- Familiarity with NLP, LLMs, data ethics, or large‑scale dataset development is preferred.
If you’re looking to do meaningful work with people who care about how we get there, we’d love to meet you. Apply now!