Data Engineer | Talent Bridge Dubai
Summary
A data engineer who builds and operates reliable batch/stream ETL/ELT pipelines (Airflow, optionally Spark) powering AI research and production, with strong data quality, lineage, and MLOps support (DVC, MLflow/W&B). Requires C++ or Java, SQL/NoSQL, PySpark, Docker, Linux, and cloud ML platform experience.
Build robust, observable data pipelines that power research and production AI. Success means high pipeline reliability (on-time SLAs), strong data quality (validation & lineage), and enabling fast experimentation. You will partner with AI/ML, analytics, and product to make data trustworthy and usable.
Responsibilities
- Architect and operate batch/stream pipelines (Airflow; Spark optional) for ETL/ELT.
- Model/manage schemas; enforce data quality and lineage/governance.
- Support ML workflows with DVC (data versioning) and MLflow or Weights & Biases.
- Build feature stores/data services; expose datasets via secure REST endpoints.
- Maintain documentation and internal catalogs; enable self-service analytics.
Qualifications
- Skills: Programming in C++ or Java; SQL & NoSQL; Pandas/NumPy; PySpark; Airflow; API development; Docker.
- MLOps: DVC; MLflow or W&B; model packaging/deployment fundamentals.
- Cloud: AWS SageMaker, Azure ML, or GCP AI experience.
- Nice to have: Unreal Engine exposure.
- Environment: Solid Linux background for development and deployment.
- Education/Experience: Proven experience building reliable pipelines in production