Senior Data Engineer
· 8-12 years of experience in Data Engineering, Big Data, Data Lake, or Lakehouse implementations.
· Hands-on experience with Databricks, Snowflake, Cloudera, Azure, AWS, GCP, Huawei, or Alibaba data platforms.
· Hands-on experience in developing Data products and marketplaces
· Strong expertise in Spark, PySpark, SQL, Python and Scala.
· Strong programing skills (Java, Scala, Python, SQL)
· Experience with Iceberg, Hudi, Delta Lake and object storage platforms.
· Experience implementing data ingestion, transformation, reconciliation and data quality frameworks.
· Experience with Trino, Dremio, Hive, Impala, Kafka, Flink, Spark Streaming and Airflow.
· Strong hands-on experience with Kubernetes, OpenShift, Docker, CI/CD, MLflow and observability tools.
· Ability to design data architectures supporting NLP and AI‑driven analytics, including ingestion, curation, and governance of unstructured data within Data Lake, Data warehouse platforms.
· Experience working with MLplatforms such as CML, Spark MLlib, and Python ML libraries (scikit‑learn, XGBoost), including model deployment.
· Develop full‑stack applications and internal engineering tools using Python, shell scripting, and modern web frameworks (e.g., Flask, React).
· Knowledge of data modelling, metadata management, lineage and governance.
· Experience exposing data through APIs, eventstreams, dashboards and BI platforms.
· Knowledge of Teradata, Netezza, Greenplum or MPPmigration programs is advantageous.
· Experience with Kubernetes, OpenShift, Terraform, Jenkins, Git and CI/CD pipelines.
EA Reg.No. 25C2690 | EA License No. R1877766
Skills
- AI
- Airflow
- Analytics
- API
- AWS
- Azure
- Bash
- CI/CD
- Data Engineering
- Data Ingestion
- Data Lake
- Data Modeling
- Data Quality
- Data Warehousing
- Databricks
- Delta Lake
- Docker
- Flask
- Flink
- GCP
- Git
- Hive
- Java
- Jenkins
- Kafka
- Kubernetes
- Lakehouse
- Machine Learning
- Metadata Management
- MLflow
- Model Deployment
- NLP
- Observability
- OpenShift
- PySpark
- Python
- React
- Scala
- scikit-learn
- Snowflake
- Spark
- SQL
- Teradata
- Terraform
- Trino
- XGBoost