Big Data Engineering Lead
Job Description
We are seeking an experienced Big Data Engineering Lead to design, develop and lead scalable enterprise data platforms, real-time data pipelines and analytics solutions. The successful candidate will work with large-volume datasets, cloud platforms, distributed processing technologies and modern AI/GenAI solutions.
Key Responsibilities
- Lead the design and development of scalable Big Data Engineering and Data Analytics platforms.
- Design and implement batch and real-time data pipelines using Python, PySpark, Spark Structured Streaming and Kafka.
- Develop data ingestion, transformation and processing workflows using AWS, S3, Airflow and orchestration tools.
- Design data models, data warehouses, data lakes/lakehouse architectures and high-performance analytics solutions.
- Develop and optimize complex SQL queries, ETL processes and data processing workflows.
- Build REST APIs and backend services using FastAPI/Flask for data and analytics applications.
- Implement data replication, archival, reconciliation, quality monitoring and governance processes.
- Deploy scalable data applications on AWS, including containerized environments and CI/CD pipelines.
- Lead development of GenAI/LLM, RAG and Agentic AI solutions for enterprise analytics and natural-language data querying.
- Provide technical leadership, mentoring and guidance to data engineering teams.
- Collaborate with business and technology stakeholders to translate requirements into scalable data solutions.
Requirements
- Degree in Computer Science, Information Technology, Engineering, Data Science or a related discipline.
- 8+ years of experience in Data Engineering / Big Data / Data Analytics.
- Strong hands-on experience with Python, SQL, PySpark, Apache Spark and Kafka.
- Experience in AWS cloud data services, including S3 and related data engineering technologies.
- Strong experience with Airflow or equivalent workflow orchestration tools.
- Experience with Snowflake, Hadoop/HDFS, Hive or other enterprise data platforms.
- Strong understanding of ETL, data warehousing, data modelling and distributed data processing.
- Experience developing REST APIs using FastAPI or Flask.
- Exposure to GenAI, LLM, RAG, LangChain/LangGraph or Agentic AI is an advantage.
- Strong analytical, problem-solving and technical leadership skills.
- Experience working in Agile software development environments.
Skills
- Agentic AI
- Agile
- AI
- Airflow
- Analytics
- API
- AWS
- CI/CD
- Cloud
- Data Analytics
- Data Engineering
- Data Ingestion
- Data Modeling
- Data Pipelines
- Data Science
- Data Warehousing
- ETL
- FastAPI
- Flask
- Generative AI
- Hadoop
- Hive
- Kafka
- Lakehouse
- LangChain
- LangGraph
- LLM
- PySpark
- Python
- REST
- Snowflake
- Spark
- SQL
- Workflow Orchestration