Pyspark Data Engineer
Summary
Tata Consultancy Services is hiring a PySpark Data Engineer (6-10 years) to design scalable PySpark/Hadoop test architectures, data validation and data quality systems for ETL pipelines. Day-to-day involves Hadoop/Hive/YARN environments, CI/CD test pipelines, performance tuning, and mentoring juniors. Work from office in Chennai/Kolkata/Hyderabad/Pune.
Dear Professionals
Greetings from Tata consultancy Services,
Job Title Pyspark Data Engineer
Experiernce: 6-10 Years
Location: Chennai / Kolkata / Hyderabad / Pune
Mode of Work : Work from Office
Job description
- Design scalable PySpark-based test architectures for ETL/data pipelines, including modular frameworks for batch processing.
- Architect end-to-end data validation systems in Hadoop environment for lineage, schema evolution
- Lead system design for Hadoop/Hive test environments, including YARN resource management, dynamic partitioning.
- Exposure to Zephyr-Jira-ServiceNow integrated test management systems with experience on API-driven automation.
- Design CI/CD test pipelines for PySpark/Hadoop jobs, incorporating artifact management, parallel execution, and blue-green deployments.
- Create data quality system designs using PySpark integrated with Hive metadata services.
- Design testing platforms, test data generators
- Mentor juniors on PySpark testing basics, contribute to testing strategy discussions
- Spark session configurations for memory and core allocations for both local and cluster manager settings
- Data handling with distributed file systems like HDFS and writing back to hive tables
- Implementation of Partitioning, caching techniques in organizing code for transformation pipelines
- Performance tuning implementation like salting, minimizing shuffling