Python Developer
Summary
An experienced PySpark developer/lead role in Chennai focused on designing, building, and optimizing large-scale data pipelines and processing solutions with Apache Spark, Python, and Spark SQL/streaming with Kafka, deployed on cloud data platforms like AWS, Azure, or GCP.
Experience- 4-15 Years
Location: Chennai
Job Summary: An experienced PySpark Developer / Leads to design, develop, and maintain large-scale data processing solutions using Apache Spark and Python. The ideal candidate should have a strong background in data engineering, data processing, and cloud-based data platforms, big data ecosystems and performance optimization
Responsibilities:
- Data Processing: Design, develop, and maintain data processing solutions using PySpark, Apache Spark, and Python.
- Data Pipeline Development: Develop and optimize data pipelines using PySpark, Apache Spark, and cloud-based data platforms.
- Data Integration: Integrate data from various sources, including relational databases, NoSQL databases, and cloud storage.
- Data Transformation: Develop and implement data transformation logic using PySpark, Apache Spark, and Python.
- Collaboration: Work with cross-functional teams to identify and prioritize project requirements, provide technical guidance, and ensure data quality.
Required Skills:
- PySpark: In-depth knowledge of PySpark, Apache Spark, and Python.
- Data Processing: Strong understanding of data processing concepts, including data ingestion, data transformation, and data storage.
- Cloud Experience: Experience with cloud-based data platforms, including AWS, Azure, or Google Cloud.
- Expertise in DataFrames & Spark SQL, Spark Streaming with Apache Kafka real-time data ingestion pipelines
- Very good conceptual understanding of Multithreading, distributed computing concepts of Pyspark
- Programming: Proficiency in programming languages, including Python, Java, or Scala.
- Communication: Excellent communication and collaboration skills.
Preferred Skills:
- Apache Spark Certifications: Relevant certifications, such as Apache Spark Certification or Cloudera Certified Spark Developer.
- Big Data: Experience with big data technologies, including Hadoop, HBase, or Cassandra.
- Machine Learning: Experience with machine learning frameworks, including TensorFlow, PyTorch, or Scikit-learn.
- Data Science: Experience with data science tools, including Jupyter Notebook, Pandas, or NumPy