Data Engineer
Numentica Data Engineer
Summary
Onsite Data Engineer in Saidapet (Chennai) who designs, builds, and maintains scalable ETL pipelines and cloud data infrastructure on AWS/GCP/Azure. Core stack includes SQL, Python/Java/Scala, and big data frameworks like Apache Spark, Hadoop, and Kafka, with a focus on data quality, governance, and pipeline performance.
- Design, develop, and maintain scalable ETL pipelines to process and transform large-scale datasets.
- Integrate structured and unstructured data from multiple sources, ensuring quality, security, and consistency.
- Collaborate with data scientists, analysts, and software engineers to deliver well-structured and accessible datasets.
- Build and optimize data infrastructure using big data technologies such as Apache Spark, Hadoop, and Kafka.
- Deploy and manage cloud-based data solutions on AWS, GCP, or Azure.
- Monitor and troubleshoot data pipeline performance, ensuring reliability and efficiency.
- Implement data governance, security, and compliance best practices.
- Drive automation, testing strategies, and continuous improvements in data engineering workflows.
- 5 years of related experience with a Bachelor’s degree or equivalent work experience.
- Advanced proficiency in SQL and experience with relational and NoSQL databases (PostgreSQL, MySQL, MongoDB, etc.).
- Strong programming skills in Python, Java, or Scala for data processing and automation.
- Deep expertise in ETL processes, data modeling, and data warehousing.
- Hands-on experience with big data frameworks such as Apache Spark, Hadoop, or Kafka.
- Proficiency in cloud platforms (AWS Redshift, Google BigQuery, Azure Synapse) and data infrastructure automation.
- Experience optimizing data pipeline performance and scalability.
- Strong problem-solving skills with the ability to work on complex, large-scale datasets.
- Knowledge of data governance, security, and compliance best practices.
- Excellent leadership, collaboration, and communication skills to work effectively across teams.