Data Engineer
Summary
Senior Data Engineer (6-8 yrs) designing and maintaining scalable batch and event-driven ETL/ELT data pipelines using Python, advanced SQL, and AWS data services (S3, Glue, Lambda, Redshift, RDS), with data-quality, security, and analytics-readiness work. Hybrid role in Hyderabad, contract-to-hire through an IT consulting/staffing firm; Azure data services and Power BI/Tableau exposure are preferr
- Jobseeker Video Testimonials
- Employee Glassdoor Reviews
We are an IT Solutions Integrator/Consulting Firm helping our clients hire the right professional for an exciting long-term project. Here are a few details.
Experience:6-8 Years
Requirements
Job Description – Senior Data Engineer
Role Summary
We are looking for an experienced and hands-on Data Engineer to design,
develop, and maintain scalable data pipelines and analytics solutions. The
successful candidate will have strong programming capabilities in Python,
advanced SQL expertise, and practical experience working with AWS-based data
engineering services.
The role involves integrating data from multiple structured and unstructured
sources, transforming and organizing large datasets, implementing data-quality
controls, and preparing reliable data layers for analytics, reporting, and
business intelligence solutions.
Key Responsibilities
· Design, develop, test, and maintain scalable batch and
event-driven data pipelines.
· Build end-to-end ETL/ELT workflows using Python, SQL, and
AWS data services.
· Develop reusable Python components for data ingestion,
transformation, validation, reconciliation, and automation.
· Write and optimize complex SQL queries, stored procedures,
views, functions, and data-transformation logic.
· Integrate data from databases, APIs, enterprise
applications, flat files, and cloud-based file systems.
· Work with file-based datasets in formats such as CSV, JSON,
XML, Parquet, and Excel.
· Query and analyse data stored in databases.
· Work with relational databases hosted on Amazon RDS, Azure
Synapse, including Microsoft SQL Server or PostgreSQL.
· Implement incremental data loads, change-data-capture
patterns, error-handling mechanisms, restartability, and audit controls.
· Perform data validation, reconciliation, profiling,
cleansing, and quality monitoring.
· Apply appropriate security controls, including IAM-based
access, encryption, logging, and secure handling of sensitive data.
· Collaborate with Power BI or Tableau developers to deliver
analytics-ready datasets and semantic models.
Mandatory Technical Skills
Programming
· Strong hands-on programming experience in Python.
· Good understanding of object-oriented programming, modular
development, exception handling, logging, and reusable coding practices.
· Prior exposure to another programming language, preferably
Java, is desirable.
· Experience using Python libraries for data processing and
database integration, such as Pandas, PySpark, Boto3, SQLAlchemy, or equivalent
libraries.
SQL and Databases
· Strong proficiency in writing and optimizing advanced SQL
queries.
· Sound knowledge of relational database concepts,
normalization, denormalization, indexing, and database performance.
· Hands-on experience with Complex joins and subqueries, Window
and analytical functions, Stored procedures, functions, and views, Query-performance
optimization & Dimensional and relational data modelling
AWS Data Engineering
· Hands-on experience with several of the following services:
- Amazon S3
- AWS Glue
- AWS Lambda
- AWS Step Functions
- Amazon Redshift
- Amazon RDS
- AWS IAM
Preferred Skills
· Hands-on experience or working knowledge of Microsoft Azure
data services, including:
o Azure Data Factory
o Azure Synapse Analytics
o Microsoft Fabric
o Azure Data Lake Storage
o Azure Functions
· Exposure to business intelligence and reporting tools,
particularly Microsoft Power BI or Tableau
Benefits
Skills
- Analytics
- API
- Automation
- AWS
- Aws Glue
- Azure
- Azure Data Factory
- Azure Functions
- Azure Synapse
- Cloud
- Data Engineering
- Data Ingestion
- Data Lake
- Data Modeling
- Data Pipelines
- Data Quality
- ELT
- ETL
- Event Driven Architecture
- IAM
- Java
- JSON
- Lambda
- Microsoft Fabric
- OOP
- pandas
- Parquet
- PostgreSQL
- Power BI
- PySpark
- Python
- RDS
- Redshift
- S3
- SQL
- SQL Server
- Tableau
- XML