Data Engineer
Summary
Bacancy Services is hiring a Data Engineer (3+ years) in Ahmedabad to design, build, and maintain scalable ETL/ELT data pipelines and data platforms. Day-to-day work centers on Python, SQL, data warehouses like Snowflake, and Airflow orchestration, supporting analytics, BI, and ML teams.
- Design, develop, and maintain end-to-end ETL/ELT data pipelines
- Build and optimize data ingestion frameworks from multiple sources (APIs, databases, files, streaming
- Work with data warehouses and lakehouse architectures (Snowflake, Redshift, BigQuery, Databricks, etc.)
- Develop and maintain batch and near-real-time pipelines
- Ensure data quality, validation, monitoring, and observability
- Optimize SQL queries and data models for performance and scalability
- Collaborate with analytics, BI, and ML teams to support reporting and advanced analytics
- Implement data governance, security, and access controls
- Automate workflows using orchestration tools (Airflow, Luigi, etc.)
- Troubleshoot and resolve data pipeline and production issues
- 3+ years of hands-on experience as a Data Engineer
- Strong proficiency in Python and SQL
- Experience with ETL/ELT pipeline development
- Hands-on experience with data warehouses (Snowflake preferred)
- Experience with Apache Airflow or similar orchestration tools
- Strong understanding of data modeling (star/snowflake schemas)
- Experience working with large-scale datasets
- Familiarity with Git-based version control
- Experience working in agile environments
- Experience with Apache Spark / PySpark
- Exposure to Kafka or streaming platforms
- Knowledge of cloud platforms (AWS / Azure / GCP)
- Experience with Docker & Kubernetes
- Familiarity with data quality frameworks (Great Expectations, Soda, etc.)
- BI tools experience (Power BI, Looker, Tableau)
- Infrastructure-as-Code tools (Terraform, CloudFormation)
- CI/CD pipelines for data workflows
Skills
- Agile
- Airflow
- Analytics
- API
- AWS
- Azure
- BigQuery
- CI/CD
- Cloud
- CloudFormation
- Data Governance
- Data Ingestion
- Data Modeling
- Data Pipelines
- Data Quality
- Databricks
- Docker
- ELT
- ETL
- GCP
- Git
- Infrastructure as Code
- Kafka
- Kubernetes
- Lakehouse
- Looker
- Machine Learning
- Observability
- Power BI
- PySpark
- Python
- Redshift
- Snowflake
- Spark
- SQL
- Tableau
- Terraform
- Version Control