Senior Data Engineer
Summary
Senior Data Engineer for the VARTA SENSE banking analytics platform in Noida: designs and operates scalable batch/incremental data pipelines, data quality controls, and ML feature pipelines for bank transaction data, with secure deployment in restricted bank environments. Core stack: Python, SQL, Spark/PySpark, Airflow, CDC, Docker and monitoring tools.
- Design and develop scalable batch, incremental and production-grade data pipelines for large volumes of transaction and customer data.
- Build standardized and reusable data models and canonical schemas for transactions, customer attributes, offers, exposure and outcome events.
- Develop configurable ingestion and mapping frameworks to integrate data received from different banks and source systems.
- Manage pipeline orchestration including dependencies, checkpointing, retries, partial failures, backfills and safe replay mechanisms.
- Implement deduplication, late-arriving data handling, incremental loads, CDC, watermarking and merge/upsert processing.
- Establish strong data-quality controls, automated testing, source-to-target reconciliation and schema validation.
- Ensure data lineage and schema evolution are properly managed across pipelines.
- Proactively identify and prevent issues such as missing data, duplicate records, incorrect aggregations, silent data loss and double counting.
- Maintain reliable and auditable data-processing standards across production environments.
- Work closely with the ML Lead to develop and maintain analytical and ML feature pipelines.
- Ensure consistency between training and production/inference datasets.
- Maintain point-in-time correctness and prevent data leakage within feature pipelines.
- Support versioned analytical features and reusable feature definitions.
- Optimize large-scale data-processing workloads through effective partitioning, query optimization, join strategies, memory management and I/O optimization.
- Benchmark workloads against expected transaction volumes and available infrastructure.
- Identify bottlenecks and continuously improve processing speed, stability and infrastructure utilization.
- Ensure data pipelines remain scalable as transaction volumes and customer deployments increase.
- Develop and support pipelines for deployment within bank-controlled on-premise and private-cloud environments.
- Ensure pipelines can operate in restricted or offline environments without public-internet dependency at runtime.
- Implement appropriate controls for PII handling, encryption, tokenization, masking, access management and secure connectivity.
- Establish monitoring and alerting for pipeline failures, data-quality issues, infrastructure bottlenecks and processing delays.
- Define and maintain backfill, replay, recovery and disaster-recovery procedures for critical pipelines.
- Work closely with ML, Architecture, Backend Engineering, DevOps and Infrastructure teams to design and deliver reliable data solutions.
- Collaborate with development partners to reproduce, transition and operationalize data pipelines within FCI.
- Support integration between data platforms and downstream application / ML services.
- Participate in technical discussions, architecture reviews and production troubleshooting.
- Build reusable adapters and configuration layers so that onboarding a new bank can be managed through configuration and mapping rather than changes to the core product code.
- Maintain clear technical documentation, runbooks and operational procedures.
- Document architecture decisions, troubleshooting steps and recovery procedures.
- Enable alternate engineering resources to independently operate and troubleshoot critical pipelines.
- Participate in code reviews, Git-based development, CI/CD practices and continuous improvement of engineering standards.
Requirements
- 5–9 years of relevant experience in Data Engineering, with hands-on ownership of production-grade data pipelines.
- Strong hands-on expertise in Python and Advanced SQL.
- Strong understanding of SQL concepts including window functions, query optimization and query-plan analysis.
- Practical experience with distributed data-processing platforms such as Apache Spark / PySpark or equivalent technologies.
- Experience building and managing production pipelines using Apache Airflow or an equivalent orchestration framework.
- Strong understanding of data modelling, ETL/ELT, batch processing and incremental data pipelines.
- Experience working with Parquet and other columnar data formats, partitioning, schema evolution and reconciliation.
- Practical experience with CDC / incremental ingestion technologies and concepts.
- Exposure to technologies such as Kafka, Debezium or equivalent event / CDC platforms would be preferred.
- Experience implementing automated Data Quality frameworks using tools such as Great Expectations, dbt tests, Soda or equivalent.
- Understanding of Data Lineage and Metadata Management using OpenLineage, Marquez or similar self-hosted solutions would be an advantage.
- Experience with lakehouse technologies such as Delta Lake, Apache Iceberg or Apache Hudi would be preferred.
- Exposure to Feature Store platforms such as Feast or equivalent would be beneficial.
- Working knowledge of Git, CI/CD, Docker and Linux environments.
- Experience with monitoring tools such as Prometheus, Grafana or equivalent platforms.
- Strong understanding of production support, troubleshooting, performance tuning and failure recovery.
- Ability to clearly explain technical design decisions, performance trade-offs and production engineering challenges.
- Strong collaboration skills with Data Science, ML, Backend Engineering, Architecture and Infrastructure teams.
Benefits
- Cashless medical insurance for employees, spouses, and children
- Accidental insurance coverage
- Life insurance coverage
- Retirement benefits including Provident Fund (PF) and Gratuity
- ESI*
- Complementary meal coupons
- Company-paid transportation
- Sodexo benefits for income tax savings
- Paternity & Maternity Leave Benefit
- National Pension Saving
- EL encashment
- Sick Leave
Skills
- Airflow
- CI/CD
- Cloud
- Data Engineering
- Data Lineage
- Data Modeling
- Data Pipelines
- Data Quality
- Data Science
- dbt
- Debezium
- Delta Lake
- DevOps
- Docker
- ELT
- ETL
- Feature Engineering
- Git
- Grafana
- Iceberg
- Kafka
- Lakehouse
- Linux
- Machine Learning
- Metadata Management
- Parquet
- Prometheus
- PySpark
- Python
- Spark
- SQL
- Test Automation