Senior Data Engineer
Summary
Senior Data Engineer to design, build, and optimize enterprise scale data platforms using Databricks, Python, and PySpark. This role focuses on developing high-performance batch and real-time data pipelines, implementing modern data engineering frameworks, and delivering reliable, governed data solutions.
Wissen Technology is hiring for Senior Data Engineer
About Wissen Technology:
At Wissen Technology, we deliver niche, custom-built products that solve complex business challenges across industries worldwide. Founded in 2015, our core philosophy is built around a strong product engineering mindset—ensuring every solution is architected and delivered right the first time. Today, Wissen Technology has a global footprint with 2000+ employees across offices in the US, UK, UAE, India, and Australia. Our commitment to excellence translates into delivering 2X impact compared to traditional service providers. How do we achieve this? Through a combination of deep domain knowledge, cutting-edge technology expertise, and a relentless focus on quality. We don’t just meet expectations—we exceed them by ensuring faster time-to-market, reduced rework, and greater alignment with client objectives. We have a proven track record of building mission-critical systems across industries, including financial services, healthcare, retail, manufacturing, and more. Wissen stands apart through its unique delivery models. Our outcome-based projects ensure predictable costs and timelines, while our agile pods provide clients with the flexibility to adapt to their evolving business needs. Wissen leverages its thought leadership and technology prowess to drive superior business outcomes. Our success is powered by top-tier talent. Our mission is clear: to be the partner of choice for building world-class custom products that deliver exceptional impact—the first time, every time.
Job Summary:
We are hiring a Senior Data Engineer to design, build, and optimize enterprise scale data platforms using Databricks, Python, and PySpark. This role focuses on developing high-performance batch and real-time data pipelines, implementing modern data engineering frameworks, and delivering reliable, governed data solutions that power analytics, reporting, and AI-driven applications across financial services ecosystems.
Must Have Skills:
- Python (7+ years) for developing scalable, modular, and production-ready data engineering applications
- PySpark & Apache Spark (5+ years) including DataFrame API, Spark SQL, Structured
- Streaming, performance optimization, partitioning strategies, joins, caching, and handling data skew
- Databricks (4+ years) including Delta Lake, Unity Catalog, Databricks Workflows, notebooks, jobs, and end-to-end data engineering capabilities
- ETL/Data Engineering/Data Pipeline Development (7+ years) building large-scale batch and real-time data processing solutions
- Advanced SQL & Data Modeling including dimensional modeling, slowly changing dimensions (SCD), schema evolution, and query optimization
- Apache Airflow (3+ years) for workflow orchestration, DAG development, dependency management, scheduling, monitoring, retries, and backfills
- dbt (Data Build Tool) including layered architecture, incremental models, snapshots, macros (Jinja), testing, documentation, and deployment best practices
- Cloud Platforms (AWS/Azure/GCP) with hands-on experience building and operating cloud native data solutions.
Good to Have:
- Microsoft Fabric for data integration, engineering, and analytics workloads
- AI-assisted development tools such as GitHub Copilot, Claude Code, or equivalent code generation platforms
- Delta Lake advanced optimization techniques and Lakehouse architecture expertise
- CI/CD pipeline implementation using Jenkins, GitHub Actions, Azure DevOps, or GitLab CI
- Containerization and orchestration using Docker and Kubernetes
- Infrastructure as Code (Terraform, CloudFormation, ARM Templates, or equivalent)
- Financial Services, Banking, Capital Markets, or AML domain expertise
- Data Quality, Data Observability, and Data Governance frameworks
Professional Attributes & Qualifications: - Education: BE/BTech/ME/MTech in Computer Science, Information Technology, Engineering, or a related discipline
- Leadership: Experience mentoring junior engineers, conducting code reviews, and contributing to technical design discussions
- Ownership: Strong track record of driving end-to-end delivery of data engineering solutions from architecture through production deployment
- Problem Solving: Excellent analytical and troubleshooting skills for distributed systems, large-scale data processing, and performance optimization
- Quality Focus: Commitment to engineering excellence, coding standards, testing automation, reliability, security, and maintainability
- Communication: Strong verbal and written communication skills with the ability to collaborate effectively with business, analytics, and engineering stakeholders
- Collaboration: Proven ability to work in cross-functional, globally distributed teams while influencing technical decisions
Key Responsibilities:
- Design, develop, and optimize large-scale batch and streaming data pipelines using Python, PySpark, Databricks, and cloud-native technologies
- Build and maintain scalable data processing workloads in Databricks with a strong focus on performance, reliability, cost optimization, and maintainability
- Develop robust dbt models using layered architecture, incremental processing, snapshots, macros, testing frameworks, and documentation standards
- Design and maintain Apache Airflow DAGs for workflow orchestration, operational monitoring, dependency management, retries, and observability
- Implement data governance, lineage, access controls, data quality validation, monitoring, and privacy standards across enterprise data platforms
- Optimize Databricks and Spark workloads through partitioning strategies, query tuning, caching, file optimization, join optimization, and efficient compute utilization
- Collaborate closely with analytics teams, product owners, architects, and business stakeholders to transform requirements into high-quality, trusted datasets
- Conduct code reviews, mentor junior engineers, and drive adoption of engineering best practices and architectural standards
- Build and enhance CI/CD pipelines to automate testing, deployment, monitoring, and operational readiness of data platforms.
Wissen Sites:
- Website:
- LinkedIn:
- Wissen Leadership: Leadership Team | Wissen
- Wissen Live:
- Wissen Thought Leadership: