Lead / Senior Data Engineer (Azure Databricks)
Summary
Lead/Senior Data Engineer who designs, builds, and optimizes large-scale data pipelines on Azure Databricks using PySpark, Delta Lake, and SQL migrations, while setting engineering standards, running code reviews, and mentoring a distributed team. Secondary stack includes .NET, Angular, MongoDB, and AI tooling like LangChain and RAG.
About the Role
We are looking for an experienced Lead / Senior Data Engineer to join a high-performing technology team delivering innovative enterprise and AI-driven solutions within a global professional services environment.
The team designs, develops, and deploys modern technology platforms and intelligent tools that support complex business and operational processes. You will work alongside experts in software engineering, data, AI, change management, and project delivery on initiatives ranging from solution design and development to deployment and continuous improvement.
This is an exciting opportunity to work on large-scale, modern cloud and data engineering projects using Azure, Databricks, Spark, Python, and AI-powered technologies.
Technology Stack
Azure Cloud
Azure Databricks
Apache Spark / PySpark
Microservices Architecture
.NET 8
ASP.NET Core
Python
MongoDB
Azure SQL
Angular
GitHub and AI-assisted development tools
LangGraph
LangChain
RAG Pipelines
Multimodal LLMs
Your Responsibilities
Define and promote engineering best practices and coding standards across the project
Conduct thorough code reviews to ensure high code quality and adherence to established standards
Work independently while collaborating closely with cross-functional and distributed teams
Provide technical guidance and clear direction to team members
Support the coordination of day-to-day technical activities and delivery priorities
Communicate regularly with business and project stakeholders
Design, develop, and maintain robust, scalable, and high-performance Spark applications
Write clean, maintainable, and efficient code following modern software engineering principles
Optimize applications and data-processing workflows for performance and scalability
Ensure efficient and reliable data handling across large-scale data platforms
Collaborate with engineering, product, data, and business teams to deliver high-quality solutions
Identify, investigate, and resolve complex technical issues
Create and maintain comprehensive technical documentation for code, processes, architectures, and workflows
Experience & Technical Skills
5+ years of hands-on experience in software development or data engineering
Practical experience developing SQL stored procedures and migrating legacy SQL logic into Spark SQL or PySpark environments
Extensive hands-on experience with PySpark and Azure Databricks
Strong knowledge of Delta Tables, cluster management, and workflow automation
Proven experience optimizing Apache Spark performance in production environments
Strong Python programming skills, particularly for complex data processing and manipulation using libraries such as Polars or Pandas
Solid understanding of columnar data-storage formats, particularly Parquet
Practical experience working with Delta Lake and modern lakehouse architectures
Strong expertise in data processing, transformation, and analytical workflows
Strong analytical and problem-solving skills with a detail-oriented mindset
Ability to balance engineering standards and structured processes with pragmatic delivery requirements
Solid understanding of microservices architecture and scalable distributed systems
Nice to Have
Experience working with Azure Cloud services or other major cloud platforms
Experience with cloud-based services such as messaging platforms, Data Lake, object storage, caching solutions, and related managed services
Familiarity with FastAPI
Experience with containerization and orchestration technologies such as Docker and Kubernetes
Exposure to AI-driven applications, RAG architectures, LLM orchestration frameworks, or modern generative AI solutions