Principal Data Engineer
Summary
Principal Data Engineer (VP grade, permanent, London) at HKEX/LME shaping the strategic data platform for high-volume trading data: building a Spark/Trino/Iceberg lakehouse, Kafka/Flink streaming, and self-service data capabilities on Kubernetes/OpenShift with Terraform and GitOps.
Shift Pattern:
Standard 40 Hour Week (United Kingdom)Scheduled Weekly Hours:
40Corporate Grade:
C - Vice PresidentReporting Line:
(UK Division) Information TechnologyLocation:
UK-LondonWorker Type:
PermanentThe purpose of this role is to help shape the LME’s strategic data platform which can reliably process, store, and serve high-volume, high-velocity trading data at scale, while providing engineering teams with reusable, self-service capabilities that accelerate the delivery of trading and analytical solutions.
Success in this role means delivering a highly scalable and resilient data platform that becomes the engineering foundation for all data-intensive trading workloads, enabling faster innovation, improved reliability, reduced operational complexity, and better utilisation of market and trading data across the organisation.
Strong Required Demonstrable Experience of:
A candidate for this position must have got working experience in IT organisation as a technologist who has evolved from a strong background in hands on software development:
10+ years of experience with software development, preferably working with JVM based languages (Java/Scala/Kotlin).
Provide expertise on the development of highly scalable data platforms, data warehouses (SQL server, Postgres), and modern data lakes.
Solve complex scalability, reliability, latency, and performance challenges associated with large-scale distributed systems.
Provide self-service platform capabilities that allow engineering teams to onboard, process, discover, and consume data efficiently.
Influence technology strategy and roadmaps across Trading Technology, Data Engineering, and Platform Engineering functions.
Build high-volume distributed data platforms capable of processing billions of events per day and managing terabytes scale datasets.
Develop modern Data Lakehouse platforms leveraging S3-compatible object storage, Apache Iceberg, Spark, and Trino.
Build real-time streaming integration using Apache Kafka and Apache Flink for low-latency event processing, enrichment, and data distribution.
Drive performance engineering and scalability across data ingestion, storage, streaming, query, and analytics platforms.
Design and build self-service data platform capabilities including APIs, data products, metadata services, and developer enablement frameworks.
Implement cloud-native platform engineering practices using Kubernetes/OpenShift, Terraform, CI/CD, GitOps, and Infrastructure as Code.
Lead metadata, lineage, and data discovery capabilities using technologies such as DataHub, OpenMetadata, or Apache Atlas to improve platform transparency and usability.
Ensure platform security and regulatory compliance through fine-grained access controls, encryption, secrets management, auditing, and secure-by-design engineering principles.
Bonus for knowledge of:
Automation/configuration management using toolsets such as Puppet, Chef, Ansible or equivalent
Docker and Kubernetes.
Knowledge and understanding of financial markets.
Setting up CI/CD pipelines using Bamboo.
Confluent certified developer for Apache kafka.
Dollar Universe
Personal Qualities:
Ability to work under pressure with changing priorities, with a view to resolving issues innovatively, and meeting key stakeholders expectations.
A dynamic and self-motivated attitude
Accountable and proactive
Able to provide leadership and motivate team demonstrating strong interpersonal skills
Must display strong analytical skills and attention for detail.
Demonstrates pragmatic judgement, balancing risk and business value to reach decisions which are well informed and actionable.
Skills
- Analytics
- Ansible
- API
- Automation
- CI/CD
- Cloud
- Cloud Native
- Data Engineering
- Data Ingestion
- Distributed Systems
- Docker
- Flink
- GitOps
- Iceberg
- Infrastructure as Code
- Java
- JVM
- Kafka
- Kotlin
- Kubernetes
- Lakehouse
- OpenShift
- PostgreSQL
- Puppet
- Regulatory Compliance
- Scala
- Secrets Management
- Spark
- SQL
- SQL Server
- Terraform
- Trino