Quantexa Data Engineer
Summary
Design and build large-scale data pipelines using Quantexa, Apache Spark, Scala, and Elasticsearch to power compliance and financial crime solutions for banking clients.
- Design, develop, and maintain scalable data engineering solutions using Apache Spark, Scala, Hadoop, and Quantexa.
- Build high-performance data pipelines to process structured and unstructured datasets.
- Develop Spark applications utilizing RDDs, DataFrames, Datasets, and Spark SQL.
- Design and implement data transformation, enrichment, aggregation, and validation processes.
- Integrate Elasticsearch with Spark applications for indexing, search, and analytics.
- Optimize Spark jobs and Elasticsearch queries for maximum performance and scalability.
- Develop and maintain robust, fault-tolerant distributed data processing applications.
- Deploy and manage applications on OpenShift Container Platform (OCP) using Kubernetes.
- Collaborate with DevOps teams to implement CI/CD pipelines and automate deployments.
- Support SIT, UAT, deployment activities, and production cutover.
- Implement monitoring, logging, and performance optimization using industry-standard tools.
- Ensure data quality, governance, lineage, metadata management, and regulatory compliance.
- Troubleshoot complex data processing and system integration issues.
- Prepare technical documentation and support knowledge transfer activities.
- Work closely with Solution Architects, Technical Leads, and Application Delivery Managers to deliver enterprise-grade solutions.
- Bachelor's or Master's Degree in Computer Science, Information Technology, Software Engineering, or related discipline.
- Minimum 5+ years of Data Engineering experience.
- Experience working in Banking, Financial Services, Compliance, AML, or Financial Crime projects is highly preferred.
- Strong experience working in Agile development environments.
- Quantexa Certified Data Engineer or Data Architect
- Hands-on experience with Quantexa Platform
- Apache Spark
- Scala
- Hadoop Ecosystem (HDFS, Hive, Pig)
- Elasticsearch
- OpenShift Container Platform (OCP)
- Kubernetes
- DevOps & CI/CD
- Docker
- Jenkins
- Git / BitBucket
- Data Integration
- Distributed Data Processing
- Apache Spark
- Hadoop
- HDFS
- Hive
- Pig
- Spark SQL
- RDD
- DataFrames
- Datasets
- Scala
- Python
- Java
- Elasticsearch
- Data Indexing
- Query Optimization
- OpenShift (OCP)
- Kubernetes
- Docker
- Jenkins
- Ansible
- BitBucket
- CI/CD Pipelines
- Grafana
- Prometheus
- Splunk
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
- Banking & Financial Services
- Compliance & AML Solutions
- Financial Crime Analytics
- Quantexa Implementation Projects
- Large-scale Enterprise Data Platforms
- Distributed Computing
- Performance Tuning
- Data Governance & Metadata Management
- Excellent analytical and problem-solving skills
- Strong debugging and troubleshooting capabilities
- Experience with enterprise system integration
- Strong understanding of software architecture and design principles
- Excellent communication and stakeholder management skills
- Ability to work independently and within cross-functional teams
- Strong documentation and technical writing skills