Quantexa Data Engineer #sn
Summary
Build and optimize big-data pipelines using Spark, Scala, Elasticsearch, and OpenShift to support analytics and compliance in a finance context.
Overview
We are seeking a talented and experienced Data Engineer with expertise in Hadoop, Scala, Spark, Elasticsearch, OpenShift Container Platform (OCP), and DevOps practices to join our team. You will design, develop, and optimize big data solutions using Apache Spark, Scala, and Elasticsearch, collaborating with cross‑functional teams to build scalable data processing pipelines and search applications. Knowledge and experience in the Compliance / AML domain is a plus. Experience with Quantexa software is required.
Responsibilities
- Implement data transformation, aggregation, and enrichment processes to support analytics and machine learning initiatives.
- Collaborate with cross‑functional teams to understand data requirements and translate them into data engineering solutions.
- Design, develop, and implement Spark Scala applications and data processing pipelines to handle large volumes of structured and unstructured data.
- Integrate Elasticsearch with Spark to enable indexing, querying, and retrieval of data.
- Optimize and tune Spark jobs for performance and scalability; ensure efficient data processing and indexing in Elasticsearch.
- Implement data transformations, aggregations, and computations using Spark RDDs, DataFrames, and Datasets, integrating them with Elasticsearch.
- Develop and maintain scalable, fault‑tolerant Spark applications, adhering to best practices and coding standards.
- Troubleshoot and resolve issues related to data processing, performance, and data quality in the Spark‑Elasticsearch integration.
- Monitor and analyze job performance metrics, identify bottlenecks, and propose optimizations in both Spark and Elasticsearch components.
- Ensure data quality and integrity throughout the data processing lifecycle.
- Design and deploy data engineering solutions on OpenShift Container Platform (OCP) using containerization and orchestration techniques.
- Optimize data engineering workflows for containerized deployment and efficient resource utilization.
- Collaborate with DevOps teams to streamline deployment processes, implement CI/CD pipelines, and ensure platform stability.
- Implement data governance practices, data lineage, and metadata management to ensure data accuracy, traceability, and compliance.
- Monitor and optimize data pipeline performance, troubleshoot issues, and implement necessary enhancements.
- Implement monitoring and logging mechanisms to ensure the health, availability, and performance of the data infrastructure.
- Document data engineering processes, workflows, and infrastructure configurations for knowledge sharing and reference.
Education
- Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related field.
- At least 2 years of relevant IT experience, preferably in a Compliance Domain of a Finance Institution.
Essential
- Must be a Quantexa certified data engineer / data architect and proficient with the software.
- Experienced developer with strong system integration/interfacing skills.
- Typically assigned to specific projects and expected to support the Application Delivery Manager (ADM) in all relevant technical matters end‑to‑end within the project.
- Depending on the project, duties may include coding, scripting, building new systems (where necessary) and interfaces, and providing environment support during SIT/UAT for new system builds.
- Ensure work is adequately documented and transferred to the production team post‑cutover.
- Collaborate with senior developers and the system architect to formulate technical solutions that satisfy security, regulatory, and architectural standards.
- Proven experience as a Data Engineer, working with Hadoop, Spark, and data processing technologies in large‑scale environments.
- Proficiency in Scala and familiarity with functional programming concepts.
- Experience with Quantexa tooling is highly preferred.
- In‑depth understanding of Apache Spark architecture, RDDs, DataFrames, and Spark SQL.
- Strong expertise in designing and developing data infrastructure using Hadoop, Spark, and related tools (HDFS, Hive, Pig, etc.).
- Experience with containerization platforms such as OpenShift and container orchestration using Kubernetes.
- Proficiency in programming languages used in data engineering, such as Spark, Python, Scala, or Java.
- Knowledge of DevOps practices, CI/CD pipelines, and infrastructure automation tools (e.g., Docker, Jenkins, Ansible, Bitbucket).
- Experience with monitoring tools such as Grafana, Prometheus, and Splunk is a plus.
- Experience with data integration and enterprise data platforms is desirable.