freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Analyst – Knowledge Graph Initiative

Summary

Senior Data Analyst validates large datasets across Databricks, PostgreSQL, and graph databases to ensure accuracy and consistency for compliance and risk use cases.

Outstaff Hiring: Remote – Contract (Full-time, long-term)

About the Role
We’re looking for a Senior Data Analyst to support the enterprise Knowledge Graph initiative.
The project connects large-scale company, individual, and relationship data to support compliance, KYC, credit risk, sanctions screening, beneficial ownership analysis, and corporate structure research.
In this role, you’ll work directly with data engineers, product managers, and analysts to validate data, verify query results, investigate discrepancies, and assess data quality across multiple database technologies.
This is a hands-on role focused on delivering clear evidence, including validated results, documented defects, root-cause analysis, and statistical assessments.

Location: Remote (EU)
Engagement: Full-time, long-term contract
Start Date: ASAP
Language: English
Time Zone: Ability to overlap with Eastern US hours

What You’ll Do
Validate large datasets loaded into Databricks, PostgreSQL, and graph databases
Confirm data completeness, accuracy, consistency, and structural integrity
Compare query results across different database technologies
Ensure the same queries produce correct and consistent results across platforms
Identify and investigate data discrepancies and quality issues
Determine whether issues come from ETL pipelines, schema mapping, source data, or database-specific behaviour
Develop validation test cases and define expected query results
Maintain data-quality reports, defect logs, and resolution tracking
Support the assessment of compliance and business use cases
Help determine whether each use case requires a graph database or can be handled using a traditional relational database
Apply statistical methods to assess datasets and benchmark results
Analyse distributions, variance, outliers, sampling quality, and measurement reliability
Document findings clearly for engineering and product teams
Work independently with minimal supervision as part of a cross-functional engineering team

Must-Have Requirements
Strong experience in data analysis, data validation, or data quality roles
Experience validating large and complex datasets
Strong SQL skills
Experience working with PostgreSQL or another relational database
Ability to identify discrepancies and perform root-cause analysis
Experience with ETL pipelines and data transformation processes
Understanding of common data-quality issues during data loading and migration
Experience working with large-scale datasets where manual validation is not sufficient
Practical knowledge of statistical analysis, including:Distribution analysis
Outlier detection
Variance analysis
Sampling validation

Ability to analyse benchmark results and separate real differences from normal performance variation
Experience working across multiple databases or query technologies
Experience with Databricks, Spark, Delta Lake, or similar distributed data platforms
Strong written communication and documentation skills
Ability to work independently in a remote environment
Experience working within a cross-functional engineering team

Nice-to-Have
Hands-on experience with graph databases such as:Neo4j
TigerGraph
NebulaGraph
ArangoDB
Apache AGE

Understanding of graph data models, nodes, edges, and relationship structures
Experience with Cypher or other graph query languages
Knowledge of Knowledge Graphs or ontology concepts
Experience with MongoDB
Experience with data visualisation or graph analysis tools
Experience in financial services, compliance, KYC, AML, or risk
Understanding of beneficial ownership, sanctions screening, PEP data, or corporate ownership structures
Experience with large company and entity relationship datasets
Familiarity with Bureau van Dijk, Orbis, or similar data sources

Technical Stack
Databricks
PostgreSQL
SQL
Spark / Delta Lake
Neo4j
TigerGraph
MongoDB
Graph databases
Statistical analysis
ETL pipelines
Data validation and quality testing

Soft Skills
Strong analytical and problem-solving skills
Excellent attention to detail
Clear and structured communication
Ability to explain data issues to technical and non-technical stakeholders
Independent and proactive working style
Strong ownership and accountability
Comfortable working with changing technologies and requirements
Team-oriented approach

Key Deliverables
Documented validation test cases and expected results
Dataset validation reports for each technology
Cross-platform query comparison results
Root-cause analysis for identified discrepancies
Data-quality defect log and resolution tracking
Statistical analysis of benchmark results
Analytical support for graph database use-case assessment

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available