Data QA Engineer
Summary
Designs and executes test strategies for AWS-based cloud data platforms, validates ETL/ELT pipelines, ensures data integrity, and builds automated testing frameworks using AWS Glue, EMR, Redshift, DynamoDB, SQL, PySpark, and Python.
JOB RESPONSIBILITIES
We are seeking a Data QA Engineer to design and execute test strategies across our AWS-based cloud data platforms and other platforms. You will validate complex ETL/ELT pipelines, ensure data integrity across data lakes and warehouses, and build automated testing frameworks for large-scale datasets.
- Pipeline Validation: Design and execute end-to-end data testing strategies for batch and streaming pipelines using AWS Glue, EMR
- Data Quality Assurance: Verify data accuracy, completeness, schema drift, and transformation logic between raw sources (S3) and destination analytical stores (Redshift, DynamoDB).
- Querying & Analytics: Write complex SQL scripts and PySpark jobs to perform data reconciliation, checksum validations, and boundary testing.
- CI/CD Integration: Integrate automated data test suites into AWS CodePipeline or GitHub Actions for continuous testing in deployment pipelines.
- Defect Tracking: Identify, log, and monitor data anomalies using tools like Jira, collaborating directly with Data Engineers and Analytics teams.
- Automation: Build and maintain automated data testing suites using Python, PyTest, and frameworks like Great Expectations or AWS Deequ. (Optional, nice to have)
JOB QUALIFICATIONS
Experience 1 - 3 years in Data Quality Engineering, Data Testing, or Data Engineering
Technical
AWS Services Deep experience with S3, Amazon Redshift, Athena, and EMR
Skills
Advanced SQL (window functions, aggregations) and Python (Pandas, PySpark)
Testing Tools
Experience with Great Expectations, pytest, dbt test, or custom SQL/Python test frameworks
Databases
Hands-on experience with Relational (PostgreSQL, MySQL) and NoSQL (DynamoDB) databases