Senior Big Data Engineer
Summary
The Senior Big Data Engineer will design, build, and maintain large-scale ETL/ELT data pipelines using Python and SQL. The role involves supporting data integration, troubleshooting pipeline performance, and communicating technical requirements to stakeholders in a government-focused environment.
Our client is seeking an experienced Big Data Engineer to design, develop, and support large-scale data pipelines and transformation processes. The ideal candidate will have strong SQL, Python, ETL/ELT, and data engineering experience, along with the ability to communicate technical concepts clearly to clients, stakeholders, and cross-functional teams.
Java experience is helpful for supporting existing systems and integrations but is not the primary focus of this position.
Responsibilities
- Design, build, maintain, and optimize ETL/ELT pipelines that move and transform large-scale data across source and target systems.
- Write and optimize complex SQL queries, stored procedures, data models, and validation processes.
- Develop data transformation, integration, and automation solutions primarily using Python.
- Use Java as needed to support existing applications, services, and system integrations.
- Parse, transform, and validate XML-based data formats.
- Work within Linux and shell environments to support scripting, scheduled jobs, cron processing, log reviews, and troubleshooting.
- Implement data validation, error handling, monitoring, and traceability controls.
- Troubleshoot pipeline failures, data discrepancies, performance issues, and integration defects.
- Collaborate with client, security, infrastructure, data engineering, and delivery teams.
- Communicate technical issues, risks, and recommendations clearly to technical and nontechnical stakeholders.
- Participate in Agile planning, status reporting, documentation, issue tracking, and cross-team coordination.
Required Qualifications
- Bachelor’s degree from an accredited college or university or equivalent professional experience.
- Five or more years of professional experience in data engineering, database development, ETL/ELT development, or a related data integration role.
- Strong hands-on SQL experience, including complex queries, performance tuning, stored procedures, and data modeling.
- Experience designing, building, optimizing, and supporting large-scale data pipelines.
- Strong Python programming experience supporting data transformation, integration, automation, or pipeline operations.
- Experience parsing, transforming, and validating XML-based data.
- Experience working in Linux and shell environments, including scripting, scheduled jobs, cron, and log troubleshooting.
- Experience implementing data validation, error handling, pipeline monitoring, and production troubleshooting.
- Strong written and verbal communication skills.
- Ability to explain technical concepts clearly to clients and stakeholders.
- Ability to collaborate across security, infrastructure, engineering, and delivery teams.
- Self-motivated and proactive, with a positive and collaborative approach.
- U.S. citizenship is required.
- Must be willing and able to undergo a background investigation for a Public Trust suitability determination.
Preferred Qualifications
- Experience using Python as the primary language for data pipeline automation, transformation, monitoring, or production troubleshooting.
- Experience supporting Java-based applications, services, or integrations.
- Working knowledge of AWS data services, including S3, Glue, Lambda, RDS, or Redshift.
- Exposure to XBRL or other structured financial, regulatory, or compliance-driven data formats.
- Experience supporting federal, government, or compliance-driven data environments.
- Experience with Oracle, PostgreSQL, Redshift, or other relational database platforms.
- Familiarity with reporting, analytics, or dashboarding tools such as Power BI or MicroStrategy