Database Administrator
Summary
A hybrid Data Scientist/ETL Engineer role supporting the IRS Compliance Data Warehouse: the hire designs, builds, and refactors ETL pipelines (Unix/Linux shell, SQL, Python) loading data into Sybase IQ and Postgres, and applies statistical/ML techniques (Pandas, NumPy, scikit-learn) to produce predictive models and insights.
Brillient Corporation is seeking a Data Scientist / ETL Engineer to support a mission-critical IRS data warehousing initiative. In this hybrid role, you will design, build, and maintain ETL processes that move data into the Compliance Data Warehouse (CDW). You will also apply statistical and machine learning techniques to that data to produce predictive models and actionable insights. You'll work closely with data architects, business analysts, data quality specialists, and mission stakeholders to make sure data is accurate, consistent, and analytically useful.
This role supports CDW operations within RAAS. That work includes data analysis, process improvement recommendations, and research and evaluation of emerging technologies to improve RAAS data availability, analytics, and value to stakeholders. The CDW is a non-IT data warehouse containing all of the IRS's return, entity, information return, regulatory, enforcement, web, and security data. The team provides technical guidance and executes work across ETL, database administration, SQL and Bash development, SAS administration, system security, COTS ETL development, metadata, data quality, customer service, web design, and intergovernmental data exchanges.
Key Responsibilities
ETL Design and Development
- Extract data and tables from Unix/Linux systems, and transform and load them into Sybase IQ and Postgres.
- Refactor existing ETL jobs and build new solutions as needed. This includes building an object-oriented, Unix-based framework of scripts and stored procedures to run and monitor multiple ETL and statistics processes.
- Develop and run shell, SQL, and Python scripts to support data processing, automation, and analytical workflows.
- Build repeatable, scalable data pipelines and analytical processes.
- Troubleshoot errors in shell and SQL, and configure and operate SSH clients across servers.
Data Science and Analytics
- Prepare, clean, transform, and explore data, and engineer features, using Pandas and NumPy.
- Analyze large, complex datasets and tables to find trends, patterns, relationships, and actionable insights.
- Develop, implement, evaluate, and maintain statistical and machine learning models with scikit-learn or comparable frameworks.
- Validate model results, assess performance, and find ways to improve accuracy and reliability.
- Conduct exploratory and statistical analysis to support data-driven decision-making.
- Develop and document analytical workflows and models in Jupyter Notebook.
- Work with structured and unstructured data from multiple sources and formats.
Optimization, Performance, and Data Quality
- Optimize ETL processes for efficient, timely data processing.
- Monitor ETL jobs and resolve performance bottlenecks and failures promptly.
- Keep data accurate and intact through validation, cleansing, and auditing.
- Work with the Data Quality team to resolve data inconsistencies and issues.
Collaboration and Documentation
- Work with Data Architects to design data models and schemas that meet business needs.
- Translate business and mission requirements from analysts, SMEs, and stakeholders into ETL and analytical solutions.
- Document ETL processes, workflows, data dictionaries, methodologies, models, assumptions, and results.
- Help users move data into Sybase IQ and provide metadata for the metadata repository.
- Present complex technical findings clearly to both technical and non-technical audiences.
- Take part in Scrum and client meetings, and keep Kanban board cards up to date.
Environment and Continuous Improvement
- Work in Linux/Unix and Windows environments, including containerized and distributed platforms (OpenShift, Kubernetes).
- Keep up with emerging ETL, data science, ML, and AI technologies, and recommend improvements to processes and infrastructure.