Spark Data Engineer
Summary
On-site data engineer in Tysons Corner, VA who builds and manages data pipelines, processing data with Apache Spark and Python on AWS, working with relational databases (PostgreSQL/Oracle/MySQL), Linux shell scripting, and varied file formats, with unit testing and pipeline orchestration tooling.
BT-454 – Spark Data Engineer
Location: Tysons, VA
Required skills:
- Demonstrates experience building and managing data pipelines
- Demonstrates experience with Python
- Demonstrates experience with cloud computing, using AWS services
- Demonstrates experience processing data using Apache Spark
- Demonstrates experience with an RDBMS (PostgreSQL, Oracle, MySQL) and writing SQL queries
- Demonstrates experience with Linux and shell scripting
- Demonstrates experience analyzing data in different file formats like CSV, XML, JSON, Avro, Parquet, etc.
- Demonstrates experience writing and validating unit tests
Highly desired:
- Experience with NiFi, Apache AirFlow, or an equivalent solution or tool for orchestrating data pipelines
- Experience with Java or Scala
- Experience administering an EMR/Spark cluster
- Experience conducting performance tuning of a Spark job
- Experience developing cloud-based security solutions
- Experience following a configuration management process to review and deploy code as part of releases
Skills
As published by greenhouse · 9 questions
Basics
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location
Short answers (3)
- Preferred First Name optional
- LinkedIn Profile optional
- Website optional
Pick from a list (6)
- Are you a US citizen?
- Will you now or in the future require employment visa sponsorship?
- Do you have an active Poly security clearance?
- What is your Poly Type?
- Do you live within Virginia, Maryland, or DC area?
- There is no option for remote work and all support is fully on-site. Which location do you prefer?