Data Engineer – Analytic Platform & Data Pipelines
Role Overview
3GIMBALS is seeking a Data Engineer to design, build, and maintain the data pipelines and infrastructure that power our unclassified PAI/CAI-based analytic platform. This role is responsible for ingesting, transforming, and curating large volumes of structured and unstructured data from diverse open and commercial sources; building resilient, automated ETL/ELT workflows; and ensuring data is high-quality, well-governed, and analysis-ready for the downstream analytics, knowledge graph, and modeling teams. The ideal candidate is comfortable working with messy, multi-source data at scale within secure development environments.
Key Responsibilities
Data Pipeline Development & Ingestion
- Design and build scalable batch and streaming pipelines to ingest structured and unstructured data from PAI/CAI sources, APIs, and third-party feeds
- Develop ETL/ELT workflows to normalize, enrich, and transform heterogeneous data into standardized schemas
- Build and maintain automated ingestion connectors for web, document, geospatial, and tabular data sources
- Manage data orchestration and scheduling using tools such as Airflow, Dagster, or Prefect
Data Modeling & Storage
- Design and maintain data models, schemas, and storage layers across relational, NoSQL, and object stores
- Build and maintain data lakes/lakehouses and curated, analysis-ready data marts
- Optimize partitioning, indexing, and query performance for large datasets
- Support entity resolution and data linking in coordination with the knowledge graph and modeling teams
Data Quality, Governance & Lineage
- Implement data validation, quality checks, and monitoring across pipelines
- Establish data lineage, cataloging, and metadata management
- Enforce data governance, provenance tracking, and source attribution appropriate for PAI/CAI data
- Document datasets, schemas, and pipeline logic for downstream consumers
Security & Compliance
- Ensure pipelines and data stores meet security requirements for operation in sensitive environments
- Implement encryption, access control, and secure data-handling practices
- Support Authority to Operate (ATO) processes and compliance frameworks
Required Qualifications
Technical Expertise
- 4+ years of data engineering experience building and operating production data pipelines
- Strong programming skills in Python and SQL (Scala or Java a plus)
- Experience with distributed data processing frameworks (Spark, Dask, or similar)
- Hands-on experience with workflow orchestration tools (Airflow, Dagster, Prefect)
- Proficiency with relational and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, etc.)
- Experience with cloud data platforms and services (AWS, Azure, or GCP)
Data & Infrastructure
- Experience designing data models, warehouses, and lakehouse architectures
- Familiarity with data formats and serialization (Parquet, Avro, JSON, GeoJSON)
- Understanding of data quality, lineage, and governance practices
- Experience with containerization (Docker) and CI/CD for data workflows
Domain Knowledge
- Experience working with large-scale, heterogeneous, or open-source datasets
- Understanding of data provenance and source-attribution requirements
Preferred Qualifications
- Active security clearance or ability to obtain one
- Experience in government, defense, or intelligence contracting environments
- Familiarity with PAI/CAI (publicly and commercially available information) data sources
- Experience with geospatial data processing (PostGIS, GDAL, or similar)
- Knowledge of graph data structures and preparing data for knowledge graphs
- Experience with streaming platforms (Kafka, Kinesis)
- Familiarity with federal compliance frameworks (FedRAMP, FISMA, NIST 800-53)
Technical Environment
- Languages: Python, SQL (Scala/Java a plus)
- Processing: Spark, Airflow/Dagster/Prefect, streaming frameworks
- Storage: PostgreSQL, Elasticsearch, object storage / data lake, Parquet
- Infrastructure: Docker, Kubernetes, cloud platforms (AWS GovCloud, Azure Government)
- Security: Encryption at rest and in transit, RBAC, secure data handling
This role is central to the platform: the data engineering team delivers the clean, trustworthy, well-documented data that every analytic, knowledge graph, and risk-modeling capability depends on.