Data Engineer

Summary

Build and own the data pipelines, models, and infrastructure that power Virtualitics' AI-native readiness platform, turning high-volume government and defense data into analysis-ready inputs using Python, SQL, Spark, and AWS. Washington, D.C.-based role with remote flexibility; an active SECRET clearance is required.

Virtualitics is the category leader in AI-native readiness applications for defense, government, and critical infrastructure. Founded on a decade of Caltech research in partnership with NASA/JPL, we are led by scientists, strategists, and servicemembers united by a single mission: to solve the world’s most complex, mission-critical challenges with AI.

Our Readiness AI solutions deliver operational certainty — giving leaders and operators a clear picture of what’s ready, what’s at risk, and what to do next. By identifying risks early, diagnosing root causes, and recommending prioritized actions with transparent, explainable AI, we help organizations move from data complexity to decision advantage.

Behind that impact is relentless innovation. Inventors at heart, we hold 15+ U.S. patents and are leading the shift toward agent-driven readiness. But what truly sets us apart is our culture — relentless about results, grounded in transparency, and driven by compassion for the mission and the people it serves.

If you’re motivated by impact, inspired by technical depth, and ready to build AI that performs where it matters most — you’ll find your mission here.

Our team is excited to find our next Data Engineer to join the company.

As a Data Engineer at Virtualitics, you will build and own the data foundation that our AI-native readiness applications run on — the pipelines, models, and infrastructure that turn messy, high-volume government and defense data into reliable, analysis-ready inputs for our platform. You will work shoulder-to-shoulder with Machine Learning engineers, Data Scientists, and Platform teams, adapting ingestion and transformation workflows to the specific, often unique data environments of each customer. This is a chance to work close to the mission and see your pipelines directly shape how analysts and decision-makers solve hard problems.

This is a Washington, D.C.–based role with on-site work at customer sites and secure facilities as required, plus remote flexibility.

*Candidates must possess an active US Government Security Clearance (SECRET or higher).



Responsibilities

  • Design, build, and maintain scalable, secure data pipelines for ingestion, transformation, and delivery across cloud and hybrid environments.

  • Model and manage structured and unstructured data to support AI/ML workloads, analytics, and application features.

  • Partner with Full Stack Engineers, Machine Learning Engineers, and Data Scientists to productionize data workflows, feature pipelines, and model inputs.

  • Build and improve automation for data quality, validation, lineage, and monitoring.

  • Optimize storage, query performance, and cost across relational and object stores.

  • Implement and enforce data security, access controls, and governance in line with government and DoD compliance requirements.

  • Support ingestion and integration of data within DoD and commercial customer environments, including deployments into secure and disconnected settings.

  • Investigate and resolve data, pipeline, and performance issues.

  • Collaborate with Platform Engineers, AI Engineers, DevSecOps, and QA to ensure reliable, scalable, and secure data architectures.

Skills & Qualifications

  • Active US Government Security Clearance (SECRET or higher) — required.

  • BS in Computer Science, Engineering, or related field.

  • 4+ years of experience in data engineering, or building production data pipelines.

  • Strong programming skills in Python and SQL.

  • Experience with distributed data processing frameworks (e.g., Spark).

  • Strong experience with relational databases (PostgreSQL, MySQL) and data modeling.

  • Advanced AWS experience (S3, RDS, EMR, Lambda, IAM, CloudWatch); comfortable working within secure cloud environments.

  • Experience building for data quality, validation, and observability.

  • Familiarity with containerized workloads (Docker, Kubernetes) and Infrastructure as Code (Terraform).

Nice to Have

  • Experience deploying into DoD environments (Platform One, Palantir FedStart, etc.).

  • Knowledge of NIST 800-53, NIST 800-171, and CMMC 2.0 controls.

  • Experience supporting Machine Learning or Data Science teams with feature pipelines and ML data infrastructure.

  • Experience with data pipeline and orchestration tools (e.g., Airflow, dbt, Dagster, or similar).

  • Experience with streaming data (Kafka, etc.) and real-time processing.

  • Experience with data warehousing and lakehouse architectures (i.e. AWS Athena, Redshift, Databricks, Stardog, etc.).

  • Startup experience on a growing team, and a desire to mentor more junior engineers.

We Offer

At Virtualitics, you’ll join a high-performance team of engineers, scientists, strategists, and servicemembers building AI that operates in the world’s most demanding environments. The problems we solve matter — and so does the opportunity to grow while solving them.

You’ll have meaningful ownership from day one, with the ability to accelerate your career alongside a company scaling rapidly in national security and critical infrastructure.

We offer highly competitive compensation, meaningful equity participation, and fully paid medical, dental, and vision coverage for you and your dependents. Our benefits also include unlimited PTO and flexible work arrangements, with remote flexibility and hybrid options for team members based in the Los Angeles or Washington, DC areas.

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available