Fantom Corporation is a mission-focused organization supporting critical programs across the defense and intelligence community. We partner with our customers to deliver high-impact technical solutions while fostering a culture built on trust, expertise, and long-term career growth.
We are seeking a Data & Software Engineer to join a small, agile team developing advanced data solutions for a custom application supporting mission-critical operations. This role is focused on designing, building, and optimizing scalable data pipelines and ETL workflows while leveraging modern cloud technologies, distributed data processing frameworks, and AI/ML capabilities.
The ideal candidate has strong Python programming skills, experience building production-grade data pipelines, and expertise with Apache Spark, AWS cloud services, relational databases, and modern data engineering best practices.
Responsibilities
Design, develop, and maintain scalable, end-to-end data pipelines and ETL/ELT workflows supporting enterprise applications
Build and optimize distributed data processing solutions using Apache Spark and PySpark
Develop robust Python applications and automation scripts to support data engineering initiatives
Deploy and manage data pipelines using workflow orchestration platforms such as AWS Step Functions or Apache Airflow
Containerize and deploy data applications within AWS cloud environments using Docker, Podman, or similar technologies
Design, optimize, and maintain PostgreSQL and MySQL databases, including schema design, indexing, and query optimization for analytical workloads
Collaborate with stakeholders to gather requirements, assess technical feasibility, and design scalable data solutions
Troubleshoot and resolve data quality issues, pipeline failures, and performance bottlenecks
Implement and support data governance, security, privacy, and compliance best practices
Develop and maintain technical documentation, architecture diagrams, and engineering standards
Support large-scale data migration and platform modernization initiatives
Integrate AI/ML services, machine learning models, and advanced analytics capabilities into enterprise data platforms
Utilize Git and CI/CD pipelines to support automated testing, deployment, and version control
Required Qualifications
Must be fully cleared with a recent polygraph
Must be willing and able to work fully onsite at the location listed in this posting
5+ years of experience as a Data Engineer, Software Engineer, or similar technical role
Demonstrated experience building production-scale data pipelines and ETL/ELT workflows
Advanced programming experience with Python, including libraries such as Pandas and NumPy
Strong experience with Apache Spark and PySpark for distributed data processing
Experience with workflow orchestration tools such as Apache Airflow or AWS Step Functions
Experience deploying containerized applications using Docker, Podman, or similar technologies
Experience developing cloud-native applications within AWS environments
Hands-on experience with AWS services including Amazon S3, AWS Lambda, and AWS Step Functions
Strong experience with PostgreSQL and MySQL, including schema design, performance tuning, and query optimization
Advanced SQL skills supporting large-scale analytical workloads
Experience with Git and CI/CD practices for data engineering and software deployment
Understanding of data governance, privacy, security, and compliance principles
Strong analytical, troubleshooting, and problem-solving skills
Experience collaborating directly with stakeholders to gather requirements and deliver technical solutions with minimal oversight
#CJ
Desired Qualification
- Experience designing and implementing Lakehouse architectures using Apache Iceberg
- Experience configuring and supporting enterprise data platform technologies, including:
- Apache Ranger
- Trino
- Apache Polaris or Unity Catalog OSS
- Apache Superset
- Experience with Infrastructure as Code (Terraform or AWS CloudFormation)
- Proficiency with Bash scripting for automation and operational support
- Experience tracking data lineage using OpenLineage or similar tools
- Working knowledge of Java
- Experience implementing data quality frameworks, automated testing, and validation strategies
- Experience supporting enterprise data modernization and large-scale migration initiatives
- Experience integrating AI/ML services, including OCR, NLP, speech-to-text, language detection, translation services, topic modeling, Large Language Models (LLMs), and Retrieval-Augmented Generation (RAG) pipelines
- Experience processing geospatial data using technologies such as H3, PostGIS, or similar frameworks
- Experience with NoSQL databases such as DynamoDB
- Experience developing engineering standards, documentation, and reusable design patterns
Fantom Corp is a Software Development, Agile Cloud, Software Development, Cyber Security (Risk Management, Assessments & Authorization (A&A)), Data, AI Platform (Computer Vision Models), Podcasting Media Services, and IT Services provider. Established in 2015, Fantom Corp serves Federal customers with top-notch Cybersecurity Architects, Data Scientists/Analysts, Software Engineers/Developers, DevSecOps Engineers, Project Managers, Identity, Credential Access Management (ICAM) services , and Cloud-certified practitioners. We excel in delivering emerging technologies such as Artificial Intelligence (AI) and Machine Learning (ML) with a focus on identifying trends, object detection, and classification of structured and unstructured data. Fantom Corp possesses mastery in all aspects of digital audio production. We lead in the ideation and creation of efforts for clients who want to harness the power of podcasting. We guide them in selecting the right show format for their needs and goals. As a Small Business, we possess the innovation, speed and flexibility to meet your requirements.