Data Architect / Engineer (US-based)
Summary
Staffy is hiring a Data Architect/Engineer (100% remote, US-based, 1099 contract) to own the data layer of a graph-RAG platform: defining schemas, taxonomies and graph models, and building Python ETL pipelines that keep a knowledge graph in sync with enterprise content like SharePoint, on a Supabase/PostgreSQL backend.
About the company
We are a young and fast-growing recruiting company with five years of experience working across Latin America and the United States. We partner closely with teams and founders to help them build strong, high-impact teams through recruitment, outsourcing, and team-building services.
Our culture is built on effective communication, trust, and transparency. We believe great work happens when people feel heard, supported, and empowered to grow. Today, our team is made up of more than 80 professionals working across different projects throughout the region, collaborating remotely and learning from each other every day.
About the role
We are looking for a Data Architect / Engineer to join a highly motivated and experienced team and take ownership of both the data structures behind a graph-RAG platform and the pipelines that bring enterprise content into it.
In this role, you will define schemas, taxonomies, and graph models for knowledge domains, build resilient ingestion and ETL pipelines from enterprise content sources, and ensure the knowledge graph remains accurate and up to date as source documents evolve.
The ideal candidate combines strong data engineering and architecture expertise with experience in graph databases, Python-based data pipelines, enterprise content ecosystems, and RAG/LLM architectures. You should also bring a solution-focused mindset, strong problem-solving skills, and the ability to contribute to architectural decisions in a fast-moving environment.
Responsibilities
- Define schemas, taxonomies, and graph models for different knowledge domains
- Design how documents, versions, and machine-readable representations are structured, including how facts and relationships evolve over time
- Build resilient ETL and data ingestion pipelines from enterprise content sources such as SharePoint and the broader Microsoft ecosystem
- Process and integrate multimodal enterprise content, including presentations, documents, transcripts, and other sources
- Keep knowledge graphs synchronized as source documents are created, modified, or removed
- Evaluate and contribute to the evolution of the Supabase/PostgreSQL backend architecture
- Establish data quality, lineage, governance, security, and privacy standards across data pipelines
- Build monitoring, alerting, validation, and error-handling mechanisms for ingestion processes
- Document data models, schemas, pipeline architecture, and technical decisions
- Optimize pipelines for cost, throughput, scalability, and reliability
- Collaborate with QA and cross-functional teams to validate the quality and accuracy of ingested content
- Contribute to architectural decisions related to knowledge representation and RAG/LLM retrieval quality
Requirements
- 5+ years of professional experience in Data Engineering, Data Architecture, or related roles
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
- Strong experience with data modeling across relational and graph paradigms, including Neo4j and PostgreSQL/Supabase
- Strong hands-on experience building ETL/data pipelines in Python, including large-scale, incremental, and change-data ingestion
- Experience working with enterprise content ecosystems, particularly SharePoint and ideally Power Platform or Power Automate
- Understanding of knowledge representation, taxonomies, ontologies, and versioned/temporal data
- Understanding of how data structures and retrieval architectures impact RAG and LLM answer quality
- Strong SQL skills and experience with data orchestration tools such as Airflow or equivalent technologies
- Familiarity with cloud-based data services and modern data architectures
- Understanding of data governance, security, privacy, lineage, and data quality practices
- Strong documentation, communication, and cross-functional collaboration skills
- Ability to work independently, make sound technical decisions, and operate effectively in ambiguous environments
Nice to have
- Experience designing and operating graph-RAG or knowledge graph platforms
- Experience with Microsoft Power Platform / Power Automate
- Experience with multimodal enterprise content processing
- Experience working with cloud data platforms and large-scale data environments
- Previous experience in technology consulting or client-facing environments
- Experience in Life Sciences, Healthcare, or other regulated industries
- Experience optimizing data pipelines for large-scale workloads and cloud costs
Benefits
- 100% remote position within the United States, USD 105,000 - 115,000 annual compensation
- Type of contract: Independent Contract (1099)
- People First culture
- Referral Program
- GYM discount
- Birthday-day gift
- Points Program
