Data Architect / Engineer
Summary
Senior data architect/engineer owning the data foundations of a Graph-RAG, AI-powered platform: defining schemas, taxonomies, and graph models, and building Python ETL/ingestion pipelines over enterprise content (SharePoint/Microsoft ecosystem), using Neo4j, PostgreSQL/Supabase, and orchestration like Airflow.
About the company
We are a young and fast-growing recruiting company with five years of experience working across Latin America and the United States. We partner closely with teams and founders to help them build strong, high-impact teams through recruitment, outsourcing, and team-building services. Our culture is built on effective communication, trust, and transparency. We believe great work happens when people feel heard, supported, and empowered to grow. Today, our team is made up of more than 80 professionals working across different projects throughout the region, collaborating remotely and learning from each other every day.
About the role
We are looking for a Senior Data Architect / Engineer to join an innovative team building and evolving the data foundations behind a Graph-RAG and AI-powered platform. In this role, you will own both the data structures and pipelines that support the platform, defining how knowledge is modeled, represented, ingested, and continuously updated. You will design schemas, taxonomies, and graph models while building resilient ingestion and ETL pipelines that process enterprise content from sources such as SharePoint and the Microsoft ecosystem. The ideal candidate combines strong data engineering and architecture expertise with experience in graph databases, knowledge representation, large-scale data pipelines, and the data foundations required to support high-quality RAG and LLM retrieval.
Responsibilities
- Define schemas, taxonomies, and graph data models for different knowledge domains
- Design structures for documents, versions, machine-readable representations, and the evolution of facts over time
- Build and maintain ETL and data ingestion pipelines using Python, including incremental and change-based ingestion at scale
- Integrate enterprise content sources such as SharePoint and Microsoft ecosystem platforms, processing documents, presentations, transcripts, and other multimodal content
- Keep the knowledge graph synchronized as source documents evolve and new information becomes available
- Work with graph and relational databases, including Neo4j, PostgreSQL, and Supabase, and contribute to architectural decisions regarding their long-term use
- Define and implement data quality, lineage, governance, security, and privacy standards across data pipelines
- Validate ingested data and knowledge representations in collaboration with QA and engineering teams
- Build monitoring, alerting, retry mechanisms, and error handling for ingestion and processing jobs
- Optimize pipelines for cost, throughput, scalability, and reliability
- Document data models, schemas, architectures, and pipeline processes
- Collaborate with Data Engineering, AI, QA, Product, and other technical teams to ensure data structures support RAG and LLM retrieval quality
Requirements
- 5+ years of professional experience in Data Engineering, Data Architecture, or related roles
- Strong experience with data modeling across relational and graph paradigms, including hands-on experience with Neo4j and PostgreSQL/Supabase
- Strong Python skills and proven experience building ETL/data pipelines, including large-scale, incremental, and change-data ingestion
- Experience working with enterprise content ecosystems, particularly SharePoint and preferably Power Platform or Power Automate
- Solid understanding of knowledge representation, taxonomies, ontologies, and versioned/temporal data
- Understanding of how data models and retrieval structures impact RAG and LLM quality
- Strong SQL skills and experience with data orchestration tools such as Airflow or equivalent
- Experience working with cloud-based data services and modern data architectures
- Familiarity with data governance, lineage, security, and privacy practices
- Strong analytical and problem-solving skills, with the ability to make architectural and technical decisions
- Excellent communication and documentation skills, with the ability to collaborate across multidisciplinary teams
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
Nice to have
- Experience designing data foundations for RAG, Graph-RAG, AI agents, or LLM-powered applications
- Experience with Microsoft Power Platform, Power Automate, or additional Microsoft data services
- Experience with multimodal enterprise content ingestion and document processing
- Experience with temporal databases, event-driven architectures, or knowledge graphs
- Previous experience working in technology consulting or client-facing environments
- Experience in Life Sciences, Healthcare, or Pharmaceutical industries
- Familiarity with cloud platforms such as Azure or AWS
- Experience optimizing data platforms for high-volume workloads and production environments
Benefits
- People First culture
- Referral Program
- Free access to streaming platforms
- Free access to Spotify Premium
- GYM discount
- Travel discount
- E-Learning discount
- Birthday-day gift
- Points Program
