Data Engineer (AI Agents)
Summary
Sigmatic is hiring a remote Data Engineer to build and run data pipelines — dbt transforms on a multi-tenant PostgreSQL warehouse and Python ETL on AWS — feeding an AI agent layer (AWS Bedrock, Qdrant) over healthcare, accounting, and operations data. The role emphasizes rigorous verification of AI-assisted code.
- Transformation — dbt on PostgreSQL: staging, intermediate and marts, with contracts and tests enforced in CI
- Warehouse — multi-tenant PostgreSQL on AWS RDS, schema per tenant
- Extraction and loading — Python on AWS Lambda, Glue PySpark, Step Functions, EventBridge, with SAM and Terraform
- Sources — EHR and clinical systems, accounting (QuickBooks, NetSuite), scheduling, supply chain, IoT and scanner event streams
- AI layer — AWS Bedrock, a generated semantic catalogue, Qdrant retrieval, Python agent services
- Engineering — GitHub with gitflow, mandatory review, CI gates, Linear, SOC 2 Type II
- Five or more years in data engineering or analytics engineering, having owned pipelines a business depended on
- Strong SQL and Python: production transformations, testing, debugging, performance work
- Serious dbt experience. Multi-tenant or multi-source projects are what we most want to hear about
- ETL and ELT against messy real sources: APIs, databases, files, cloud storage, with full and incremental loads, CDC and idempotent reruns
- A way of working that already assumes AI assistance, and a clear account of how you verify what you did not write
- Precision in writing. Much of your output is read by a model before a person sees it, so unambiguous descriptions of what a column means are a deliverable, not an afterthought
- A record of finishing: edge cases, validation, documentation, deployment, production support