freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer (Canadian Citizen or PR)

Summary

Build and optimize data pipelines, vector databases, and real-time embedding systems for AI-powered search and agent memory.

Skillset Requirements

• ETL & Data Modeling: Designing pipelines for structured/unstructured data, normalization, deduplication, and semantic consistency.

• Vector Databases: pgvector, Redis, Azure AI Search, hybrid search, and index optimization.

• Distributed Data Systems: Kafka, Spark, Flink, or similar event-driven architectures.

• Data Governance: Zero-trust access, privacy controls, compliance, and auditability.

• Real-time Embedding Updates: Event-driven refresh pipelines for RAG and agent memory systems.

• Chunking & Embeddings: Semantic chunking, metadata tagging, and embedding model selection.

• Search Infrastructure: BM25, hybrid search, inverted indexes, and ranking algorithms.

• Performance Tuning: High-throughput read/write optimization.

• Data Quality & Lineage: Validation, schema enforcement, and lineage tracking (e.g., Great Expectations, OpenLineage).


See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available