Data Engineer
Summary
Designs and maintains scalable ETL/ELT pipelines and real-time data processing for AI/ML use cases using pgVector, Azure AI Search, Redis, and cloud platforms.
We are seeking an experienced Data Engineer with 5–7 years of experience in developing scalable ETL/ELT processes, data pipelines, and real-time data processing solutions. The ideal candidate will have hands‑on experience with vector databases and search technologies, including pgVector, Azure AI Search, and Redis, along with a strong understanding of data governance and compliance.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines.
- Build batch and real-time data processing solutions for structured and unstructured data.
- Develop data ingestion, transformation, validation, and integration workflows.
- Work with pgVector, Azure AI Search, Redis, and other vector search technologies.
- Support data pipelines for AI/ML, Generative AI, RAG, embeddings, and semantic search use cases.
- Optimize pipelines for performance, scalability, reliability, and data quality.
- Implement monitoring, validation, error handling, and recovery mechanisms.
- Work with relational and NoSQL databases.
- Implement data governance, security, privacy, compliance, and access controls.
- Maintain data lineage, metadata, documentation, and auditability.
- Collaborate with Data Scientists, Software Engineers, Architects, and business stakeholders.
- Troubleshoot pipeline failures and resolve data-quality and performance issues.
Required Skills
- 5–7 years of experience as a Data Engineer.
- Strong hands‑on experience with ETL/ELT and data pipeline development.
- Experience with real‑time/streaming data processing.
- Hands‑on experience with one or more of:
- pgVector
- Redis
- Strong Python and SQL skills.
- Experience with data modeling and database technologies.
- Experience working with structured and unstructured data.
- Knowledge of data governance, data quality, security, and compliance.
- Experience with cloud data platforms, preferably Microsoft Azure.
- Experience with APIs and data integration.
- Familiarity with Git and CI/CD practices.
Preferred Skills
- Experience with Generative AI, RAG, embeddings, and vector search.
- Experience with Azure Data services.
- Knowledge of Kafka or other streaming technologies.
- Experience with Spark/PySpark.
- Knowledge of data cataloging, lineage, and metadata management.
- Experience with data pipeline monitoring and observability.
- Experience working in a regulated industry such as banking or financial services.