Mid/Senior Data Engineer

Summary

Build and maintain data pipelines, storage (including vector databases), and a metrics layer to power AI agents for customer support data, using big data tools, Python/PySpark, and cloud data warehouses.

We're transforming our customer support data from static, disconnected reports into a live, AI-powered system. We're looking for an experienced Data Engineer to build the pipelines, storage, and logic that power our new AI agents — turning messy customer chats and tickets into clean, structured information our AI can reliably use.


Job Responsibilities

  • Ingest and clean large volumes of unstructured support data from Bliss, Salesforce, Sprinklr, and JIRA
  • Design storage and retrieval systems (including vector databases) that give our AI accurate, relevant context
  • Build a centralized, reliable metrics layer for AI-driven analytics
  • Build robust, monitored, failure-resistant pipelines — no fire drills
  • Partner with ops, product, and engineering; push back on poor data practices at the source


Required Skills

  • 5+ years in data engineering with big data tools (Hadoop, Hudi, Spark, Presto, Pinot, Flink, Kafka) and cloud data warehouses
  • Strong hands-on Python and PySpark
  • Proven use of AI tools (e.g., Claude, Codex) to accelerate development — automating checks, parsing messy text, building data logic layers
  • Direct experience with vector databases (Pinecone, Milvus, Weaviate, pgvector) and AI-feeding data pipelines
  • Bonus: experience building metric or semantic layers


If this opportunity aligns with your experience and interests, I'd be happy to discuss it further. Please send your updated resume to sjain@pksi.com.

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available