AI Data Engineer

Role Overview

We are hiring a hands-on engineer with strong capabilities in data pipelines and applied AI (GenAI) to build and support our AI-driven platform and Datalake platform.

Key Responsibilities

  • Own the architecture, operation, and continuous improvement of the enterprise Data Lake and AI Platform.
  • Design, implement, and maintain scalable data pipelines using AWS Glue, Apache Iceberg, Redshift, and related cloud-native technologies.
  • Ensure data quality, governance, lineage, observability, security, and platform reliability.
  • Establish standards and best practices for data ingestion, transformation, storage, and consumption.
  • Design, build, and deploy AI Agents, AI Advisors, and GenAI-powered business solutions.
  • Develop Retrieval-Augmented Generation (RAG)architectures leveraging enterprise knowledge and data assets.
  • Design multi-agent workflows to automate business processes and improve user productivity.
  • Evaluate emerging AI technologies and identify opportunities to enhance AI capabilities across the organization.

Core Skills (Must-Have)

1. Data Engineering Fundamentals

  • Strong hands-on experience in ETL/ELT pipeline development
  • Proficient in data transformation, cleaning, and modeling
  • Solid experience with SQL and working with large datasets
  • Familiar with Airflow, AWS Glue, S3, Redshift, Lambda
  • Understanding of data quality, lineage, and reliability concepts

2. Programming & Backend Development

  • Strong proficiency in Python (preferred) or similar backend language
  • Experience building RESTful APIs and backend services
  • Ability to write clean, maintainable, production-grade code

3. GenAI / LLM Capabilities

  • Hands-on experience working with LLMs (e.g. OpenAI, Claude, or QWEN)
  • Understanding of Retrieval-Augmented Generation (RAG) architecture
  • Experience with embeddings, vector databases, and prompt orchestration
  • Ability to connect enterprise data with LLMs in a secure and scalable way

4. Data Storage & Systems

  • Experience with relational databases (e.g. MySQL, PostgreSQL)
  • Familiarity with NoSQL / document stores
  • Understanding of data lake / warehouse concepts

5. Deployment & Platform Skills

  • Experience with Docker and containerization
  • Basic familiarity with Kubernetes / AWS / OpenShift or similar platforms
  • Understanding of CI/CD practices for backend or data applications

Good-to-Have Skills

  • Experience with streaming data (Kafka or equivalent)
  • Exposure to machine learning workflows
  • Experience with API gateways, authentication, and security practices
  • Familiarity with cloud platforms (AWS)
  • Prior experience in financial services / trading systems

Key Attributes

  • Able to operate as a hybrid engineer across data and AI domains
  • Strong problem-solving and system design thinking
  • Comfortable working in ambiguous, fast-moving environments
  • Focus on delivering working solutions, not just prototypes

Scope (High-Level)

  • Build and maintain data pipelines
  • Enable AI/GenAI use cases (e.g. AI Advisor, Research Chatbot etc.)
  • Integrate AI capabilities into applications and services