AI Data Engineer
Role Overview
We are hiring a hands-on engineer with strong capabilities in data pipelines and applied AI (GenAI) to build and support our AI-driven platform and Datalake platform.
Key Responsibilities
- Own the architecture, operation, and continuous improvement of the enterprise Data Lake and AI Platform.
- Design, implement, and maintain scalable data pipelines using AWS Glue, Apache Iceberg, Redshift, and related cloud-native technologies.
- Ensure data quality, governance, lineage, observability, security, and platform reliability.
- Establish standards and best practices for data ingestion, transformation, storage, and consumption.
- Design, build, and deploy AI Agents, AI Advisors, and GenAI-powered business solutions.
- Develop Retrieval-Augmented Generation (RAG)architectures leveraging enterprise knowledge and data assets.
- Design multi-agent workflows to automate business processes and improve user productivity.
- Evaluate emerging AI technologies and identify opportunities to enhance AI capabilities across the organization.
Core Skills (Must-Have)
1. Data Engineering Fundamentals
- Strong hands-on experience in ETL/ELT pipeline development
- Proficient in data transformation, cleaning, and modeling
- Solid experience with SQL and working with large datasets
- Familiar with Airflow, AWS Glue, S3, Redshift, Lambda
- Understanding of data quality, lineage, and reliability concepts
2. Programming & Backend Development
- Strong proficiency in Python (preferred) or similar backend language
- Experience building RESTful APIs and backend services
- Ability to write clean, maintainable, production-grade code
3. GenAI / LLM Capabilities
- Hands-on experience working with LLMs (e.g. OpenAI, Claude, or QWEN)
- Understanding of Retrieval-Augmented Generation (RAG) architecture
- Experience with embeddings, vector databases, and prompt orchestration
- Ability to connect enterprise data with LLMs in a secure and scalable way
4. Data Storage & Systems
- Experience with relational databases (e.g. MySQL, PostgreSQL)
- Familiarity with NoSQL / document stores
- Understanding of data lake / warehouse concepts
5. Deployment & Platform Skills
- Experience with Docker and containerization
- Basic familiarity with Kubernetes / AWS / OpenShift or similar platforms
- Understanding of CI/CD practices for backend or data applications
Good-to-Have Skills
- Experience with streaming data (Kafka or equivalent)
- Exposure to machine learning workflows
- Experience with API gateways, authentication, and security practices
- Familiarity with cloud platforms (AWS)
- Prior experience in financial services / trading systems
Key Attributes
- Able to operate as a hybrid engineer across data and AI domains
- Strong problem-solving and system design thinking
- Comfortable working in ambiguous, fast-moving environments
- Focus on delivering working solutions, not just prototypes
Scope (High-Level)
- Build and maintain data pipelines
- Enable AI/GenAI use cases (e.g. AI Advisor, Research Chatbot etc.)
- Integrate AI capabilities into applications and services