Data Engineering & AI Intern
Posted Updated
As a Data Engineering & Artificial Intelligence Intern, you will work closely with data scientists, data engineers, software engineers, and business stakeholders to design, build, and deploy data-driven and AI-powered solutions that address real-world challenges. You will gain hands-on experience in developing data pipelines, managing data platforms, and building machine learning solutions that support operational excellence and business innovation.
Roles and Responsibilities:
- Perform data exploration, cleaning, feature engineering, and model evaluation using structured and unstructured datasets.
- Develop and maintain data ingestion, ETL/ELT pipelines, and workflows to support AI model development and business analytics.
- Assist in designing, building, and optimizing scalable data pipelines and data models for analytics and machine learning applications.
- Support the development of machine learning and AI models for use cases related to transportation industry.
- Assist in building AI pipelines, dashboards, and data products to deliver actionable insights for operations and management.
- Collaborate with cross-functional teams to understand business problems and translate them into AI solutions.
- Participate in proof-of-concepts (POCs) and pilots, and support deployment into production environments where applicable.
- Research new techniques, tools, and algorithms to enhance the team's capabilities and contribute to innovation.
- Assist in organizing and managing data repositories, documenting data sources, methodologies, and findings, and ensuring data security and privacy.
Qualifications/Experience:
- Bachelor's degree or Postgraduate degree in Artificial Intelligence, Data Engineering, Data Science, Computer Science, Statistics, or related fields.
- Strong foundation in Python and familiarity with data science libraries (e.g. pandas, numpy, scikit-learn, PyTorch / TensorFlow)
- Understanding of machine learning concepts (supervised/unsupervised learning, model validation, evaluation metrics)
- Familiarity with Large Language Models (LLMs), Generative AI concepts, and prompt engineering.
- Good knowledge of SQL and relational databases, with experience querying and manipulating large datasets.
- Familiarity with data engineering concepts such as ETL/ELT processes, data pipelines, data warehousing, and data modelling.
- Exposure to cloud platforms (e.g. Microsoft Azure, AWS, or Google Cloud Platform) and data engineering services is an advantage.
- Familiarity with data orchestration or workflow tools (e.g. Apache Airflow, Azure Data Factory) is an advantage.
- Basic understanding of version control (Git) and software development best practices.
- Strong analytical and problem-solving skills with attention to detail.
- Ability to communicate effectively and collaborate within multidisciplinary teams.
- Self-motivated, eager to learn, and passionate about AI, data engineering, and digital transformation.
Skills
- AI
- Airflow
- Analytics
- AWS
- Azure
- Azure Data Factory
- Cloud
- Data Engineering
- Data Ingestion
- Data Modeling
- Data Pipelines
- Data Science
- Data Warehousing
- ELT
- ETL
- Feature Engineering
- GCP
- Generative AI
- Git
- LLM
- Machine Learning
- Model Evaluation
- NumPy
- pandas
- Prompt Engineering
- Python
- PyTorch
- scikit-learn
- SQL
- Statistics
- TensorFlow
- Version Control