Data Engineer
Summary
Data Engineer at G2G (OffGamers), a global gaming marketplace in Kuala Lumpur. You'll build and run ETL/ELT pipelines on a serverless AWS lakehouse (Glue/Spark, S3, Athena/Iceberg), maintain datamarts and Tableau dashboards, and deliver stakeholder reporting — a hybrid data engineering/analytics role open to fresh graduates through 3 years' experience.
G2G is a global gaming marketplace serving millions of buyers and sellers worldwide. Our Data team owns the company’s analytics platform end to end — a fully serverless AWS lakehouse processing billions of records across multiple brands and AWS accounts — and delivers the pipelines, datamarts, and Tableau dashboards that drive daily business decisions.We are hiring a Data Engineer whose role spans both data engineering (ETL/ELT, data modeling, pipeline operations) and data analytics (reporting, dashboards, stakeholder insights). You will work directly with senior engineers on a modern, AI-assisted engineering workflow, and see your work used by the business every day. We welcome candidates from fresh graduates through mid-level (up to 3 years of experience); scope and ownership will be matched to your level.
- Develop, maintain, and optimize ETL/ELT pipelines using AWS Glue (Spark), Glue workflows, and triggers.
- Build ingestion for structured and semi-structured data from databases (AWS DMS / CDC), APIs, and file sources into the S3 data lake.
- Develop data models and curated datamarts in Athena/Iceberg, maintaining source-to-target mappings based on business rules.
- Implement data validation and quality checks; monitor production pipelines and respond to alerts.
- Investigate and resolve pipeline failures and data quality incidents through structured root-cause analysis.
- Optimize SQL queries, Athena scan volumes, and Glue job configurations for performance and AWS cost efficiency.
- Analytics & Reporting:
- Build, extend, and maintain Tableau dashboards.
- Translate stakeholder requests into well-defined metrics, datasets, and reports.
- Develop and operate automated reporting so business teams receive accurate, timely data.
- Validate report accuracy and investigate discrepancies raised by business users.
- Platform & Practices:
- Use AI tooling (e.g., Claude, MCP integrations) to accelerate development, operations, and reporting workflows.
- Document pipelines, data mappings, processes, and incident resolutions.
- Handle data responsibly: follow PII, security, and access-control practices (IAM, KMS, scoped datamarts).
- Contribute to engineering standards, code reviews, and continuous improvement within the Data team.
- Bachelor’s degree in Computer Science, Data Engineering, Information Systems, or a related field.
- 0–3 years of relevant experience. Fresh graduates are welcome — demonstrable project ownership (academic, personal, or internship) is a must-have.
- Proficiency in SQL and working knowledge of Python (depth expected to match experience level).
- Solid understanding of ETL/ELT, data modeling, and data warehouse / data lake concepts.
- Strong attention to data quality, accuracy, and detail.
- Strong analytical and problem-solving skills; able to troubleshoot data workflows to root cause.
- Good communication and documentation skills; able to work with both technical and non-technical stakeholders.
- Willingness to learn AWS cloud data technologies and AI-assisted engineering practices.
- Ability to work independently and take ownership of assigned work, with support scaled to your level.
- Nice to have:
- Hands-on experience with AWS services such as S3, Glue, Athena, Lambda, or DMS.
- Experience with Apache Spark / PySpark or open table formats (Apache Iceberg, Delta Lake, Hudi).
- Experience with BI tools such as Tableau or Power BI.
- Experience using AI coding tools (Claude, Copilot, Codex) or exposure to MCP / LLM integrations.
- Exposure to streaming platforms (Kinesis, Kafka) or workflow orchestration tools.
- Understanding of data governance, PII handling, security, and compliance principles.
- Familiarity with Git, CI/CD, or Infrastructure as Code.
- Experience in e-commerce, marketplace, or gaming domains.
Skills
- AI
- Analytics
- API
- Athena
- AWS
- Aws Glue
- CI/CD
- Cloud
- Data Analytics
- Data Engineering
- Data Governance
- Data Lake
- Data Modeling
- Data Quality
- Data Warehousing
- Delta Lake
- E-commerce
- ELT
- ETL
- Git
- IAM
- Iceberg
- Infrastructure as Code
- Kafka
- Kinesis
- Lakehouse
- Lambda
- LLM
- MCP
- Power BI
- PySpark
- Python
- Serverless
- Spark
- SQL
- Tableau
- Workflow Orchestration