Data Engineer
Data Engineer -
• Hands-on production experience with Apache Doris (or a comparable MPP OLAP engine -
StarRocks, ClickHouse, Greenplum).
• Strong, deep SQL expertise - complex analytical queries, window functions, CTEs, query
optimization.
• Solid understanding of Doris architecture (FE/BE), the three data models
(Duplicate/Aggregate/Unique), partitioning, bucketing, and tablet/replica management.
• Experience with Apache Iceberg (or comparable open table formats - Delta Lake, Hudi) and
lakehouse / external-catalog federation.
• Practical experience with Snowflake and/or Apache Spark- enough to read, understand, and
migrate existing workloads.
• Understanding of distributed / MPP query execution: join distribution, runtime filters, memory
management, and data skew.
• Hands-on production experience with Trino (or PrestoSQL/Presto)
• Understanding of distributed query execution: MPP architecture, join distribution, memory/spill behaviour,
partition pruning, and predicate pushdown
• Experience with cloud object storage and columnar file formats (Parquet, ORC)
• Proficiency in at least one programming language (Python, Java, or Scala) for tooling, UDFs, and
Automation
• Version control (Git) and CI/CD for data pipelines.