Senior Data Engineer
Transflo Senior Data Engineer
- Architect, build, and evolve a scalable enterprise data warehouse on Amazon Redshift, applying industry-standard concepts including star schemas, snowflake schemas, normalization, denormalization, referential integrity, and performance optimization strategies
- Design and implement bronze, silver, and gold data layer architecture (medallion architecture): raw ingestion, cleansed and standardized intermediate layers, and curated, business-ready data products optimized for analytics consumption
- Develop dimensional data models, fact and dimension tables, slowly changing dimensions (SCDs), and aggregate structures that support BI tooling, ad-hoc analytics, and downstream API consumption
- Apply rigorous data modeling practices including schema design, constraint definition, indexing strategy, sort keys, distribution keys, and query plan optimization within Redshift and connected systems
- Build, own, and maintain robust batch and streaming data pipelines that ingest data from disparate source systems including REST APIs, flat files, IBM DB2, MySQL, Amazon Aurora, Amazon DynamoDB, and PostgreSQL
- Implement real-time and near real-time data streaming architectures using AWS-native services such as Kinesis Data Streams, Kinesis Firehose, MSK (Managed Kafka), and EventBridge to support low-latency data delivery requirements
- Design pipeline frameworks for data extraction, transformation, and loading (ETL/ELT) using tools such as AWS Glue, dbt, Apache Airflow, or equivalent orchestration platforms
- Ensure pipeline reliability, idempotency, fault tolerance, and automated recovery; build alerting and observability into every data workflow from day one
- Own data quality end-to-end: design and implement automated profiling, cleansing, deduplication, standardization, and validation frameworks that enforce data integrity at each layer of the medallion architecture
- Build and continuously evolve tooling and processes to support data governance including data cataloging, lineage tracking, metadata management, access controls, and data classification
- Define and enforce data contracts between source systems and the warehouse, establishing clear SLAs for freshness, completeness, and accuracy
- Partner with data consumers — Data scientists, Data analytics engineers, BI developers, product managers, and external API clients — to understand consumption patterns and ensure data products meet quality and performance expectations
- Support the architecture and buildout of a reliable, scalable Data as a Service (DaaS) product, enabling external and internal consumers to access curated Transflo data via governed APIs and data sharing mechanisms
- Contribute to the data platform infrastructure using infrastructure-as-code practices (Terraform), ensuring all data infrastructure is version-controlled, reproducible, and auditable
- Design for scale: apply partitioning strategies, workload management (WLM) tuning, concurrency scaling, and caching patterns to sustain performance under high-traffic analytical and operational workloads
- Champion security and compliance best practices across the data platform: column-level security, row-level access controls, encryption, and audit logging
- Collaborate with software engineers, mobile platform teams, and DevOps to ensure upstream application data is well-structured, well-documented, and reliably delivered to the data platform
- Leverage AI-assisted development practices and tooling to accelerate pipeline development, automate data quality checks, and improve engineering velocity
- 5+ years of professional data engineering experience with a track record of building and operating production-grade data warehouses and pipeline infrastructure
- Expert-level experience with Amazon Redshift including cluster sizing, WLM configuration, distribution and sort key optimization, vacuuming, and query plan analysis
- Deep proficiency in SQL for complex analytical queries, window functions, CTEs, stored procedures, and performance tuning across Redshift and ANSI-compatible engines
- Hands-on experience ingesting data from heterogeneous source systems: REST APIs, IBM DB2, MySQL, Amazon Aurora (MySQL and PostgreSQL-compatible), Amazon DynamoDB, PostgreSQL, and file-based sources (CSV, JSON, Parquet, Avro)
- Proven experience designing and implementing medallion (bronze/silver/gold) or equivalent layered data architectures at enterprise scale
- Strong working knowledge of star schema and snowflake schema design, dimensional modeling theory, slowly changing dimensions, and fact table granularity decisions
- Experience building real-time or near real-time data pipelines using streaming technologies such as Amazon Kinesis, Apache Kafka (or Amazon MSK), or equivalent
- Proficiency with ETL/ELT orchestration tools such as AWS Glue, dbt, Apache Airflow, or AWS Step Functions
- Demonstrated experience implementing data governance practices: data catalogs (AWS Glue Data Catalog, Apache Atlas, or equivalent), lineage, metadata tagging, and access control frameworks
- Infrastructure-as-code experience with Terraform for provisioning and managing data infrastructure on AWS
- Strong Python skills for pipeline development, data transformation logic, and automation scripting
- Deep understanding of data reliability engineering: idempotency, exactly-once processing, late-arriving data handling, schema evolution, and SLA-driven pipeline design
- Experience in the transportation, logistics, trucking, or fleet management industry, or with high-volume transactional SaaS platforms processing operational telemetry data is a huge plus
- Experience building DaaS or data product offerings including governed external data APIs, Redshift Data Sharing, or AWS Data Exchange integrations
- Knowledge of columnar storage formats (Parquet, ORC) and lakehouse patterns using Amazon S3 as a data lake layer in conjunction with Redshift Spectrum or AWS Glue
- Familiarity with BI and analytics consumption tools such as Tableau, Power BI, Amazon QuickSight, or Looker and how data model design decisions impact end-user query performance
- Experience with data observability platforms such as Monte Carlo, Great Expectations, or dbt tests for automated data quality monitoring
- Contributions to reusable data platform tooling, shared dbt packages, or internal data engineering frameworks
- Experience working in fully remote, distributed engineering teams