Data Architect
Summary
Designs and builds scalable data platforms, pipelines, and architectures using Kafka, Pub/Sub, Kubernetes, and other tools; leads PODs, makes architecture decisions, mentors engineers, and deploys GCP resources as code.
Key Responsibilities:
Data Engineering Architecture (Primary)
- Design and build scalable data platforms, pipelines, and architectures
- Define batch and streaming frameworks (Kafka, Pub/Sub, OpenShift/Kubernetes, GitLab CI/CD, ArgoCD, Optional - Spring Batch)
- Establish data ingestion, transformation, and storage patterns
- Design event-driven and real-time data solutions
Data Processing & Integration
- Build real-time data pipelines and streaming integrations
- Standardize data formats (Avro, JSON, Parquet) and connectors
- Enable API-driven (optional) and event-based system integrations
- Batch Processing, MongoDB, Redis, Helm Charts, Kibana /Elasticsearch, Artifactory
Data Quality & Governance
- Ensure data quality, consistency, and reliability
- Implement data validation, lineage, and traceability
- Define failure handling and recovery strategies
Automation & Optimization
- Drive automation using Python / Bash scripting
- Continuously improve pipeline performance and scalability
- GCS , Ingress / API Gateways, Enterprise Platform Services
Requirements
Required Skills:
- POD Leadership - Act as technical lead for the POD
- Drive architecture decisions, design reviews, and best practices
- Collaborate with data, platform, and product teams
- Mentor engineers and build data engineering capability
- Ensure smooth work progression (To Do → In Progress → Done) and remove blockers
- Streaming is of lower priority as we are migrating out of real-time for our monetization platforms.
- Deploy data engineering GCP resources as code (IaaC)
Benefits
What We Offer:
- Competitive salaries and comprehensive health benefits.
- Flexible work hours options.
- Professional development and training opportunities.
- A supportive and inclusive work environment.