Data Engineer
Summary
Builds and maintains cloud-based data pipelines, warehouses, and AI/ML-ready datasets for Ford’s global analytics and AI initiatives, focusing on scalability, data quality, and GenAI infrastructure.
We are looking for a hands-on Data Engineer with 3+ years of experience building production-grade data pipelines, cloud data platforms, and automated data workflows. You are comfortable working across structured, semi-structured, and unstructured data; you understand the importance of data quality, lineage, security, and cost optimization; and you are excited to build the data foundation required for modern AI, ML, and GenAI use cases.
- Understand business, analytics, and AI use cases and translate them into scalable data engineering solutions.
- Design, build, and maintain reliable batch and streaming data pipelines for ingestion, transformation, validation, and publishing.
- Develop curated, reusable, and well-documented data products that support BI dashboards, analytics applications, ML models, and GenAI-enabled solutions.
- Implement strong data quality checks, observability, lineage, metadata management, and monitoring practices to improve trust in enterprise data assets.
- Write clean, modular, and well-tested code using Python, SQL, and modern data engineering frameworks.
- Use cloud-native technologies such as BigQuery, Dataflow, Dataproc, Cloud Composer/Airflow, Dataform, DBT, Spark, or equivalent tools to deliver resilient data solutions.
- Enable AI/ML and GenAI teams by preparing high-quality feature datasets, vector-ready datasets, document corpora, and governed data access patterns.
- Partner with data scientists, ML engineers, product owners, and business stakeholders to support experimentation, model deployment, and production analytics.
- Apply DataOps practices including CI/CD, version control, automated testing, reusable templates, release management, and production support standards.
- Optimize pipeline performance, storage usage, compute cost, and reliability across cloud-based data platforms.
- Support data governance, privacy, access control, and compliance expectations for enterprise and AI-ready data assets.
- Stay current with advances in cloud data engineering, AI data infrastructure, orchestration, data quality, and GenAI-enabling technologies.
Minimum Qualifications:
- Bachelor’s or Master’s degree in Computer Science, Data Engineering, Information Systems, Engineering, Statistics, Mathematics, or related technical field.
- 3+ years of hands-on experience in data engineering, ETL/ELT development, data warehousing, or cloud-based data platform delivery.
- Strong proficiency in SQL and Python for data extraction, transformation, automation, testing, and production support.
- Experience designing and operating scalable pipelines on cloud platforms such as Google Cloud Platform, AWS, Azure, or equivalent enterprise data ecosystems.
- Experience with modern data platforms and tools such as BigQuery, Spark, Dataflow, Dataproc, Airflow/Cloud Composer, Dataform, DBT, or similar technologies.
- Good understanding of data modeling, dimensional modeling, partitioning, clustering, performance tuning, and cost optimization.
- Working knowledge of data quality frameworks, monitoring, alerting, metadata, lineage, and production support practices.
- Familiarity with Git, CI/CD, agile delivery, code reviews, documentation, and reusable engineering standards.
Strong communication skills with the ability to explain technical solutions clearly to engineering, analytics, and business stakeholders.
Preferred Qualifications:
- 5+ years of experience delivering enterprise data engineering solutions in cloud-native environments.
- Experience building data products for AI/ML, GenAI, semantic search, retrieval-augmented generation, feature engineering, or model monitoring use cases.
- Experience working with unstructured data such as documents, logs, text, images, transcripts, or embeddings, and preparing them for downstream AI consumption.
- Hands-on experience with DataOps, MLOps enablement, pipeline observability, automated testing, and production incident resolution.
- Experience migrating legacy workflows from Hadoop, Alteryx, or on-premise platforms to modern cloud services.
- Experience with APIs, microservices, event-driven architectures, streaming data, or real-time analytics.
- Cloud certifications in Google Cloud Platform, AWS, Azure, or relevant data engineering technologies.
- Experience mentoring junior engineers, defining engineering standards, or contributing reusable platform accelerators.
Skills
- Agile
- AI
- Airflow
- Alteryx
- Analytics
- API
- Automation
- AWS
- Azure
- BigQuery
- CI/CD
- Cloud
- Cloud Native
- Data Engineering
- Data Governance
- Data Modeling
- Data Pipelines
- Data Quality
- Data Warehousing
- dbt
- Dimensional Modeling
- ELT
- Embeddings
- ETL
- Event Driven Architecture
- Feature Engineering
- GCP
- Generative AI
- Git
- Hadoop
- Machine Learning
- Metadata Management
- Microservices
- MLOps
- Model Deployment
- Observability
- Python
- RAG
- Semantic Search
- Spark
- SQL
- Statistics
- Test Automation
- Version Control