Machine Learning Engineer (DE)
ClearDemand Machine Learning Engineer (DE)
The role requires someone who can bridge data engineering and ML engineering by ensuring that high-quality, well-structured, scalable data flows are available for model training, inference, evaluation, and business reporting.
Key Responsibilities
- Build, maintain, and optimize data pipelines that support ML/ Model workflows, Decision systems, model inference, and validation processes.
- Work with large-scale structured and semi-structured datasets across product catalogs, competitor data, customer data, match data, and validation data.
- Develop robust ETL / ELT workflows using Python, SQL, Spark or equivalent distributed processing frameworks.
- Support feature generation, data reconciliation, data quality checks, and data observability for ML systems.
- Collaborate with ML Engineers and Data Scientists to prepare clean, reliable datasets for model training, evaluation, and inference.
- Improve pipeline performance, scalability, reliability, and cost efficiency.
- Build reusable data processing components and automation utilities.
- Debug data discrepancies across upstream and downstream systems.
- Support operational reporting, dashboards, metric pipelines, and audit workflows.
- Work closely with engineering teams to productionize ML data flows and ensure smooth handoffs between data systems and ML systems.
Required Skills
- Strong programming skills.
- Hands-on experience with data processing pipelines and large datasets.
- Good understanding of ETL / ELT design, data modeling, partitioning, incremental processing, and data validation.
- Experience with tools or platforms such as Spark, Airflow, AWS Glue, Apache Hudi, Databricks, BigQuery, Snowflake, Redshift, Athena, or similar technologies.
- Ability to debug data quality issues and trace data across systems.
- Experience building monitoring, reconciliation, or data quality frameworks.
- Good understanding of APIs, file formats, and data storage formats
- Strong with cloud platforms, preferably AWS and GCP.
- Basic understanding of ML workflows such as training data preparation, inference data generation, evaluation datasets, and feature engineering.
- Experience working with ML pipelines or MLOps workflows.
- Ability to write clean, modular, maintainable, and well-documented code.
Good to Have
- Exposure to vector databases, embeddings, search systems, or retrieval-based systems.
- Experience with product catalog data, e-commerce data, taxonomy, attributes, or entity matching.
- Understanding of model evaluation datasets and metric generation.