Data Scientist
Summary
The Data Scientist will build and maintain machine learning models and data pipelines within a Databricks Lakehouse environment to support marketing intelligence. The role involves data preparation, feature engineering, and model evaluation using Python, PySpark, and SQL.
Data Scientist
(Machine Learning & Databricks)
π° Salary | $2,000 USD/month (Bi-weekly payments)
π Location | Remote from Latin America
π Schedule | Full-time
πΌ Employment Type | Contractor
πΊπΈ Perks & Time Off | U.S. holidays off, generous PTO, internal training, and mentorship
About the Client
Latino Legends is partnering with a high-growth U.S. marketing intelligence company to find top-tier nearshore talent. Our client does not just report on dataβthey build with it, driving marketing decisions for over 1,000 dealer clients across digital and direct channels. They operate on a cloud-native Databricks Lakehouse platform and foster an AI-forward culture where modern tooling accelerates analysis and deployment.
The Role
As a Data Scientist, you will work primarily inside Databricks building datasets, features, and machine learning models that power marketing intelligence. You will spend your time on data preparation, exploratory analysis, feature engineering, and model training/evaluation, supported by senior engineers and architects on the production side.
Key Responsibilities
Data & Modeling (80%)
Build and maintain data pipelines and transformations in Databricks using Python, PySpark, and SQL.
Perform exploratory data analysis to understand marketing, campaign, and customer datasets.
Engineer features, prepare training datasets, and train, evaluate, and tune machine learning models.
Support model deployment and monitoring in production alongside senior engineers.
Investigate and resolve data quality issues in source feeds and pipelines.
Leverage AI coding tools efficiently while maintaining code quality and deep understanding.
Collaboration & Quality (20%)
Participate in sprint planning, standups, and retrospectives within an Agile/SCRUM framework.
Write tests, data validation checks, and participate in code reviews.
Document datasets, features, model assumptions, and results.
Present findings and recommendations to both technical and business stakeholders.
Core Qualifications
Required:
Experience: 1-3 years of hands-on experience in data science, machine learning, or analytics engineering (including internships, co-ops, research, or complex projects).
Education: Degree in Computer Science, Statistics, Mathematics, Engineering, Data Science, or comparable practical experience.
Languages & Core Tooling: Solid proficiency with Python (pandas, NumPy, scikit-learn) and SQL (joins, aggregations, window functions).
Data Science Fundamentals: Solid understanding of supervised learning, train/test splits, feature engineering, cross-validation, and evaluation metrics (ROC AUC, RMSE, precision, recall).
Statistics & Data Handling: Working knowledge of descriptive statistics, hypothesis testing, data profiling, and data cleaning.
Modern Workflow: Proficiency with Git, collaborative development, and AI-assisted development tools (Cursor, GitHub Copilot, Claude, etc.).
Communication: Strong written and spoken English skills to effectively explain complex findings to non-technical audiences.
Nice to Have:
Hands-on experience with Databricks (Unity Catalog, clusters, jobs, notebooks) or PySpark.
Exposure to MLflow, Delta Lake (medallion architecture), or Azure cloud services.
Experience with time-series forecasting, uplift modeling, or audience segmentation.
Familiarity with visualization tools (Power BI, Tableau, Plotly) or CI/CD pipelines (GitHub Actions, Azure DevOps).
Interest in Large Language Models (LLMs) and generative AI applications.
Why Apply Through Latino Legends?
Modern Stack: Hands-on experience in Databricks on a modern lakehouse platform with zero legacy code clutter.
Mentorship & Career Growth: Work side-by-side with senior architects who review your code and help you transition from notebook experiments to production-level machine learning.
High Impact: Your models directly inform millions of dollars in active marketing spend.
Distributed Culture: Enjoy a collaborative, cross-border culture across the U.S. and Latin America.