Data Scientist
Summary
Build and validate predictive models in Python (pandas, scikit-learn, TensorFlow/PyTorch) and SQL, then translate insights into dashboards and business decisions.
Must Have
- 3–5 years of experience in data science, analytics, or a quantitative modeling role.
- Hands‑on proficiency in Python (pandas, NumPy, scikit‑learn) and SQL for building and evaluating models end to end.
- Demonstrated ability to design experiments, validate model performance, and guard against overfitting and data leakage.
- Clear, reproducible, and well‑documented analytical practice.
- Ability to communicate findings clearly to both technical colleagues and business stakeholders.
Nice to Have
- Experience with visualization tools such as matplotlib, seaborn, Power BI, or Tableau.
- Exposure to cloud analytics environments, ideally Oracle Cloud Infrastructure (OCI).
- Experience with time‑series analysis or natural‑language data.
- Familiarity with collaborative workflows and code review.
- Exposure to working with engineering teams on model handoff.
- Data science certifications.
Responsibilities
- Conduct exploratory data analysis to uncover patterns, relationships, and opportunities in client and operational data.
- Build, train, and evaluate predictive and descriptive models using Python (pandas, scikit‑learn) and SQL.
- Build and evaluate deep learning models, such as neural networks, using frameworks like TensorFlow or PyTorch where they suit the problem.
- Perform feature engineering, data cleaning, and dataset preparation for modeling.
- Design and analyze experiments, including A/B tests, to measure the impact of changes and interventions.
- Prototype analytical solutions and iterate on them based on stakeholder and senior data scientist feedback.
- Create clear visualizations, dashboards, and summaries that translate analysis into business insight.
- Validate model performance using appropriate metrics and guard against overfitting and data leakage.
- Document methodology, assumptions, and results to ensure reproducibility and knowledge sharing.
- Collaborate with senior data scientists, engineers, and analysts on larger initiatives.
- Support the preparation of analytical reports and presentations for clients and internal teams.
- Maintain and improve existing analytical code and notebooks.
Qualifications
- Bachelor's degree in a quantitative field (Statistics, Mathematics, Computer Science, Engineering, Data Science) or equivalent experience; Master's an advantage.
- Proficiency in Python for data analysis (pandas, NumPy, scikit‑learn) and working knowledge of SQL.
- Solid grounding in statistics and core machine‑learning techniques (regression, classification, clustering).
- Working knowledge of deep learning concepts and frameworks such as TensorFlow or PyTorch.
- Experience with feature engineering and preparing real‑world, messy datasets for modeling.
- Ability to communicate findings clearly to technical and business audiences.
- Understanding of model evaluation metrics and validation techniques.
- Familiarity with version control (Git) and reproducible analysis practices.