Data Scientist- Air Quality

Open 34d reposted 2× · 2 open copies

We are seeking a talented Data Scientist with competency in Data Engineering capabilities to design, build, and deploy some advanced analytical solutions across a range of environmental sectors. This role is ideal for a professional who can develop machine learning models, support the engineering of robust data pipelines, and contribute to the broader technical architecture.

The successful candidate will be able to apply analytical techniques, work with complex and large-scale datasets, collaborate with cross-functional teams, and support the development of high-quality data infrastructure enabling reliable, scalable data-driven decision-making.

Data Science & Advanced Analytics
• Build, validate, and deploy sophisticated predictive and statistical models using modern Python-based libraries. (e.g., scikit-learn, TensorFlow, PyTorch)
• Conduct exploratory data analysis, statistical modelling, hypothesis testing, and insight generation to support strategic business decisions.
• Develop high-quality, interpretable data visualisations using Matplotlib, Seaborn, and similar tools.
• Translate complex analytical outcomes into clear, actionable insights for both technical and non-technical stakeholders.
Data Engineering & Pipeline Development
• Design and implement robust, scalable data orchestration pipelines using tools such as Airflow or Dagster.
• Work with big data technologies (Spark, Hadoop, Dask) to process and analyse large-scale datasets efficiently.
• Define and implement data models using ER diagrams, normalisation techniques (e.g., star schema, Snowflake), and modern ORM frameworks (e.g., SQLAlchemy).
• Provision, administer, and optimise relational databases, including schema design, indexing, and access management.
• Build high-performance APIs using frameworks such as FastAPI or Flask, following best RESTful design practices.

Software Engineering and DevOps Integration
• Develop modular, maintainable, object-oriented Python code using best practice design patterns and testing practices (CI/CD, unit tests, code reviews).
• Containerise and deploy applications using Docker and Kubernetes; support infrastructure provisioning through Infrastructure-as-Code tools (e.g., Terraform, Ansible).
• Contribute to system architecture discussions, microservices design, and integration with cloud platforms notably Azure.
Collaboration and Leadership
• Work collaboratively across technical and non-technical teams to embed data-driven decision-making.
• Support large-scale analytical or data engineering projects, adherence to best practices in code quality, reproducibility, and documentation.
• Shape AI use cases with internal stakeholders and, where needed, client counterparts (problem framing, value hypothesis, data readiness, and delivery roadmap), and communicate trade-offs clearly.
Essential Technical Skills and Experience
• Python programming expertise, including functions, classes, modules, packages, and advanced OOP features.
• Solid statistical knowledge (degree-level or equivalent).
• Demonstrable experience in machine learning model development and evaluation.
• Proficiency with relational databases and database design.
• Experience designing and maintaining data pipelines and big-data processing workflows.
• Familiarity with containerisation, Kubernetes, CI/CD, and cloud infrastructure.
• Experience deploying models and data pipelines into production environments.
• Good understanding of data modelling, APIs, and software engineering best practices.
Professional Competencies
• Strong problem-solving ability, capable of resolving ambiguous or complex analytical challenges.
• Ability to communicate complex technical topics clearly to diverse stakeholders.
• Advocates for high-quality software development practices.
• Comfortable mentoring team members and leading aspects of project delivery.
• Ability to collaborate with cross-functional teams.
Desirable Skills
• Experience with microservice architectures.
• Knowledge of cloud cost optimisation and resource management.
• Ability to configure and interpret monitoring and alerting systems.
• Familiarity with data ethics, governance, and privacy best practices.
• Some background in developing LLM, generative AI and Responsible AI based solutions

  • Bachelor’s/master’s degree in a quantitative field (e.g., Mathematics, Computer Science, Engineering, Economics, Data Science) or equivalent industry experience.
  • Evidence of continued professional development in data science, data engineering, or cloud technologies.

BGV:

  • Employment with WSP India is subject to the successful completion of a background verification (“BGV”) check conducted by a third-party agency appointed by WSP India.

  • Candidates are advised to ensure that all information provided during the recruitment process — including documents uploaded — is accurate and complete, both to WSP India and its BGV partner”.