freehire launches on Product Hunt on 26 August.

Follow →

Data Scientist

Summary

Build and deploy NLP/LLM models for automotive applications, using transformers, LangChain, and cloud tools to power chatbots, QA, and semantic search.

Data Scientist

Location: Chennai, Tamil Nadu, India; Hybrid

ML/DL Skills:

  • High familiarity in the use of DL theory/practices in NLP applications

  • Comfort level to code in ADK, A2A, AgentSkills, Ontology, Huggingface, LangGraph, LangChain, Chainlit, Tensorflow and/or Pytorch, Scikit-learn, Numpy and Pandas

  • Comfort level to use two/more of open source NLP modules like SpaCy, TorchText, fastai.text, farm-haystack, and others

NLP Skills:

  • Knowledge in fundamental text data processing (like use of regex, token/word analysis, spelling correction/noise reduction in text, segmenting noisy unfamiliar sentences/phrases at right places, deriving insights from clustering, etc.,)

  • Have implemented in real-world BERT/or other transformer fine-tuned models (Seq classification, NER or QA) from data preparation, model creation and inference till deployment

Python Project Management Skills

  • Familiarity in the use of Docker tools, pipenv/conda/poetry env

  • Comfort level in following Python project management best practices (use of setup.py, logging, pytests, relative module imports,sphinx docs,etc.,)

  • Familiarity in use of Github (clone, fetch, pull/push,raising issues and PR, etc.,)

Cloud Skills and Computing:

  • Use of GCP services like BigQuery, Cloud function, Cloud run, Cloud Build, VertexAI,

  • Good working knowledge on other open source packages to benchmark and derive summary

  • Experience in using GPU/CPU of cloud and on-prem infrastructures

  • Skillset to leverage cloud platform for Data Engineering, Big Data and ML needs.

Deployment Skills:

  • Use of Dockers (experience in experimental docker features, docker-compose, etc.,)

  • Familiarity with orchestration tools such as airflow, Kubeflow

  • Experience in CI/CD, infrastructure as code tools like terraform etc.

  • Kubernetes or any other containerization tool with experience in Helm, Argoworkflow, etc.,

  • Ability to develop APIs with compliance, ethical, secure and safe AI tools.

UI:

  • Good UI skills to visualize and build better applications using Gradio, Dash, Streamlit, React, Django, etc.,

  • Deeper understanding of javascript, css, angular, html, etc., is a plus.

Data Engineering:

  • Skillsets to perform distributed computing (specifically parallelism and scalability in Data Processing, Modeling and Inferencing through Spark, Dask, RapidsAI or RapidscuDF)

  • Ability to build python-based APIs (e.g.: use of FastAPIs/ Flask/ Django for APIs)

  • Experience in Elastic Search and Apache Solr is a plus, vector databases.

Ford Global Data Insights and Analytics (GDIA) is looking for professionals experienced in NLP/LLM/GenAI, who are hands-on and can employ many NLP/Prompt engineering techniques from traditional statistical/ML NLP to DL-based sequence models and transformers in their day-to-day work. You'll be working alongside leading technical experts from all around the world, on a variety of products involving Sequence/token classification, QA/chatbots, translation, semantic/search and summarization, among others.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available