Point your AI agent at freehire and let it find you a job.

Get the CLI →

jobgether

NewBe an early applicant

AI Engineer — LLM / VLM

Posted Updated
Discussion

Summary

Build and deploy production-grade AI applications powered by LLMs and vision-language models — including RAG pipelines, fine-tuning (LoRA/QLoRA), AI agents, and optimized model serving — for a partner company hiring via Jobgether. Core stack: Python, PyTorch, Hugging Face, FastAPI, Docker, and AWS/Azure/GCP, working remotely from India (Kolkata-based).

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a AI Engineer — LLM / VLM based in India.

As an AI Engineer specializing in LLMs and VLMs, you will design, develop, and deploy production-grade AI solutions across text and multimodal use cases. You will work with advanced foundation models, retrieval-augmented generation, prompt engineering, fine-tuning, and AI agent architectures. The role combines deep machine learning expertise with strong software engineering practices to turn AI prototypes into reliable production services. You will build solutions capable of understanding text, images, PDFs, charts, tables, and other complex documents. You will also focus on model evaluation, inference optimization, safety, latency, and cost efficiency. Working closely with ML engineers, software engineers, and product teams, you will help deliver practical AI capabilities from experimentation through production.

Accountabilities:

  • Design, develop, and deploy production-ready AI applications powered by Large Language Models and Vision-Language Models, selecting appropriate models and architectures based on business and technical requirements.
  • Build end-to-end RAG pipelines covering document ingestion, chunking, embedding generation, retrieval, reranking, and response generation, with a focus on accuracy, relevance, scalability, and reliability.
  • Work with leading foundation models and multimodal models, including GPT, Claude, Gemini, Llama, Mistral, Qwen, and comparable technologies, evaluating their suitability for different use cases.
  • Develop multimodal AI solutions that can process and reason across text, images, PDFs, charts, tables, and complex business documents.
  • Apply prompt engineering, supervised fine-tuning, LoRA/QLoRA, and other model adaptation techniques to improve model performance for specific applications.
  • Build AI agents and tool-calling workflows where they provide meaningful value, designing systems that can combine reasoning, retrieval, external tools, and business processes.
  • Optimize model inference and serving for latency, throughput, memory consumption, scalability, and cost, balancing technical performance with production requirements.
  • Develop production APIs and AI services using Python, FastAPI, Docker, and cloud infrastructure, while applying sound software engineering, version control, testing, and deployment practices.
  • Create evaluation frameworks and metrics to measure AI system accuracy, relevance, hallucination rates, latency, safety, and overall production quality.
  • Collaborate with ML engineers, software engineers, and product teams to transform experimental AI concepts into robust, maintainable, and production-ready solutions.
  • Requirements

    • Bring strong Python programming skills and solid software engineering fundamentals, with the ability to build clean, maintainable, and production-ready AI applications.
    • Have hands-on professional experience working with Large Language Models and/or Vision-Language Models and a strong understanding of how modern generative AI systems are developed and deployed.
    • Demonstrate a solid understanding of Transformers, attention mechanisms, tokenization, embeddings, model inference, and the underlying concepts that enable modern language and multimodal models.
    • Have practical experience with PyTorch and Hugging Face Transformers, including working with or adapting open-source and foundation models.
    • Bring hands-on experience designing and implementing RAG systems and vector-search solutions, including knowledge of retrieval and embedding strategies.
    • Demonstrate knowledge of prompt engineering and LLM evaluation methodologies, with an understanding of how to assess model quality, hallucinations, relevance, and safety.
    • Have experience developing APIs and REST services and working with Git, Docker, CI/CD, and modern software development practices.
    • Be familiar with vector databases or vector-search technologies such as FAISS, Milvus, Pinecone, Weaviate, or pgvector.
    • Understand cloud-based AI infrastructure and deployment, with experience in AWS, Azure, GCP, or comparable cloud environments.
    • Have experience with multimodal models such as Qwen-VL, LLaVA, Gemini, GPT vision models, or similar technologies, along with an understanding of image preprocessing and document or image understanding.
    • Experience with OCR, document intelligence, image classification, object detection, or visual question answering is a plus, as is the ability to build integrated pipelines combining vision, language, and retrieval.
    • Bring strong analytical and problem-solving abilities, curiosity for emerging AI technologies, and the ability to collaborate effectively with technical and product stakeholders in an experienced engineering environment.
    • Benefits

      • Remote work from India, with the role based in Kolkata and structured as a remote position.
      • Full-time employment within a technology-focused environment working on advanced AI and generative AI applications.
      • Opportunities to work hands-on with LLMs, VLMs, RAG, AI agents, multimodal AI, fine-tuning, and model serving.
      • Exposure to leading AI technologies and models, including GPT, Claude, Gemini, Llama, Mistral, Qwen, and other emerging foundation models.
      • The opportunity to build production-grade AI systems rather than limiting your work to experimentation and prototypes.
      • Broad technical exposure across machine learning, software engineering, cloud infrastructure, APIs, vector databases, evaluation frameworks, and AI deployment.
      • The opportunity to collaborate with ML engineering, software engineering, and product teams while contributing directly to the development of practical AI capabilities.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Skills

See also

AI Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available