Agentic AI Engineer
Role
As our Agentic AI Engineer, you will build the intelligence layer that sits on top of our knowledge graph: the part of APAS that turns structured regulatory and operational data into deterministic, citation-backed answers and, over time, autonomous agentic workflows. This is a specialist role. We need someone who lives in large language models, retrieval architectures, and agent design, not a generalist ML engineer picking this up as one of several things they do.
Every model you build depends on data structured below you, and every decision you make about grounding, hallucination control, and citation accuracy has to hold up to the same standard the rest of APAS is built on: deterministic, defensible, provable.
This is not a job you'll do in a silo. Expect tight and transparent communication loops across our team.
Must Have
Minimum 3 years of hands-on experience building and shipping production ML/AI systems, focused specifically on LLMs, NLP, or generative AI
Minimum 3 years of production Python experience, with frameworks like PyTorch, LangChain, or Hugging Face Transformers
Deep, practical experience with retrieval-augmented generation (RAG) architectures, including chunking strategies, embeddings, vector search, and hybrid retrieval that combines structured and unstructured data
Proven experience building agentic systems: multi-step reasoning, tool use, task planning, and orchestration (LangGraph, LlamaIndex, AutoGen, or a custom-built framework)
Strong grounding in prompt engineering and techniques for hallucination control, citation accuracy, and deterministic output, especially in high-stakes or regulated domains
Experience integrating LLM applications with knowledge graphs, ontologies, or structured data sources such as SPARQL/RDF. You don't need to own the graph, but you need to work fluently with it
Experience working with function-calling and tool-use models — open-source (e.g., Hermes, and other fine-tuned Llama/Mistral variants built for agentic use) as well as proprietary (e.g., Claude, GPT) — and knowing when a self-hosted, fine-tuned model is the right call versus a foundation model API
Experience fine-tuning or adapting open-source or foundation LLMs for domain-specific applications
Experience designing evaluation frameworks for LLM output quality, accuracy, and safety, not just spot-checking outputs by hand
Comfortable working with cloud infrastructure (AWS or Azure) and collaborating with DevOps on deployment, even though you won't own infrastructure yourself
Strong communicator who can explain model behavior and trade-offs clearly to non-technical stakeholders, including our go-to-market and leadership team
Nice to Have
Experience in regulated, compliance-heavy, or government and public-sector domains, where deterministic and defensible AI output genuinely matters
Familiarity with knowledge graph query languages (SPARQL, Cypher) beyond a surface level
Experience building digital-twin or simulation-style AI features
Prior startup experience, especially pre-revenue or early-stage, and comfort building from zero with a high degree of ambiguity
Open-source contributions to LLM tooling, agent frameworks, or RAG libraries
Exposure to the water utility, civil engineering, or environmental regulatory space
Location
India, Remote
Compensation
$1,500 - $2,500 per month
Skills
As published by ashby · 10 questions
Basics
Your Resume, Full Name (First and Last), Email Address, Location
Short answers (7)
- Phone Number
- LinkedIn Profile
- How did you hear about this opportunity?
- Annual Base Salary Expectations
- When can you start?
- What is your GitHub? optional
- Share with us one agentic AI system you built and shipped to production. What did the agent do, and what tools/frameworks did you use (e.g., LangGraph, function-calling models, vector search)? What broke or didn't work as expected, and how did you fix it?
Pick from a list (3)
- Do you have 3+ years of hands-on experience building and shipping production ML/AI systems, with a focus on LLMs, NLP, and/or generative AI?
- Have you personally built and deployed an agentic system that uses multi-step reasoning and tool use (e.g., with LangGraph, LlamaIndex, AutoGen, or a custom framework), rather than only working with single-turn prompting or chatbot-style LLM applications?
- Are you currently authorized to work in India?