#EG AI / LLM Specialist
Summary
Build and productionize LLM-powered features: design prompts, integrate foundation models, automate evaluation/benchmarking, and monitor quality, safety, and cost for GenAI solutions.
This role is an AI/LLM Engineer focused on prompt engineering, model integration, evaluation, and production quality within NCS AI Central’s Forward Deployed Engineering model. This role helps take GenAI solutions from POC/POV to production by building prompts, integrating foundation models, creating automated evaluation and benchmarking frameworks, and continuously monitoring quality, safety, hallucination, bias, drift, latency, and cost.
You will work closely with AI Architects, Solution Architects, AI Engineers, and Testers to provide evidence-based model recommendations, support production readiness, and maintain reusable internal assets such as prompt libraries, evaluation templates, and benchmark datasets.
What will you do?
1. Model Integration & Prompt Engineering
- Design, test, and optimise prompts and prompt chains for production use cases — balancing accuracy, latency, and cost.
- Integrate foundation models into applications via APIs and gateways; advise on model/version selection for a given use case alongside AI Architects.
- Support light fine-tuning and instruction-tuning work (LoRA/PEFT and similar techniques) where a use case calls for it, in partnership with AI Engineers.
2. Evaluation Framework & Benchmarking
- Design and maintain evaluation harnesses (golden datasets, benchmark suites) that measure LLM/agentic system accuracy, consistency, and safety.
- Run comparative model benchmarking (accuracy, latency, cost-per-query) to feed technical evidence into model-selection decisions led by AI/Solution Architects.
- Build repeatable, automated regression suites that run on every prompt, model, or pipeline change, wired into CI.
3. Quality & Adversarial Testing
- Proactively red-team AI systems — adversarial prompting, edge-case and jailbreak testing — to surface failure modes before clients do.
- Detect and quantify hallucination, bias, and drift in production and pre-production systems, producing clear, defensible metrics (not qualitative impressions).
4. Governance & Reporting
- Feed evaluation evidence into the PRR (Production Readiness Review) Evaluation & Quality pillar, supporting engagement teams at the Scale gate.
- Maintain evaluation and prompt-version documentation and reporting standards that align with client compliance and audit needs (e.g., government AI governance requirements).
5. FDE & Development/Maintenance Coverage
- During FDE engagements: rapidly prototype prompts and model integrations, and stand up lightweight evaluation harnesses to compare candidate models/approaches during POC/POV, giving the team fast, evidence-based go/no-go signals.
- During system development & maintenance engagements: own ongoing prompt/model tuning and run continuous evaluation and regression monitoring on live production systems, flagging quality degradation over time.
- Contribute reusable prompt libraries, evaluation templates, and benchmark datasets back into the shared internal asset library for reuse across engagements.
6. Collaboration
- Provide technical benchmark evidence and integration recommendations to AI Architects and Solution Architects, who own the final client-facing model recommendation and proposal.
- Partner with AI Engineers and Testers to distinguish functional QA (does it work) from output-quality evaluation (is it right), and to hand off tuned prompts/models cleanly into production builds.
The ideal candidate should possess:
- 3+ years working hands-on with LLMs across prompt engineering, model integration, and evaluation — not evaluation alone.
- Practical experience designing and optimising prompts and prompt chains for production applications, and integrating models via APIs/gateways.
- Strong grasp of evaluation methodologies — accuracy/hallucination/toxicity metrics, human-in-the-loop evaluation, A/B testing.
- Hands-on scripting ability (Python) to build and automate evaluation harnesses and integration/testing pipelines.
- Statistical literacy — able to design a representative test/benchmark set and interpret results rigorously, not anecdotally.
- Clear, structured written communication — able to translate evaluation results and model recommendations into a defensible report for both engineering and client audiences.
- Working knowledge of the China AI model/tech stack (e.g., DeepSeek, Qwen, GLM, Kimi, MiniMax) — deployment patterns, licensing, and self-hosting requirements.
Preferred Qualifications
- Hands-on fine-tuning/instruction-tuning experience (LoRA/PEFT or similar) on open-weight models.
- Experience with LLM evaluation tooling (RAGAS, DeepEval, TruLens, promptfoo) or building custom eval frameworks.
- Exposure to red-teaming/adversarial testing practices for generative AI systems.
- Familiarity with regulated-sector AI governance expectations (Healthcare, Government, Financial Services).
- Prior experience supporting client-facing presales or solutioning conversations with technical evidence (without owning the proposal).
- Hands-on benchmarking or integration experience with Chinese open-weight models (DeepSeek, Qwen, GLM) alongside Western models.
Tech Stack (Illustrative)
- Languages: Python (primary), SQL
- Prompt & Integration: LangChain/LlamaIndex, model gateways (LiteLLM, Bedrock, Azure OpenAI), prompt-versioning tools
- Fine-Tuning: LoRA/PEFT, Hugging Face Transformers (where applicable)
- Eval Tooling: RAGAS, DeepEval, TruLens, promptfoo, custom harnesses
- LLM Runtime: OpenAI/Azure OpenAI/Bedrock/Vertex APIs; DeepSeek/Qwen/GLM (China stack)
- Data & Reporting: Pandas, Jupyter, BI/reporting tools for evaluation dashboards
- CI Integration: GitHub Actions/GitLab CI for automated regression evaluation
Why Join NCS?
Grow with Us
- Work on cutting-edge AI products that shape the future of technology
- Collaborate with talented, passionate teams across research, engineering, and design
- Access continuous learning opportunities and career development pathways
Make an Impact
- Transform AI research into products that solve real problems for clients and users
- Drive innovation in a leading Technology Services Firm with regional presence
- Contribute to building a better future through responsible, human-centred AI
Thrive in Our Culture
- Experience a human-to-human approach where relationships and collaboration matter
- Be part of Team NCS, where bold ideas meet practical execution
- Enjoy a supportive environment that values diversity, inclusion, and respect
We are driven by our AEIOU beliefs—Adventure, Excellence, Integrity, Ownership, and Unity—and we seek individuals who embody these values in both their professional and personal lives. We are committed to our Impact: Valuing our clients, Growing our people, and Creating our future.
Together, we make the extraordinary happen.
Learn more about us at ncs.co and visit our LinkedIn career site.
Scam Alert
We are aware of fraudulent job offers and impersonations of NCS recruiters. Phishing emails using convincing-looking but fake addresses are also commonly used to trick you into thinking that they come from official NCS sources.
Please note that all official communications from NCS Group will only be sent from verified corporate email addresses. Always check that the sender’s email address ends with the genuine NCS domain, @ncs.com.sg and beware of extra letters, symbols or misspellings. When in doubt, verify the sender’s identity by contacting us at reachus@ncs.com.sg.