LLM Platform Engineer
Summary
As an LLM Platform Engineer at Schemata, you build and scale the AI platform that turns documents, video, imagery, audio, and 3D scene data into real-time, grounded guidance for training and maintenance applications. Day to day: multimodal knowledge ingestion pipelines, RAG retrieval improvements, evals, and production inference/deployment across cloud, on-prem, and air-gapped environments, mainly
About Schemata
At Schemata, we are transforming the $400 B virtual‑training and simulation market by fusing 3D computer vision, neural rendering and large multimodal models inside highly regulated industries. Our platform delivers photorealistic, intelligent 3D experiences, demanding robust spatial reasoning, high‑performance data pipelines and seamless integration between traditional graphics and AI‑driven perception.
About the Role
We are seeking a highly skilled LLM Platform Engineer to join our team full‑time. You will play a foundational role in designing, building and optimizing the AI systems that turn heterogeneous knowledge — technical documentation, video, imagery, voice, and 3D scene data — into grounded, real‑time guidance for next‑generation training and maintenance applications.
This is a high‑impact, cross‑functional role: you will work end‑to‑end from evaluating emerging methods to production inference and performance optimization, ensuring our platform retrieves the right knowledge, reasons over it reliably, and responds in real time across diverse deployment environments.
Core Responsibilities
Design and build multimodal knowledge base ingestion: pipelines that convert documents, video, images, audio, and structured data into a unified, queryable semantic layer with rich embeddings and metadata, including bringing in external sources and messy real‑world formats as they arise.
Improve our retrieval‑augmented generation (RAG) systems: chunking and indexing strategies, hybrid retrieval, reranking, grounding controls, and context assembly for multimodal queries.
Build the knowledge base lifecycle: handle versioned documentation updates, incorporate SME corrections and annotations, and design validation and freshness mechanisms so the knowledge base improves continuously.
Scale the feedback loop: capture user signals (ratings, corrections, query patterns), route them into measurable improvements to retrieval and content, and build a durable picture of how the AI is performing in the field.
Own the deployment and delivery lifecycle of AI system updates across cloud, on‑premises, and air‑gapped environments: packaging, release management, configuration, and monitoring.
Strengthen the reliability and latency of AI features: profiling inference paths, caching, fallback behavior, and graceful degradation in production.
Collaborate with spatial computing, mobile, and product engineers to connect the knowledge base with 3D scene representations and ship mission‑critical features.
Essential Skills & Experience
4+ years of software engineering experience, with at least 1–2 years building production LLM systems (RAG pipelines, agents, or LLM‑powered products with real users).
Hands‑on experience designing knowledge ingestion and retrieval systems: embedding models, vector and hybrid search, chunking strategies, and semantic data modeling across more than one modality.
Strong Python and backend engineering fundamentals: APIs, data pipelines, async systems, and working in a cloud environment (AWS preferred).
A metrics‑driven approach: experience building evals or benchmarks for LLM systems and using them to drive iteration, not just report scores.
Demonstrated ability to take ambiguous problems from prototype to reliable, maintainable production services.
Experience deploying and operating production systems: CI/CD, release management, and monitoring.
Nice to Have
Experience fine‑tuning open‑weight models (LoRA/PEFT, distillation) or deploying quantized models for offline, edge, or on‑premises environments.
Research background or publications in retrieval, multimodal learning, or agentic systems, or a track record of translating recent research into shipped capabilities.
Voice pipeline experience: streaming STT/TTS, latency optimization for real‑time conversational systems.
Familiarity with 3D or spatial data (scene graphs, point clouds, 3DGS) and grounding language models in spatial representations.
MLOps and inference infrastructure experience: GPU serving, model versioning, cost/latency optimization at scale.
Defense, aerospace, energy or other regulated‑industry experience; active or ability to obtain U.S. security clearance.
Why Join Us?
Competitive salary that reflects your experience and track record
Meaningful equity stake in a high-growth, venture-backed defense tech startup, so you share in the upside you help create
Comprehensive health coverage: medical, dental, and vision insurance
401(k) plan
Paid parental leave
High visibility and real impact: Collaborate with world-class engineers and researchers in a high-ownership environment.
Skills
As published by ashby · 3 questions · 1 written answer
Basics
Name, Email, Resume
Short answers (1)
- LinkedIn URL:
Pick from a list (1)
- Do you require sponsorship to work in the US?
Written answers (1)
- Cover Letter