Senior AI Engineer
Summary
Build and ship AI-powered features using LLMs, focusing on reliability and integration; work with Linux VMs, NVIDIA GPUs, and local model serving like vLLM.
We are looking for a pragmatic, end-to-end AI Engineer to build and ship AI-powered product features using pre-trained foundation models. In this role, your focus will be on reliability, measurable quality, and seamless integration.
Beyond core AI application development, you possess an intermediate ability to navigate Linux environments and manage GPU infrastructure. You should be comfortable operating on Linux VMs with NVIDIA GPUs, running local models (e.g., vLLM), and troubleshooting the ML stack. We are not looking for a pure DevOps or systems specialist, but rather a highly resourceful engineer who can confidently navigate the infra layer to get the job done.
Key Responsibilities
Product & Feature Delivery
End-to-End Ownership: Ship AI features from initial design and prototyping to production deployment and continuous iteration.
- Robust Integrations: Build reliable model integrations utilizing structured outputs, rigorous validation, and deterministic fallbacks.
- Advanced AI Patterns: Develop RAG, agentic workflows, and multimodal features embedded with explicit guardrails.
- Data Pipelines: Design RAG and retrieval pipelines with sensible chunking, strategic indexing, and strict access control.
Quality, Evals & Operations
- Root-Cause Debugging: Diagnose and resolve quality issues at the source (retrieval, data pipeline, prompt engineering) rather than relying on blunt-force retries.
- Production Instrumentation: Implement evaluation suites and production monitoring to track quality, latency, and cost trade-offs.
- On-Call Support: Instrument production behavior for owned domains and participate in the team's on-call rotation.
️ Focus Area: Linux & GPU Infrastructure
- Environment Management: Work comfortably within Ubuntu VMs and manage Python virtual environments cleanly.
- Local Model Serving: Run local/self-hosted LLMs (e.g., vLLM) and spin up containerized workloads.
- Stack Troubleshooting: Debug and resolve issues across the ML stack, including PyTorch/TensorFlow, CUDA/cuDNN, and underlying driver conflicts.
- NVIDIA Tooling: Utilize NVIDIA GPU tooling efficiently (nvidia-smi, CUDA toolkit, drivers, and NVIDIA Container Toolkit).
Skills & Core Competencies
- Tech Stack: Proficient in Python and JavaScript/TypeScript with solid core software engineering fundamentals.
- LLM Architectures: Strong grasp of modern LLM patterns, including context window construction, RAG, tool use (function calling), and safety basics.
- Data-Driven Mindset: Evaluation-driven engineer who reasons clearly about the trade-offs between performance, cost, and latency.
- Soft Skills: Clear communicator who thrives in collaborative environments and embraces a "disagree and commit" philosophy.
Requirements: Minimum Qualifications
- Experience: 2–3 years of professional software engineering experience shipping production-grade features.
- AI/LLM Expertise: Hands-on experience implementing LLM application patterns (prompting, structured outputs, RAG, and tool calling).
- Ops Foundations: Experience with evaluation frameworks and basic production operations (logging, metrics, incident response).
- Linux Comfort: Solid familiarity with Linux/Ubuntu environments.
- Infra Exposure: Basic exposure to local model serving, GPU workloads, or ML-stack troubleshooting—or explicit evidence of your ability to learn these systems rapidly.
- Education: Bachelor’s degree in a technical field (Computer Science, Data Science, Engineering) or equivalent practical experience.
Preferred Qualifications (Nice-to-Haves)
- Prior experience managing GPU infrastructure or deep ML-ops specialization.
- Familiarity with the NVIDIA AI Enterprise ecosystem.
- Relevant Cloud, Generative AI provider, or NVIDIA certifications.