Principal AI Engineer
Summary
Designs and builds scalable, secure AI/ML platforms for healthcare, focusing on agentic systems, RAG pipelines, and multi-modal AI applications with guardrails and observability. Leads architectural standards, cloud-native infrastructure (Azure/AWS), and AI governance for enterprise-grade solutions.
- Design and document enterprise AI platform architectures including reference implementations for agentic systems, RAG pipelines, and multi-modal AI applications with integrated guardrails, observability, and security patterns
- Define and maintain architectural standards, design patterns, and best practices for GenAI infrastructure including model serving, prompt management, vector storage, evaluation frameworks, and LLMOps pipelines.
- Lead technical evaluations and vendor assessments for AI infrastructure components (model gateways, vector databases, observability tools) and provide architectural recommendations aligned with organizational requirements.
- Collaborate with product, engineering, and platform teams to translate business requirements into scalable AI architectural solutions, ensuring consistency across multiple product implementations.
- Establish architectural governance for AI/ML workloads including security controls, compliance frameworks, cost optimization strategies, and multi-cloud deployment patterns (Azure, AWS).
- Experience with Azure OpenAI Service, Azure AI Studio, and Azure Machine Learning platforms in healthcare or regulated industries
- Familiarity with observability frameworks for AI systems (OpenTelemetry, MLFlow, Arize, LangSmith) and production monitoring strategies
- Understanding of healthcare compliance requirements (HIPAA, PHIPA) and security frameworks for AI applications
- Experience with Infrastructure as Code (Terraform, Bicep) and GitOps practices for AI platform automation
- 8+ years of experience in cloud architecture and platform engineering with AWS and/or Azure, with at least 3+ years focused on AI/ML infrastructure and GenAI solutions
- Proven track record designing and implementing enterprise-scale AI/ML platforms supporting multiple product teams and use cases
- Deep expertise in cloud-native architectures including microservices, event-driven systems, serverless patterns, and container orchestration (Kubernetes)
- Strong understanding of GenAI architectural patterns including RAG, agentic frameworks (LangGraph, CrewAI), prompt engineering, and LLM evaluation methodologies
- Experience with AI infrastructure components such as vector databases (Pinecone, Weaviate, pgvector), model serving platforms (vLLM, SGLang, Azure AI), and prompt management systems
Skills
- Agentic AI
- AI
- Automation
- AWS
- Azure
- Bicep
- Cloud
- Cloud Native
- CrewAI
- Design Patterns
- Event Driven Architecture
- Generative AI
- GitOps
- Hipaa
- Infrastructure as Code
- Kubernetes
- LangGraph
- LangSmith
- LLM
- LLMOps
- Machine Learning
- Microservices
- MLflow
- Observability
- OpenAI
- OpenTelemetry
- pgvector
- Pinecone
- Prompt Engineering
- RAG
- Serverless
- Terraform
- Vector Databases
- vLLM
- Weaviate
As published by lever · 11 questions
Basics
Resume/CV, Full name, Pronouns, Email, Phone, Current location, Current company, LinkedIn URL, Portfolio URL, Other website, Gender Identity, Do you identify as a member of the LGBTQ+ Community?, Race/Ethnicity, Veteran Status, Disability
Short answers (1)
- How did you hear about us?
Pick from a list (10)
- Are you legally authorized to work in Canada for our company?
- Would you require future sponsorship in order to work for our company in Canada?
- Have you led the design and production deployment of agentic AI systems used by external customers or enterprise users at scale?
- How many years of hands-on experience do you have building and deploying production AI/GenAI solutions? This is a Principal-level role and not intended for candidates whose experience is primarily research or proof-of-concept work
- Have you personally built and deployed solutions using at least two of the following Microsoft AI platforms?
- Do you have hands-on production development experience with Python and either Java or another enterprise-scale backend language?
- Have you designed and implemented production Retrieval-Augmented Generation (RAG) solutions including embeddings, vector databases, chunking strategies, and evaluation frameworks?
- Have you built production systems utilizing one or more of the following agent frameworks?
- Have you established technical standards, architecture patterns, or best practices that were adopted by multiple engineering teams?
- Have you built or operated systems that route requests across multiple AI models (e.g., GPT, Claude, Gemini, open-source models) based on cost, latency, or quality requirements?