Sr. ML Engineer (MLOps)
Summary
Build and deploy MLOps tooling and services for Lyra’s AI-powered mental-health platform, including training/inference pipelines, APIs, and evaluation frameworks in Python and Kubernetes.
About the Role:
We are looking for an experienced Senior ML Engineer who is eager to build the tooling and services necessary to deliver ML and generative AI based products that make a significant impact within the organization. The ideal candidate will be enthusiastic about taking ownership of their work, spearheading cross-functional projects, and providing guidance and mentorship to other team members.
Lyra is for you if you:
-
Thrive on working with brilliant teammates to solve complex, meaningful problems
-
Are passionate about making a social impact and supporting people at their most challenging moments
-
Enjoy cross-functional collaboration with physicians, therapists, data scientists, data analysts and product managers
Responsibilities
-
Build tooling and services that power the machine learning and generative AI powered solutions in production
-
Training and inference platform services
-
The model evaluation metrics and testing framework
-
Build services that expose machine learning and AI based products
-
Deploy and manage various applications that power various components of the the AI/ML SDLC process
-
And of course, you will be coding every day!
Qualifications
-
6+ years of experience deploying ML/AI solutions in production environments
-
Ability to write high-quality code in Python
-
Experience building RESTful APIs
-
Experience defining and using Protobuf messages
-
Experience working with Docker and deploying applications to Kubernetes
-
Experience with relational and low-latency databases
-
Experience working with Celery
-
A strong desire to work on ML/AI based products
-
A desire to learn new technologies quickly
-
A love of building systems from scratch
-
A thoughtful approach to balancing quality and deadlines in fast-paced settings
-
Excellent communication skills with a talent for building consensus and alignment
-
Strong organizational skills and the ability to distill complex problems into clear priorities that move the team and business forward
Preferred Qualifications
-
Experience building RAG (retrieval-augmented generation) based solutions
-
Experience setting up and maintaining vector databases
-
Experience writing production code in Java/Kotlin
-
Experience building solutions on cloud infrastructure, particularly AWS
-
Experience working with highly sensitive data in a healthcare environment
Skills
As published by lever · 9 questions · 4 written answers
Basics
Resume/CV, Full name, Pronouns, Email, Phone, Current location, Current company, LinkedIn URL, Twitter URL, GitHub URL, Portfolio URL, Other website, To which gender identity do you most identify?, To which sexual orientation do you most identify?, Do you identify as a person living with a disability?, Are you fluent in any of the following languages?, Do you identify as LGBTQIA+?
Pick from a list (5)
- Are you legally authorized to work in the United States for our Company?
- Do you now, or will you in the future, require sponsorship for employment visa status (e.g., H-1B visa status, etc.) to work legally for our Company in the United States?
- Are you an employee of Lyra, Bend or an affiliated company (Currently or Previously)?
- If yes to the previous question, please select your affiliation: optional
- If you are a Lyra employee, do you certify you have reviewed the Internal Mobility Policy and meet the requirements? optional
Written answers (4)
- If you selected "other", please share your relationship here optional
- What specific strategies (such as quantization, distillation, or guardrailing) have you implemented to optimize model latency, reduce costs, and mitigate bias or safety concerns in production?
- Describe a production RAG system or GenAI solution (e.g., LLMs, fine-tuning, or vector retrieval) you designed and deployed.
- Which vector database did you use, and how did you approach indexing, retrieval quality, and latency optimization?