Manager, AI Operations
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Manager, AI Operations based in the United States.
This is a high-impact leadership opportunity focused on making AI systems reliable, observable, secure, and scalable in production.
You will oversee operations across traditional machine learning, Generative AI, and emerging Agentic AI systems.
The role sits at the intersection of AI, infrastructure, security, governance, and data science, translating advanced models into dependable production capabilities.
You will lead a specialized team responsible for production support, monitoring, incident response, and operational readiness.
A major focus will be developing measurable and auditable practices that support ethical, responsible, and cost-efficient AI operations.
You will work in a fast-moving healthcare technology environment where operational excellence directly contributes to more efficient healthcare delivery.
This role offers the opportunity to shape the operational foundations of a growing AI portfolio while influencing how AI is deployed responsibly at scale.
Accountabilities:
- Own the deployment support, day-to-day reliability, monitoring, and production operations of AI systems spanning traditional ML, Generative AI, and Agentic AI.
- Lead incident response, troubleshooting, root-cause analysis, and remediation for AI/ML services, including systems capable of autonomous or tool-based actions.
- Partner with MLOps, Cloud Engineering, Infrastructure, Security, and IT teams to establish robust deployment pipelines and production-ready operating practices.
- Build and maintain comprehensive AI observability capabilities, including model drift detection, latency, uptime, performance degradation, logging, alerting, and agent-level tracing.
- Define production health metrics and dashboards that provide real-time visibility into the performance, reliability, and operational risks of AI systems.
- Establish processes for identifying performance issues and coordinating remediation with the appropriate AI development and data science teams.
- Serve as the operational bridge between AI development teams and enterprise security, infrastructure, and technology functions, ensuring appropriate access controls and production risk management.
- Support AI governance activities by implementing and operating measurable instrumentation for fairness, safety, bias, and other responsible-AI requirements defined by the governance function.
- Coordinate the transition of AI solutions from development into production, including structured production-readiness reviews and operational handoffs.
- Lead, develop, and mentor a small, high-leverage AI production support and observability team while establishing a strong operational and on-call culture.
- Set priorities, manage capacity, and build scalable operational processes as the AI portfolio grows.
- Lead cross-functional initiatives to establish ethical AI best practices and standard operating procedures across the design, development, and runtime layers of the AI stack.
- Establish clear visibility into AI development and runtime costs and drive targeted year-over-year efficiency improvements.
- Partner with Governance and Compliance teams to ensure AI standards are measurable, continuously monitored, and auditable through the observability platform.
- 7+ years of experience in AI/ML operations, MLOps, production data science, or a closely related discipline, including at least 2 years of people or team leadership experience.
- Hands-on experience operating traditional machine learning and Generative AI systems in production, including deployment, monitoring, troubleshooting, and incident response.
- Strong understanding of AI/ML observability practices, including model drift, performance monitoring, logging, alerting, uptime, latency, and production health metrics.
- Experience partnering with Security and Infrastructure teams on production risk, access controls, operational readiness, and secure deployment practices.
- Strong communication and stakeholder-management skills, with the ability to translate complex technical and operational risks into clear business terms.
- Demonstrated ability to lead teams, prioritize competing initiatives, manage capacity, and establish effective operational processes.
- Experience working across technical and business functions in a fast-moving, collaborative environment.
- Prior experience in healthcare, health technology, payer, or provider organizations is preferred.
- Experience operating Agentic AI systems, including tool-use monitoring, guardrails, and action-level tracing, is a strong advantage.
- Familiarity with AI/ML observability and cloud technologies such as Datadog, AWS CloudWatch, Langfuse, and native AWS capabilities is preferred.
- Experience working alongside a dedicated AI Governance function, with a clear understanding of the distinction between operational monitoring and governance policy, is desirable.
- Strong interest in responsible and ethical AI, operational risk management, cost optimization, and building scalable AI practices.
- Comprehensive medical, dental, and vision insurance.
- 401(k) retirement plan with company matching.
- Flexible paid time off.
- Paid parental leave.
- HSA and FSA options.
- Educational reimbursement program.
- Employer-paid Employee Assistance Program and mental health services.
- Professional development opportunities in a growing AI and healthcare technology environment.
- Collaborative, fast-moving culture with opportunities to make a measurable impact on AI operations and healthcare efficiency.
- Travel to the company’s Tennessee offices for training as required.
- U.S. work authorization is required; visa sponsorship or immigration support is not provided for this position.