Engineering Manager, DevOps
Summary
Lead Maven AGI’s DevOps team to design, scale, and secure cloud and on-prem Kubernetes infrastructure for an enterprise AI platform serving autonomous customer support.
Maven AGI is an enterprise AI platform founded in July 2023 by executives from HubSpot, Google, and Stripe. We build conversational AI agents for autonomous customer support at scale. Our platform unifies fragmented systems, integrates knowledge sources, and enables intelligent actions without costly infrastructure changes.
Our team includes talent from Google, Meta, Amazon, Microsoft, and Stripe, with advisors from OpenAI, Google, HubSpot, and Stripe.
The Role
We’re looking for a DevOps Manager to lead and evolve the infrastructure powering Maven AGI’s AI platform. You will manage and scale a high-performing infrastructure team while helping ensure our systems remain reliable, secure, and scalable across cloud and on-premises environments.
This is a technical leadership role that combines people management, operational ownership, and strong infrastructure judgment. You will partner closely with engineering leaders and technical leads to translate business and customer requirements into clear infrastructure priorities and execution plans.
You will also work directly with enterprise customers to understand complex deployment requirements, particularly for private-cloud and on-premises environments, and coordinate stakeholders across the organization to deliver sustainable solutions.
Leadership and Management Responsibilities:
Manage, coach, and develop a team of DevOps and infrastructure engineers
Establish clear expectations, ownership, and accountability across the team
Partner with technical leads to align technical strategy, architecture, and execution
Hire and onboard engineers as the team grows
Lead performance management, career development, and regular feedback
Own team planning, prioritization, capacity management, and delivery
Balance reliability, security, customer commitments, and long-term platform investments
Communicate infrastructure risks, trade-offs, and progress to technical and non-technical stakeholders
Build strong partnerships across Engineering, Product, Security, and Customer Success
Improve operational processes while avoiding unnecessary overhead and reducing team toil
Technical and Operational Responsibilities:
Guide the design, implementation, and operation of cloud and on-premises infrastructure across Azure, AWS, and customer-managed environments
Oversee infrastructure-as-code practices using Pulumi, Bicep, Terraform, or similar tools
Own the reliability and operation of production Kubernetes environments, including deployments, scaling, monitoring, and incident response
Drive the development and improvement of CI/CD pipelines for a large-scale monorepo
Establish consistent observability practices across metrics, logs, traces, and alerting
Advance reliability practices, including SLOs, capacity planning, disaster recovery, and runbook development
Support and scale enterprise AI deployments, including GPU infrastructure, model-serving workloads, and high-concurrency systems
Partner with engineering teams to improve developer experience, platform usability, and deployment velocity
Strengthen secrets management, access controls, and infrastructure security
Evaluate and adopt tools that improve reliability, scalability, and operational efficiency
Participate in incident response and ensure incidents lead to durable improvements
Required Qualifications:
7+ years of professional DevOps/SRE/Infrastructure experience
3+ years of experience managing teams
Deep expertise with Kubernetes in production (AKS, EKS, or GKE)
Strong infrastructure-as-code skills (Pulumi, Terraform, or Bicep)
Experience operating CI/CD systems (GitHub Actions, ArgoCD, or Jenkins)
Proficiency in at least one scripting/programming language (Python, Go, TypeScript, or Bash)
Solid understanding of IaaS providers, networking, DNS, load balancing, and TLS
Experience with monitoring and observability stacks (Datadog, Prometheus, Grafana, or similar)
Experience with multi-cloud or hybrid (cloud + on-prem) deployments
Strong communication and cross-team collaboration skills
Organized, great attention to detail, comfortable operating in a ticketing environment
Thrives in fast-paced startup environments
Nice to have:
Experience with GPU infrastructure and ML/LLM serving workloads (vLLM, TEI)
Familiarity with Temporal or other workflow orchestration systems
Security and compliance background (SOC 2, HIPAA, GDPR)
Experience managing infrastructure costs and capacity at scale
How you show up:
What unites us is our values and the passion we share to live by them:
We are customer champions. You put users at the center of your thinking, advocate for their needs, and design solutions that make their lives measurably better.
We are bold in action. You move with urgency and courage. You’re not afraid to challenge convention, take smart risks, and push boundaries in pursuit of meaningful outcomes.
We are data-driven and insight guided. You make thoughtful decisions grounded in evidence. You’re curious, analytical, and combine data with intuition to guide strategy and execution.
We are stronger together. You bring others along, value diverse perspectives, and contribute to a culture of trust and shared ownership. You believe the best ideas emerge through open dialogue and collective effort.
What We Offer:
High Impact in cutting-edge field. Be at the vanguard of AI innovation.
Competitive salary, comprehensive benefits, and meaningful equity stakes.
A diverse and welcoming work environment where everyone’s voice is heard.
MavenAGI is an equal opportunity employer that values diversity and is committed to fostering an environment where everyone feels included. Join us in changing the face of enterprise customer support.