MLOps Engineer
About AI71:
AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.
The Role:
As an MLOps Engineer you set the ML infrastructure and reliability strategy across AI71's platform, including how LLMs and other deep learning models are deployed, fine-tuned, and served at scale. You own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive multi-quarter ML infrastructure strategy. You are a force multiplier.
What You'll Do:
- Define ML infrastructure architecture across the platform: model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP)
- Set direction for ML system reliability: monitoring, latency / throughput / availability targets, and incident response across research and production environments.
- Mentor senior MLOps engineers; raise the operational bar across multiple teams.
- Drive cross-team initiatives that improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).
- Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.
- Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.
What You'll Bring:
- 10+ years of MLOps, ML infrastructure, or machine learning engineering with history of architectural ownership.
- Proven track record architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.
- Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong Python proficiency
- Mentorship record — engineers you have grown now operate independently at higher levels.
- Deep comfort architecting ML systems that run in both managed SaaS and on-premises / disconnected air-gapped environments.
- Kubernetes at architectural depth — GPU scheduling, multi-tenancy, operators, and the failure modes of distributed workloads on shared clusters.
- Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams.
Strong Preference:
- Ownership of production reliability at platform level: SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.
- Architecture-level experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.
- Deep GPU systems knowledge: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.
- Model optimization strategy at portfolio level: quantization (FP8, AWQ, GPTQ), speculative decoding, with measurable cost or latency outcomes across multiple systems.
- Experience in regulated or security-constrained environments — compliance-driven architecture, model governance, lineage, audit, and secrets management.
- On-prem / air-gap ML delivery architecture experience at scale.
- Track record maturing MLOps practice in a growing organization: standards, platform abstractions, and paved paths that outlived your involvement.
- Bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.
Nice to Have:
- Conference speaking, technical writing, or industry thought leadership.
- Open-source contributions to inference, serving, or ML infrastructure projects, particularly maintainer-level involvement.
- C/C++ or CUDA kernel experience for performance-critical paths.
- Arabic language skills.
Why AI71:
- Mission-Driven Work: Work on cutting-edge AI applications with a talented and passionate team, solving real-world challenges in critical sectors.
- Unparalleled Opportunity: This is a chance to innovate and solve real-world challenges using AI at a company with unique access to world-leading models and resources.
- Career Growth: We offer competitive compensation, benefits, and significant career growth opportunities as a foundational member of the team.
- World-Class Environment: Enjoy a flexible working environment and the latest tools & technologies needed to do your best work.
Skills
As published by greenhouse · 3 questions
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter
- Preferred First Name optional
- LinkedIn Profile optional
- Website optional