Point your AI agent at freehire and let it find you a job.

Get the CLI →

ai71 Careers Site

NewBe an early applicant

MLOps Engineer

Posted Updated
Discussion

About AI71:

AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.

The Role:

As an MLOps Engineer you set the ML infrastructure and reliability strategy across AI71's platform, including how LLMs and other deep learning models are deployed, fine-tuned, and served at scale. You own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive multi-quarter ML infrastructure strategy. You are a force multiplier.

What You'll Do:

  • Define ML infrastructure architecture across the platform: model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP)
  • Set direction for ML system reliability: monitoring, latency / throughput / availability targets, and incident response across research and production environments.
  • Mentor senior MLOps engineers; raise the operational bar across multiple teams.
  • Drive cross-team initiatives that improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).
  • Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.
  • Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.

What You'll Bring:

  • 10+ years of MLOps, ML infrastructure, or machine learning engineering with history of architectural ownership.
  • Proven track record architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.
  • Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong Python proficiency
  • Mentorship record — engineers you have grown now operate independently at higher levels.
  • Deep comfort architecting ML systems that run in both managed SaaS and on-premises / disconnected air-gapped environments.
  • Kubernetes at architectural depth — GPU scheduling, multi-tenancy, operators, and the failure modes of distributed workloads on shared clusters.
  • Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams.

Strong Preference:

  • Ownership of production reliability at platform level: SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.
  • Architecture-level experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.
  • Deep GPU systems knowledge: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.
  • Model optimization strategy at portfolio level: quantization (FP8, AWQ, GPTQ), speculative decoding, with measurable cost or latency outcomes across multiple systems.
  • Experience in regulated or security-constrained environments — compliance-driven architecture, model governance, lineage, audit, and secrets management.
  • On-prem / air-gap ML delivery architecture experience at scale.
  • Track record maturing MLOps practice in a growing organization: standards, platform abstractions, and paved paths that outlived your involvement.
  • Bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.

Nice to Have:

  • Conference speaking, technical writing, or industry thought leadership.
  • Open-source contributions to inference, serving, or ML infrastructure projects, particularly maintainer-level involvement.
  • C/C++ or CUDA kernel experience for performance-critical paths.
  • Arabic language skills.

Why AI71:

  • Mission-Driven Work: Work on cutting-edge AI applications with a talented and passionate team, solving real-world challenges in critical sectors.
  • Unparalleled Opportunity: This is a chance to innovate and solve real-world challenges using AI at a company with unique access to world-leading models and resources.
  • Career Growth: We offer competitive compensation, benefits, and significant career growth opportunities as a foundational member of the team.
  • World-Class Environment: Enjoy a flexible working environment and the latest tools & technologies needed to do your best work.

Skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available