Point your AI agent at freehire and let it find you a job.

Get the CLI →

fal

New

Machine Learning Engineer Reliability

Posted Updated 3 views
Discussion

Summary

Own the reliability, security, and safety of fal's production generative media model APIs: building observability, safe deployments (canary, rollback), incident response, and GPU capacity management. Core stack includes Python, PyTorch, Diffusers, and Kubernetes.

You will own the reliability, security, and safety of production generative media model APIs. You will build observability and safe deployment systems, lead incident response, improve GPU capacity management, and help incorporate reliability requirements into new model onboarding.

Responsibilities

  • Own availability, latency, and throughput SLOs for generative media model APIs
  • Build monitoring, alerting, and observability for ML-specific failures and model regressions
  • Harden deployments with canary releases, shadow testing, automated rollbacks, and validation gates
  • Drive secure model serving, abuse detection, rate limiting, and adversarial-use protection
  • Operationalize content moderation, safety classifiers, and inference-time guardrails
  • Lead incident response, postmortems, and recurrence-prevention work
  • Improve capacity planning, autoscaling, and GPU fleet efficiency
  • Incorporate reliability, security, and safety requirements into model onboarding

Requirements

  • Production ML or high-scale API operations
  • Diffusion models
  • Distributed systems
  • Networking
  • Observability
  • Incident management
  • Generative models
  • Security
  • ML safety
  • Python
  • PyTorch
  • Diffusers
  • Kubernetes

Benefits

  • Equity
  • Regular team events and offsites

Skills

Apply

See also

ML / AI jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available