Point your AI agent at freehire and let it find you a job.

Get the CLI →

Sakana AI K.K.

New

Software Engineer Product Infrastructure and Platform Reliability

Posted 2 views
Discussion

You will own the reliability and SLA management of production LLM products and their inference platform. You will operate cloud and GPU infrastructure, improve latency, throughput, utilization, and cost efficiency, and support deployment and rollback. You will build observability and incident-response practices, participate in on-call work, plan capacity, and translate enterprise requirements into infrastructure design.

Responsibilities

  • Manage reliability and SLAs across products
  • Operate and improve production LLM inference infrastructure
  • Improve latency, throughput, GPU utilization, and cost efficiency
  • Operate LLM serving systems and support deployment and rollback workflows
  • Build monitoring, alerting, incident response, postmortems, and prevention practices
  • Participate in the on-call rotation
  • Translate enterprise security, availability, compliance, and SLA needs into infrastructure design
  • Support capacity planning and cost management

Requirements

  • Experience designing and operating production cloud infrastructure, such as AWS or GCP
  • Experience operating containerized services, APIs, batch jobs, or model inference workloads in production
  • Experience managing and automating infrastructure with IaC or comparable tooling
  • Experience with monitoring, alerting, incident response, and postmortems
  • Ability to assess GPU, cloud, and serving resources for cost, performance, and availability
  • Experience with vLLM, TensorRT-LLM, or SGLang
  • English communication skills for technical projects and stakeholder discussions

Skills

Apply

See also

Software Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available