freehire launches on Product Hunt on 26 August.

Follow →

Member of Technical Staff

Open 16d

Summary

Optimize Morph’s inference stack—kernels, serving, routing, and autoscaling—to maximize speed, cost-efficiency, and reliability for open AI models.

Morph was a 1 person company from 0 → 10M of revenue. We will be the first 10 person $10b company.

Every employee should contribute >30M of revenue/yr to the company.

The best candidates would be top 1% at multiple parts of the inference stack, yet have breadth across the whole stack.

Morph builds high margin inference infrastructure. Our stack spans kernels, model serving, routing, autoscaling, and capacity.

What you’ll do

  • Find the gap between theoretical hardware performance and production performance
  • Trace latency and throughput regressions from the API layer down to individual kernels
  • Optimize batching, scheduling, routing, quantization, and distributed execution
  • Work on new research directions around caching
  • Work with NVLink and RoCE
  • Validate that every optimization preserves model quality and correctness

You might be a fit if you

  • Have optimized complex production systems
  • Can juggle 8+ Codex/Claude/other coding agents concurrently
  • Understand GPU performance, memory bandwidth, collectives, and inference serving
  • Are strong in Python, CuTEdsl, and comfortable navigating unfamiliar codebases
  • Care about tokens per second, tokens per dollar, and correctness equally

You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served. Small team, enormous compute, immediate production impact.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available