Point your AI agent at freehire and let it find you a job.

Get the CLI →

Cerebras Systems, Inc.

NewBe an early applicant

Full Stack LLM Engineer

Posted 4 views
Discussion

Summary

Bring up machine-learning/LLM models on Cerebras CSX wafer-scale systems: translating model architectures, lowering graphs, optimizing compilers and runtimes, and tuning performance. Core stack is Python, C/C++, PyTorch/TensorFlow, and compiler tooling like LLVM/MLIR.

You will bring up machine-learning models on CSX systems and work across model translation, graph lowering, compiler optimization, runtime integration, and performance tuning. You will debug performance and correctness issues and prototype improvements to tools, APIs, and automation flows.

Responsibilities

  • Bring up machine-learning models on CSX systems
  • Translate model architectures and lower graphs
  • Optimize compilers and tune performance
  • Integrate models with runtimes
  • Debug performance and correctness issues across model code, compiler IRs, runtimes, and hardware utilization
  • Prototype improvements to tools, APIs, and automation flows

Requirements

  • Python modeling experience
  • Compiler IR knowledge
  • Performance profiling experience
  • Debugging skills for performance, numerical accuracy, and runtime integration
  • PyTorch or TensorFlow experience
  • Knowledge of attention, mixture-of-experts, or diffusion models
  • C and C++ proficiency
  • Low-level optimization experience
  • LLVM or MLIR compiler development experience
  • Optimization techniques involving NP-hard problems

Skills

Apply

See also

Full-Stack jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available