Point your AI agent at freehire and let it find you a job.

Get the CLI →

Cerebras Systems, Inc.

NewBe an early applicant

Software Engineer GPU Inference

Posted 2 views
Discussion

Summary

Builds, deploys, and operates the GPU prefill path for Cerebras's AI inference service, working across API services, vLLM, PyTorch, ROCm, GPU nodes, and rack-scale infrastructure. Day to day, the engineer improves reliability, latency, throughput, and capacity through debugging, benchmarking, automation, and release validation.

You will build, deploy, and operate the GPU prefill path across API services, serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure. You will improve reliability, numerical correctness, observability, latency, throughput, and capacity efficiency through debugging, benchmarking, automation, and release validation.

Responsibilities

  • Design, build, deploy, and maintain the GPU prefill path
  • Establish deployment, upgrade, rollback, health-checking, capacity-management, and recovery practices
  • Define service-level indicators and objectives for GPU-backed inference
  • Profile and optimize inference latency, throughput, utilization, memory efficiency, and capacity
  • Tune model-serving scheduling, batching, caching, parallelism, admission, quantization, and graph execution
  • Diagnose failures and regressions across application, runtime, distributed-system, and hardware layers
  • Build validation infrastructure for model quality, numerical accuracy, determinism, and compatibility
  • Develop benchmarks, workload replay tools, profiling automation, dashboards, and regression gates

Requirements

  • 5+ years of software engineering experience
  • Production inference systems experience for large language models, multimodal models, or demanding GPU workloads
  • C++
  • Python
  • Multithreading
  • Concurrency
  • Memory management
  • vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent serving framework
  • GPU execution and performance optimization
  • Distributed-system debugging
  • Linux
  • Containerization
  • Kubernetes or comparable orchestration
  • Observability
  • CI/CD
  • Benchmarking
  • Technical leadership
  • Computer Science, Computer Engineering, Electrical Engineering, related degree, or equivalent practical experience

Skills

Apply

See also

Software Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available