freehire launches on Product Hunt on 26 August.

Follow →

Senior ML Engineer

Summary

Build, fine-tune, and deploy deep-learning models end-to-end: train Transformers, optimise inference with NVIDIA’s stack, and package models into production-ready containers.


We are looking for a hands-on Machine Learning Engineer to own the training, fine-tuning, and continual improvement of the models that power our metrics. You will build and curate datasets, optimise encoder and decoder models, and run systematic experiments across architectures, batch sizes, and training regimes to improve accuracy, efficiency, and cost-performance trade-offs.

You will also own the end-to-end deployment pipeline, packaging trained models into production-ready images running on NVIDIA’s inference stack and maintaining the Rust services that serve them in production. This is a highly technical role with direct ownership from dataset construction and experimentation through to production deployment.

  • Model training and fine-tuning (top priority): Train, fine-tune, and improve encoder
    and decoder models for our metrics. Own the full loop from data to evaluated,
    production-ready checkpoints.
  • Dataset and metric development: Design, build, and curate datasets for new and
    existing metrics. Define labelling schemes, manage data quality, and connect dataset
    changes to measurable model improvements.
  • Experimentation and evaluation: Run systematic experiments across model
    architectures, batch sizes, precision, and training hyperparameters. Build and maintain
    rigorous evaluation harnesses and track results to find the best accuracy/cost trade-offs.
  • Inference optimisation: Optimise trained models for production inference on the Nvidia
    stack (Triton, TensorRT / TensorRT-LLM, ONNX) using methods such as quantization,
    distillation, and precision tuning.
  • Model deployment pipeline: Maintain and improve the pipeline (Docker-based) that
    converts models to suitable formats (TensorRT, ONNX), fixes vulnerabilities, and
    containerises them for production.
  • Rust application layer (supporting): Maintain and extend the Rust services and API
    endpoints that serve our models in production.

Must-Have Skills

  • 3+ years of practical experience training and fine-tuning deep learning models, with a proven
    track record of taking models from data to production
  • PyTorch, and a solid understanding of the inner workings of Transformers and deep
    learning more broadly
  • Dataset construction and model evaluation: building datasets, defining metrics, and
    designing evaluation methodology
  • Hands-on Python programming experience
  • Nvidia Inference Stack: Triton Inference Server, TensorRT / TensorRT-LLM, ONNX
  • Docker and Kubernetes (k9s familiarity is a plus)

Nice-to-Have

  • GPU model training and experiment tracking (e.g. Weights & Biases, MLflow)
  • Inference optimisation techniques (e.g. quantization)
  • Backend development in Rust, including building and maintaining production API
    services

Soft Skills

  • Builder mindset: thrives on writing, debugging, and improving production code.
  • Collaborative, humble, and open to feedback.
  • Strong communicator who explains design decisions clearly.
  • Influences through contribution, not hierarchy.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available