freehire launches on Product Hunt on 26 August.

Follow →

Software Engineer, Machine Learning Infrastructure

Summary

Build and maintain the infrastructure that powers AI and LLM models at Stripe, from training to production deployment, enabling engineers to move models from research to live systems.

Software Engineer, Machine Learning Infrastructure at Stripe.

About the role Join the team responsible for the foundational systems that power machine learning across Stripe. You will build the infrastructure that supports the entire lifecycle of AI models, from initial data exploration to production deployment and LLM integration.

Key facts

Location: Toronto, Canada

Engagement: Full-time

Team: Machine Learning Infrastructure

What you'll do

  • Architect and maintain secure, reliable services for model training, experimentation, and LLM applications across global regions.
  • Develop libraries and internal tools that help engineers move models from research environments into production.
  • Collaborate with product and data science departments to increase developer velocity.
  • Manage technical projects that span multiple systems and operational requirements.

Requirements

  • 2+ years of professional software engineering experience focusing on distributed systems and service oriented architecture.
  • Full lifecycle development experience, including user requirements, design, implementation, testing, and production operations.
  • Practical experience with MLOps, production ML platforms, or LLM application development.
  • Background in managing high availability, low latency production systems.
  • Proven ability to partner with cross-functional teams to achieve business goals.
  • Ability to balance technical idealism with pragmatic execution.

Nice to have

  • Experience developing and deploying production AI agents.
  • Familiarity with LLM frameworks and large language models.
  • History of training and shipping ML models to address specific business challenges.

Skills & tools

  • Distributed systems
  • Service oriented architecture
  • MLOps
  • LLM frameworks
  • Machine learning model training and serving
  • High availability systems design

Practical notes

You will coordinate with teams based across the US and Canada.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available