freehire launches on Product Hunt on 26 August.

Follow →

Software Engineer, Infrastructure & Platform

Summary

Build infrastructure for AI evaluations, including sandboxed environments, scalable platforms, and orchestration systems using Python, Docker, Kubernetes, AWS/GCP, and Terraform. Focus on reliability, security, and reproducibility for AI model testing and research.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Software Engineer, Infrastructure & Platform based in the United States.

This is a hands-on engineering role focused on building the infrastructure behind advanced AI evaluations and agentic systems.
You’ll design secure, reproducible environments where AI models can execute code, use tools, and complete complex multi-step tasks.
The role sits at the intersection of backend engineering, cloud infrastructure, distributed systems, and AI safety.
You’ll build scalable evaluation platforms, orchestration systems, APIs, and tooling used by researchers, engineers, and security specialists.
A key part of the role is ensuring evaluations are reliable, observable, secure, and reproducible at significant scale.
You’ll work alongside analysts, red teamers, and domain experts to turn complex research concepts into robust technical systems.
This is a fully remote U.S.-based opportunity suited to an engineer who enjoys ambiguity, deep debugging, and technically challenging problems.

Accountabilities:

  • Design and build sandboxed evaluation environments that allow AI models to safely execute code, interact with tools and services, and perform complex tasks.
  • Develop backend services and infrastructure that support large-scale, repeatable AI and agentic evaluations.
  • Build agent scaffolding and evaluation harnesses covering tool-use loops, context management, retries, state management, token budgets, and multi-agent or subagent workflows.
  • Provision and orchestrate isolated environments using technologies such as Docker, Kubernetes, virtual machines, and cloud infrastructure.
  • Design secure approaches to networking, permissions, credentials, secrets management, and resource isolation for model-driven environments.
  • Develop APIs, internal tools, and automation that enable researchers, engineers, and subject-matter experts to efficiently create and execute evaluations.
  • Improve evaluation reliability and reproducibility through logging, observability, snapshotting, debugging capabilities, and automated testing.
  • Build infrastructure capable of running thousands of evaluation tasks reliably while capturing the artifacts and telemetry required to analyze model behavior.
  • Partner with analysts, red teamers, and technical experts to translate sophisticated evaluation concepts into dependable engineering systems.
  • Investigate failures across application, infrastructure, networking, and evaluation layers, distinguishing model limitations from problems with the underlying environment or harness.
  • Continuously improve platform scalability, security, resilience, and developer experience as evaluation requirements evolve.
  • Requirements

    • 3–5+ years of professional software engineering experience, particularly in backend, infrastructure, platform, SRE, or distributed systems engineering.
    • Strong programming skills in Python and experience developing production-quality software.
    • Proven experience designing and operating backend services, APIs, or distributed systems.
    • Hands-on experience with Docker, Kubernetes, virtual machines, or comparable container and orchestration technologies.
    • Experience working with AWS, GCP, or similar cloud infrastructure platforms.
    • Strong understanding of Linux systems, networking, authentication, permissions, and infrastructure security.
    • Experience with Infrastructure as Code and automation tools such as Terraform.
    • Excellent debugging and troubleshooting abilities across application, infrastructure, and networking layers, including complex agentic workflows.
    • Ability to build systems that are reproducible, observable, scalable, reliable, and secure.
    • Comfort working through ambiguous technical challenges where requirements and architecture may change rapidly.
    • Strong collaboration and communication skills, with the ability to work effectively with researchers, engineers, analysts, and technical specialists.
    • Genuine interest in AI systems, agentic workflows, AI security, or model evaluations; previous professional AI experience is beneficial but not mandatory.
    • Experience with developer platforms, CI/CD systems, test infrastructure, sandboxes, or ephemeral compute environments is a plus.
    • Familiarity with LLM APIs, agent frameworks, tool-calling systems, or AI evaluation infrastructure is advantageous.
    • Experience designing secure execution environments for untrusted or semi-trusted code is highly desirable.
    • Background in SRE, platform engineering, cloud infrastructure, cybersecurity, or developer tooling is a strong plus.
    • Knowledge of distributed task execution, queues, workflow orchestration, or large-scale automated testing is beneficial.
    • Familiarity with AI safety, adversarial testing, model evaluations, autonomous agents, or agentic AI concepts—including Model Context Protocol, agent benchmarks, and AI-agent security risks—is advantageous.
    • Candidates must be based in the United States and able to meet applicable work authorization requirements.
    • Benefits

      • Competitive salary: $110,000–$160,000 annually, depending on experience and location.
      • Performance bonus: Annual performance-based bonus opportunity.
      • Fully remote: Work remotely from anywhere in the United States.
      • Health coverage: Comprehensive health, dental, and vision benefits.
      • Paid time off: Generous PTO and paid holiday schedule.
      • Retirement: 401(k) plan.
      • Professional development: Support for conferences, continuing education, and leadership training.
      • Impactful work: Build infrastructure supporting advanced AI evaluation, safety, security, and research initiatives.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

What this application asks

lever

Resume/CV, Full name, Email, Phone, Current location, Current company

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available