Senior Software Engineer (ML Evaluations)
Senior Software Engineer, ML Evaluations
Artificial Agency
Location: Edmonton, Alberta (Remote possible for exceptional candidates)
Job Description
Artificial Agency, the leading player-facing agentic AI company for games, is seeking a Senior Software Engineer to own the platform our researchers and engineers use to run experiments and evaluations at scale. This is a hybrid role based in Edmonton, Alberta; remote is possible for exceptional candidates.
Reporting to the Head of Platform Engineering, the successful candidate will own the platform end to end, working directly with the teams who run their experiments on it. Every evaluation and benchmark we run goes through it, fanning out to hundreds of jobs across Linux and GPU Windows workers. You’ll build in Python (CLI, job-queueing service, Postgres data model), and own the infrastructure the jobs run on, with help from our infrastructure team.
Expect a mix of distributed-systems design, cloud infrastructure, and hands-on work with ML engineers: from troubleshooting runs that stall or fail to turning stakeholder requests into new platform features. You will make results traceable and useful—helping teams compare models, agent configurations, and software changes, investigate regressions, and understand the evidence behind each result.
Responsibilities
Own our ML evaluation platform end to end: the Python CLI and library, queueing service and API, database schema, and the web UI researchers use to manage runs.
Set technical direction and decide what the platform should and shouldn’t do.
Design and operate the job execution layer: queueing, worker leasing and recovery, concurrency limits, cancellation, timeouts, retries, artifacts, and logs. Keep runs of hundreds of jobs correct, observable, and on time.
Make evaluation results traceable and trustworthy. Record the relevant test versions, game builds, models, agent configurations, and run settings; preserve execution history; and distinguish valid evaluation outcomes from infrastructure failures, invalid tests, and incomplete runs.
Extend the data model, APIs, and web UI to support baseline comparisons, regression investigation, and access to the logs and artifacts behind reported results. Work with ML, agents, and test engineering to implement agreed evaluation and reporting methods, including repeated trials and result aggregation.
Working with our infrastructure team, own the machines the jobs run on: an auto-scaling GPU Windows fleet in AWS, containerized Linux workers on Kubernetes, and a few fixed in-office machines. This covers Terraform, machine images, GPU drivers, bootstrap scripts, secrets handling, and cost/capacity decisions.
Partner with our Senior Test Engineer and other evaluation contributors on clear interfaces for workload configuration, execution requirements, structured results, and diagnostic artifacts. Make evaluations straightforward to run through developer tools and CI workflows.
Work directly with ML, agents, game engineering, and research teams. Diagnose failed and stalled runs, unblock people quickly, and turn one-off requests into features that serve everyone.
Prioritize across competing teams and deadlines, delivering on stakeholder requests while ensuring long-term platform health. Work with the Head of Platform Engineering to resolve conflicting business priorities.
Keep runbooks accurate and actionable for other engineers to respond to incidents or outages. Maintain documentation and examples that help teams launch evaluations and integrate new workloads without routine manual intervention.
Game-specific harnesses, scenarios, and test coverage are owned by the Senior Test Engineer and integration teams, with evaluation criteria and methodology developed in partnership with ML, agents, and AI engineering. You will work together on the interfaces and reporting that connect these workloads to the platform. Evaluation worker infrastructure is part of this role, with infrastructure-team support; deploying and operating production inference services across clouds is a separate responsibility within Platform.
Requirements
An AI-first approach to engineering, with hands-on experience using AI tools for development, testing, debugging, or analysis. You are excited to make AI-driven workflows central to how you work, actively experiment with new capabilities, and adapt your approach as tools improve, while taking responsibility for the quality of what you deliver.
Bachelor’s or advanced degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience.
5+ years of professional backend development experience, with a strong focus on Python.
Experience building distributed job or workflow systems: queueing, scheduling, worker coordination, concurrency control, cancellation, and failure recovery.
Strong relational database design (e.g., PostgreSQL) and schema migration experience.
Familiarity with AWS, particularly EC2, IAM, Secrets Manager, and S3.
Windows systems administration and PowerShell scripting.
Experience designing CLIs, libraries, or internal tools for other engineers.
Proficiency with GitHub workflows and modern CI/CD, including publishing internal packages.
Experience with Docker and Kubernetes.
Experience with ML experiment tracking or evaluation harnesses is a plus.
Experience automating game engines (Unity, Unreal) in headless/batch modes is a plus.
Being a gamer who understands player expectations, real-time systems, or interactive experiences is a plus.
Skills
As published by ashby
Name, Email, Resume
