freehire launches on Product Hunt on 26 August.

Follow →

AI Harness Engineer

Continuously monitor and evaluate leading AI vendors, frontier models, and emerging AI technologies, assessing their capabilities, cost, context handling, and agentic performance. Provide clear recommendations on model suitability and adoption.

Develop and implement a vendor diversification and resilience strategy by evaluating alternative AI vendors, model routing solutions, and self-hosted LLM options. Establish reliable fallback and switching mechanisms to reduce dependency on vendors.

Build and maintain internal AI evaluation benchmarks based on real-world codebases and business scenarios to objectively measure and compare model performance.

Take ownership of and continuously improve company-wide AI development workflows and best practices, including prompt and context engineering standards, agent configurations, code review and validation processes, and safety guardrails.

Define and enforce AI security and compliance standards to prevent source code and sensitive data leakage, mitigate prompt injection risks, establish security review requirements for AI-generated code, and implement appropriate data masking and access controls.

Build and maintain internal AI harnesses and developer tooling, including shared configurations, skills and plugins, evaluation scripts, and CI integrations, enabling engineering teams to adopt AI tools efficiently.

Document AI development best practices, conduct technical training and knowledge-sharing sessions, and track AI adoption and developer productivity metrics across engineering teams.

Key Requirements

Have 5 or more years of working experience in the similar field.

Extensive hands‑on experience with agentic coding tools such as Claude Code, Codex, Cursor, or similar tools, with practical knowledge of the strengths and limitations of different frontier AI models.

Strong understanding of prompt engineering, context engineering, and model evaluation methodologies, including benchmark design and evaluation harnesses.

Strong technical curiosity and self-motivation, with the ability to independently follow developments in the rapidly evolving AI ecosystem, conduct evaluations, and translate findings into actionable recommendations.

Good understanding of key LLM and AI application security risks, including data leakage, prompt injection, and supply chain risks associated with AI-generated code, as well as relevant mitigation practices.

Experience in LLM application development, AI evaluation frameworks, developer productivity, or platform engineering.

Experience with self-hosted LLM deployment, model routing, or AI inference infrastructure.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available