Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Join a talent network for contract BI analyst roles, building tasks and evaluating AI models using SQL, Tableau/Power BI, and data warehousing.
Create and critique rubrics to evaluate biotech/pharma drugs and drug-development programs, assessing mechanisms, clinical data, and competitive positioning for investment-style analysis.
Create challenging CAD/mechanical design problems to test AI models, including expert solutions and validation against frontier language models.
Evaluate and improve frontier AI coding models by completing and reviewing data engineering tasks like ETL pipelines and data warehouses using AI agents.
Oversee AI training projects for an LLM startup, tracking annotator performance and ensuring data quality through Google Sheets and contributor communications.
Designs and writes advanced circuit-engineering free-response questions to stress-test AI models, requiring a PhD in EE and 5+ years of shipped analog/digital/power/RF designs.
Build and evaluate AI systems by designing grading criteria for data science deliverables and assessing AI-generated or human-created work against those standards.
Designs graduate-level computational problems to test AI’s ability to run scientific software, interpret results, and plan experiments in domains like RF/circuit design using tools like scikit-rf and ngspice.
Create, curate, and refine content to train large language models by summarizing, researching, and validating text for clarity and accuracy.
Leads end-to-end GenAI implementation for a Fortune 500 financial-services client, designing architecture, setting security guardrails, and managing delivery from POC to production. Works with AI/ML engineers and client stakeholders to ensure compliant, scalable AI solutions for valuation workflows.
Senior technical editor creates and refines complex AI evaluation tasks, style guides, and documentation to test and improve frontier AI agents' reliability in corporate settings.
Analyze legal documents and case law to evaluate and improve large language models' legal reasoning and accuracy.
Evaluate AI chatbot responses for retail users by crafting prompts and rating clarity, accuracy, and personalization using Google Workspace apps.
Evaluate AI research reproductions by reading papers, inspecting code outputs, and ranking attempts with written justifications for a remote freelance role.
Build and evaluate datasets of verifiable software-engineering tasks from public GitHub repos to train LLMs on realistic coding problems, including environment setup, issue triage, and test-coverage analysis.
Analyze and break down content to train and improve large language models using analytical reasoning and multilingual skills (English and Korean).
Analyze and break down content to train and improve large language models using English and Danish, solving analytical puzzles and validating claims.
Analyze and break down content to train and improve large language models using analytical reasoning and multilingual skills (English and Swedish).
Designs and solves advanced computational problems in Python for AI research, focusing on rigorous problem formulation and solution accuracy for frontier AI labs.
Evaluate and rank personalized AI responses from Gemini by designing multi-turn prompts and analyzing grounding, integration, and helpfulness for Chinese-language users.
We couldn't check your fit for this role — add a CV to your profile to see it next time.