Python Engineer - AI Data & Evaluation
Summary
3-month freelance contract (min 20 hrs/week, remote LatAm) for a Python/TypeScript engineer to design AI evaluation rubrics, review preference data, and build pipelines/infrastructure that improve model training data quality. Requires 5+ years experience and prior work at Tier 1 companies.
About Gramian
Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.
About the Role
We are looking for an experienced engineer to support advanced AI development projects focused on improving the quality and reliability of model training data. This role combines evaluation framework design, preference data review, and engineering infrastructure, requiring someone who can translate complex technical judgments into clear rubrics and scalable workflows. You will work closely with researchers and engineers to improve data quality, build supporting pipelines, and deliver concise, actionable insights in a fast-moving environment.
CONTRACT: Freelance / Contractor, short-term (3 months)
COMMITMENT: Minimum 20 hours per week, at least 4 hours per day, with 4 hours of overlap with PST working hours
LOCATIONS: Remote — LatAm
PROCESS: 1 interview round (approximately 30 minutes, technical and cultural discussion)
NOTES: Mandatory Experience working in Tier 1 companies
Responsibilities
- Design and refine evaluation rubrics and criteria for preference data.
- Review and label data while identifying subtle quality issues and inconsistencies.
- Build and maintain Python or TypeScript pipelines and infrastructure supporting data generation, collection, and evaluation.
- Develop processes that translate qualitative assessments into scalable evaluation workflows.
- Synthesize findings from data reviews into clear technical write-ups, updates, and recommendations.
- Collaborate with researchers and engineers to improve evaluation methodologies and data quality.
- Iterate on evaluation processes based on observed issues and project requirements.
- Communicate technical findings in a structured, concise, and actionable format.
Requirements
- At least 5 years of professional experience with Python or TypeScript.
- Hands-on experience building or maintaining pipelines and infrastructure using Python, or contributing to large TypeScript codebases.
- Practical experience with preference data collection, review, or rubric creation.
- Experience designing or applying evaluation criteria for AI-generated content or technical outputs.
- Strong written and verbal communication skills, demonstrated through clear technical documentation or summaries.
- Ability to explain and justify quality judgments using explicit evaluation criteria.
- Experience navigating or contributing to large codebases, such as VSCode, is an advantage.
- Prior experience working on AI, machine learning, data quality, or evaluation projects is an advantage.
Skills
As published by workable · 5 questions · 1 written answer
Basics
First name, Last name, Email, Phone, Address, Summary, Resume
Short answers (2)
- Please provide a valid and up-to-date LinkedIn profile URL that clearly reflects your professional experience, technical skills, and employment history relevant to the job. Example format: https://www.linkedin.com/in/yourprofile Applications with incomplete, inactive, or non-professional LinkedIn profiles may not be considered.
- Which Tier 1 company have you worked for?
Pick from a list (2)
- Have you worked as an engineer at a leading technology company or contributed to a large-scale software/AI project?
- Which best describes your hands-on experience with preference data collection or evaluation rubric creation?
Written answers (1)
- What is your expected hourly rate? More affordable candidates will be prioritized.