← All jobs
Verita AI Verified
remote ·

Task & Evaluation Specialist

$60/hr
Share
Software Engineering remote
Posted Sep 2, 2026

What this role does Design realistic, difficult shopping-related tasks (e.g. meal planning, restocking, gifting, travel prep) that stress-test AI models, then build the rubrics used to grade how well a model performs them.

Who we're looking for

  • Subject-matter expertise in retail, e-commerce, consumer shopping behavior, or a related domain (nutrition, logistics, home/office management, etc.)
  • Strong analytical and writing skills — able to define precise, low-variance grading criteria
  • Comfortable working with AI outputs and evaluating them critically and objectively
  • Prior experience in eval design, QA, tutoring, or curriculum/rubric design is a plus

Core responsibilities

  • Author realistic shopping tasks with a clear, concrete end goal
  • Build a grading rubric/verifier for each task that two independent reviewers would score consistently
  • Review an AI agent's attempt at the task and grade it against the rubric
  • Document where the AI succeeded or failed, with clear reasoning

Work style

  • Independent, deadline-driven work
  • High attention to detail and consistency
  • Comfortable with ambiguity — tasks are self-directed with light guidelines
Pay range
$60/hr
Share
Similar roles

You might also like

micro1 Verified New
remote · hourly
Sr. Full-Stack Software Engineer
Software Engineering
Posted Sep 16, 2026
$50–$100/hr
Alignerr Verified New
remote · hourly
Ingeniero de Software Senior (Infraestructura IA)
Software Engineering
Posted Sep 16, 2026
$20–$100/hr
Alignerr Verified New
remote · hourly
Ingeniero de Software Senior (Infraestructura de IA)
Software Engineering
Posted Sep 16, 2026
$20–$100/hr
Alignerr Verified New
remote · hourly
Engenheiro de Software Sênior (Infraestrutura de IA)
Software Engineering
Posted Sep 16, 2026
$20–$100/hr