← All jobs
Turing Verified
remote ·

AI Safety and Policy Evaluator

Pay on listing
Share
Other remote
Posted Sep 24, 2026

About Turing:

Based in San Francisco, California, Turing is the world's leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results.


Nature of Engagement:

This is a project-based independent contractor engagement. Contractors are engaged under a statement of work (SOW), determine their own methods and service hours, use their own equipment, are free to provide services to others, and are responsible for their own taxes, insurance, and expenses. No benefits, paid leave, or continuing engagement is offered or implied.


Project Description:

This project supports the evaluation and improvement of safety behavior in large language model systems. Contractors engaged on this project design adversarial test prompts, identify where model safeguards fail, and author the evaluation rubrics and written rationales used as training signal.


The work calls for creativity, analytical rigor, and applied policy judgment. Contractors design their own tests and determine their own approach to discovering model failures, then articulate why a failure occurred in a form that meets the project's specifications.


Deliverables:

Work product under this project centers on evaluating single-turn image edit requests against detailed safety guidelines and a defined category taxonomy. It may include:

  • Safety classifications of single-turn image edit prompts and their resulting outputs, applying the project's guidelines and category taxonomy (e.g., photorealistic human generation, likeness of real people, minors, self-harm, sexual or suggestive content, non-consensual intimate imagery, violence, hate and harassment, illegal activity, and benign boundary-control cases).
  • Consistent application of the taxonomy to ambiguous and borderline cases, including correct classification of benign requests that superficially resemble policy violations.
  • Clear, defensible written rationales supporting each classification decision, sufficient to serve as training signal.
  • Written documentation of identified model failures, including cases where safeguards were bypassed and subtle or edge-case policy violations.
  • Where specified by the applicable SOW: creative test prompts designed to probe the boundaries of a given safety category, and comparative evaluations or rankings of multiple model outputs against the same request.
  • Flagged gaps, ambiguities, or conflicts in the guidelines themselves, with suggested clarifications.

Project specifications, safety guidelines, and category definitions are provided as reference materials. Contractors may request clarification of specifications at any time; the manner and means of producing the deliverables remain the contractor's own.


Content Advisory

This project involves testing model safeguards across a defined set of sensitive categories. Unsafe or harmful material is encountered only as text — the prompts written and evaluated as part of the work. Any images reviewed in the course of the project are not intended to be egregious and contractors generally will not view generated imagery depicting harmful content.


The categories in scope are: photorealistic human generation; depictions of real, identifiable people and likeness; content involving minors; self-harm; sexual or suggestive content; non-consensual intimate imagery of a real person; violence and graphic content; hate, harassment, and insults; illegal activities and misconduct; and a benign boundary-control category used as a comparison baseline.


Written material on these subjects can still be distressing, and we want contractors to understand that before taking on the project. Its inclusion does not reflect the views, values, or policies of Turing or its clients. A contractor may decline or discontinue participation in all or part of any project at any time, and Turing will make reasonable efforts to offer alternative project work.

Wellbeing and risk-mitigation resources are available at no cost.


Background Sought

  • BS/BA degree or equivalent experience in a relevant field (e.g., policy, law, ethics, linguistics, journalism, computer science, or a related analytical field).
  • Prior experience in content moderation, policy analysis, AI safety evaluation, or comparable work is strongly preferred.

Skills and Resources Required to Perform the Work

  • Analytical judgment: demonstrated ability to research and evaluate nuanced, complex, and ambiguous information against defined policy criteria.
  • Creative and adversarial approach: experience in red teaming, prompt engineering, or designing challenge prompts intended to test and bypass AI safety filters.
  • Policy and taxonomy acumen: strong grounding in Trust & Safety principles as applied to LLMs (misinformation, abetting, bias and stereotypes, jailbreaks, dual-use).
  • Precision: ability to author precise, self-contained evaluation rubrics that clearly discriminate between model outputs.
  • Written communication: ability to articulate complex rationale clearly and concisely.
  • Familiarity with RLHF workflows and data annotation is a significant plus.
  • Equipment: contractor supplies their own desktop or laptop and reliable internet connection.

Engagement Terms

  • Compensation: Project-based fees set by SOW, determined by scope and specialized expertise. Required training time is separately compensated as described above.
  • Estimated volume:  Contractors set their own schedules.
  • Term: 1-2 weeks, per SOW. Turing may offer additional statements of work for subsequent project phases. Each SOW is a separate engagement, and nothing here creates an expectation of continued or ongoing work.
  • Non-exclusive: Contractors remain free to perform services for other clients.
Pay range
Pay on listing
Share
Similar roles

You might also like

Turing Verified New
Canada · remote
Personal Account Generalist Raters (Canada)
14h ago
Pay on listing
Alignerr Verified New
remote · hourly
English Writing Generalist – Quality Review
20h ago
$30–$40/hr
Terac Verified New
remote
Video Editors: Paid Interview on Creative Workflows
21h ago
$33 one-time
Terac Verified New
United States · remote
Active Donors: 30-Minute Interview on Charitable Giving
Yesterday
$105 one-time