AI delivery
RLHF Quality Reviewer and Model Evaluator
About the role
Reinforcement learning from human feedback is only as good as the judgement behind it, and at scale that judgement has to be audited. You will be the audit. You will review the work of annotators, adjudicate disagreements, and identify whether a quality failure originates in an individual, in the training, or in the guideline itself.
Fully remote across Canada, minimum twenty hours weekly.
What you will do
Review annotator preference judgements, ratings and written responses for correctness, guideline consistency and freedom from bias. Adjudicate disagreements between annotators with a documented rationale that a client can audit. Sample and audit batches, reporting quality metrics against programme thresholds. Distinguish systematic error from individual error and say which, because the remedies are entirely different. Give annotators feedback that measurably improves their next batch. Contribute to guideline refinement where the current version produces inconsistent outcomes. Support red team and safety evaluation programmes, which can involve reviewing content designed to elicit harmful model output. That work is optional, clearly flagged at assignment, and supported.
What you bring
English at C1, with rigorous written reasoning. At least two years in annotation, linguistic quality, editorial work, translation review, teaching, research or a comparable judgement-based discipline. Ability to reason carefully about ambiguous cases and to write down the reasoning. Sound judgement on safety, bias and factual accuracy. Total reliability on deadlines, since a delayed audit blocks a training run.
Nice to have
Prior RLHF or model evaluation experience. Postgraduate qualification in linguistics, philosophy, law, cognitive science or a scientific field. French at C1, which opens bilingual programmes at a premium rate. Subject depth in a specialist domain.
What we offer
Rates that rise with review tier and language scarcity. Fully remote and flexible within weekly minimums. Paid guideline training per programme. Rolling contracts with steady volume. Work that directly shapes model behaviour, and a clear route into programme leadership.
How to apply
Our process
- 1. Our talent team reviews every application against the role requirements.
- 2. Shortlisted candidates are invited to interview, which may include a role-related assessment.
- 3. We share the decision with everyone we interview, whether or not the application progresses.
Most decisions follow within a few weeks of the closing date. Urgent roles are prioritised.