Solutions d'IA
Réviseur qualité RLHF et évaluateur de modèles
À propos du poste
Reinforcement learning from human feedback is only as good as the judgement behind it, and at scale that judgement has to be audited. You will be the audit. You will review the work of annotators, adjudicate disagreements, and identify whether a quality failure originates in an individual, in the training, or in the guideline itself.
Fully remote across Canada, minimum twenty hours weekly.
Vos responsabilités
Review annotator preference judgements, ratings and written responses for correctness, guideline consistency and freedom from bias. Adjudicate disagreements between annotators with a documented rationale that a client can audit. Sample and audit batches, reporting quality metrics against programme thresholds. Distinguish systematic error from individual error and say which, because the remedies are entirely different. Give annotators feedback that measurably improves their next batch. Contribute to guideline refinement where the current version produces inconsistent outcomes. Support red team and safety evaluation programmes, which can involve reviewing content designed to elicit harmful model output. That work is optional, clearly flagged at assignment, and supported.
Votre profil
English at C1, with rigorous written reasoning. At least two years in annotation, linguistic quality, editorial work, translation review, teaching, research or a comparable judgement-based discipline. Ability to reason carefully about ambiguous cases and to write down the reasoning. Sound judgement on safety, bias and factual accuracy. Total reliability on deadlines, since a delayed audit blocks a training run.
Atouts supplémentaires
Prior RLHF or model evaluation experience. Postgraduate qualification in linguistics, philosophy, law, cognitive science or a scientific field. French at C1, which opens bilingual programmes at a premium rate. Subject depth in a specialist domain.
Ce que nous offrons
Rates that rise with review tier and language scarcity. Fully remote and flexible within weekly minimums. Paid guideline training per programme. Rolling contracts with steady volume. Work that directly shapes model behaviour, and a clear route into programme leadership.
Comment postuler
Notre processus
- 1. Notre équipe de recrutement examine chaque candidature reçue par rapport aux exigences du poste.
- 2. Les candidats retenus sont invités à une entrevue, qui peut comporter une évaluation liée au poste.
- 3. Nous communiquons la décision à chaque personne interviewée, que sa candidature progresse ou non.
La plupart des décisions suivent dans les quelques semaines qui suivent la date de clôture. Les postes urgents sont traités en priorité.
