Étude de cas
Bilingual annotation and model alignment for a global AI developer
Le défi
The client's models performed acceptably in European French and noticeably worse in Canadian French. Quebec French idiom was misread, register slipped, and anglicisms that a Montrealer would never use appeared in generated text. Their evaluation set showed the gap clearly and their vendors could not close it, because previous programmes had used European French speakers on Canadian French material and treated the difference as regional colour rather than a distinct variety.
There was a second requirement the client had struggled with: they wanted annotators in a jurisdiction with strong labour standards and enforceable contracts, after scrutiny of working conditions in parts of the global annotation industry.
Ce que Corpshore a fait
We recruited and trained a distributed Canadian workforce of Canadian French and English speakers, matched by variety rather than by language. Quebec, Acadian, Franco-Ontarian and western francophone speakers all worked on Canadian French material, with regional variation treated as signal rather than noise.
The programme ran in four phases: text annotation and preference labelling, Canadian French speech transcription including French and English code-switching, model output evaluation for accuracy, register and cultural appropriateness, and RLHF preference data operations with a layered adjudication tier.
Quality was managed as a gate rather than a report. Every batch passed layered review with measured inter-annotator agreement, and a batch that failed did not ship. Where guidelines produced inconsistent outcomes we said so and proposed revisions rather than delivering data we knew was unstable. All annotators were engaged on Canadian contracts under Canadian labour standards, with published rates and a documented progression structure, which the client audited.
Modèle de prestation
Programme-based with output pricing and quality gates, remote Canadian workforce, Corpshore AI quality operations layer, full consent and governance chain for all collected speech.
Résultats
Delivered several hundred hours of transcribed and annotated Canadian French speech plus large volumes of preference and evaluation data across four phases, all passing client acceptance. Inter-annotator agreement held above the programme threshold throughout, including through two scale increases. The client reported measurable improvement in Canadian French performance on their internal evaluation set following training on delivered data, with the largest gains on idiom and register. The programme extended twice and expanded into additional languages.
Pourquoi cela a fonctionné
Variety competence does not transfer within a language. Matching annotators to the specific variety, and refusing to treat Canadian French as a dialect of European French, is the entire difference between data that passes acceptance and data that does not.
Ce client est anonymisé de façon délibérée. Plusieurs types d'acheteurs ne peuvent être nommés sans autorisation contractuelle, et les mandats du secteur public l'interdisent souvent. Les indicateurs présentés ici proviennent des données de mandat.
Bâtissez votre équipe canadienne
Décrivez-nous le travail, les langues et la couverture dont vous avez besoin. Vous obtiendrez une réponse réfléchie en moins de six heures, ou planifiez dès maintenant un appel exploratoire.
Vous cherchez un poste plutôt qu'un partenaire ? Explorez les carrières chez Corpshore Canada
