Abstract
As large language models (LLMs) are increasingly used for machine translation, questions of quality, biases and the representation of languages (and language varieties) have become more pressing. While evaluation is often framed as a technical benchmarking task, assessing translation quality (particularly across regional varieties) requires deep contextual and sociolinguistic expertise. This creates a productive but complex space for collaboration between translation scholars and professional translators.
This paper presents a participatory research approach to evaluating LLM-generated translations in which professional translators are not merely evaluators, but co-designers of assessment frameworks. The project brings together professional translators to systematically review machine-generated translations, with a specific focus on variation (e.g. regional standards and non-dominant varieties).
Drawing on iterative evaluation workshops and structured annotation protocols, this paper reflects on three key dimensions of collaboration: (1) negotiating quality criteria between technical metrics and professional standards; (2) recognising expertise which can come in many forms; and (3) addressing power dynamics related to prestige varieties and marginalised forms of language. The results are informing recommendations for using LLMs for specialised translation purposes and a policy brief on language diversity in LLMs.
Involving professional translators in participatory research on LLM-generated translation evaluation strengthens both methodological rigor and professional legitimacy. At the same time, such collaborations require explicit attention to role clarity, ethics and shared authority in knowledge production.