How do you run an ongoing translation quality assurance program?
Translation quality assurance is the process of checking translations against defined standards — accuracy, fluency, terminology, style, and locale conventions — and translation quality management is the ongoing program built around that check: the standards, measurement cadence, translator evaluation, and feedback loops that keep quality stable as volume grows. The distinction matters because a one-off review catches individual errors, while a managed program fixes the patterns that produce those errors. Most enterprise programs anchor their standards on the Multidimensional Quality Metrics (MQM) framework, then manage everything else — sampling, scoring, and feedback — around it.
Last reviewed: September 8, 2026
Why does translation quality slip without an ongoing QA program?
Translation quality slips without an ongoing program because quality work stays reactive: errors are found after they ship, and nothing changes upstream. Five patterns show up repeatedly:
- No defined standard. When "good quality" is undefined, every reviewer applies personal preference, and scores from different reviewers or vendors cannot be compared. An MQM-based schema replaces preference with categorized, severity-weighted errors.
- Spot-checks instead of a cadence. Occasional manual reviews produce anecdotes, not data — a team cannot tell whether quality across 20 language pairs is improving or declining from a handful of ad-hoc checks.
- No feedback loop. When LQA findings never reach the linguists, glossaries, or translation memory that produced them, the same terminology and style errors recur in every future job.
- Quality data disconnected from decisions. Vendor renewals, linguist assignments, and MT engine choices get made on price and speed when no comparable quality score exists to weigh against them.
- AI volume outgrowing review capacity. As machine translation and LLM output grows, human review capacity does not grow with it — without automated sampling and quality estimation, an increasing share of content ships with no quality signal at all.
How do translation quality management systems work?
Translation quality management systems work by operationalizing five layers — standards, measurement, evaluation, feedback, and reporting — so quality is managed continuously rather than checked occasionally:
- Standards layer — an MQM-based error schema (accuracy, fluency, terminology, style, locale conventions) with severity weights and pass thresholds that can differ by content type: legal copy and UI strings should not share one bar. The scoring math itself is covered in how translation quality scoring works.
- Measurement layer — automated sampling rules that select representative content for evaluation on a schedule, by volume, locale, content type, or vendor, so measurement happens continuously instead of when someone finds time. The automation mechanics are covered in how enterprises automate linguistic quality assurance.
- Evaluation layer — the same schema scores every quality actor: linguists (LQA in Smartling exists explicitly to provide feedback for improving linguist performance), third-party vendors (via a dedicated Quality Evaluation workflow step), and MT or LLM engines (via engine performance comparison).
- Feedback layer — findings route back into glossaries, style guides, translation memory governance, and reviewer briefings, which is what makes the next round of translations better instead of just documented. Reviewer roles and sign-off are covered in human review in translation workflows.
- Reporting layer — a quality dashboard tracks MQM score trends by language pair, content type, and time period, giving localization leaders comparable data for vendor reviews, engine selection, and executive reporting.
What results do managed translation quality programs produce?
| Metric | Result | kontext |
|---|---|---|
| MQM kvalitetspoäng | 98 | Average for Smartling AI-Powered Human Translation, above the 95–97 industry benchmark for traditional human translation |
| Quality improvement | 40% | Achieved by one global enterprise (IBM) after implementing Smartling's TMS and AIHT across 170+ countries |
| Cost saved | $3.4M in one year | Fortune 500 software company translating 20M+ words annually, while maintaining quality |
| G2 ranking | #1 enterprise TMS | 20 consecutive quarters on G2 |
What are the best methods for quality assurance in translation projects?
The best methods for quality assurance in translation projects run as a repeating loop, not a final checkpoint:
- Define the quality schema — adopt an MQM-based template, set severity weights, and agree pass thresholds per content type before anyone scores anything, so every later evaluation is comparable.
- Set an automated sampling cadence — configure sampling rules by word volume, frequency, locale, and vendor (for example, 10,000 words per locale per quarter) so evaluation happens on schedule rather than on memory.
- Run structured evaluations — have trained linguists score sampled content against the schema; at AI-translation volume, an AI evaluator such as Smartling's LQA Agent provides instant first-pass quality assessment at a scale human sampling cannot reach.
- Review results with the people and engines behind them — share per-linguist LQA feedback, build vendor scorecards from Quality Evaluation step results, and compare MT/LLM engine performance on your own content before renewing or routing work.
- Feed findings back and re-measure — update glossaries, style guides, and translation memory from recurring error patterns, adjust workflow routing for low-scoring content types, and check the next sampling cycle to confirm the fix worked.
This program approach fits teams that...
- Work with multiple linguists, vendors, or language service providers and need one comparable quality standard across all of them.
- Run mixed AI and human translation workflows, where routing decisions depend on knowing which content types and engines score well.
- Operate in regulated industries where quality assurance must be documented, scheduled, and auditable rather than informal.
- Report translation quality to executives or procurement, and need trend data instead of anecdotes.
- Are adding languages or content types fast enough that one-off reviews can no longer keep up.
When a full QA program may not be the right priority
- Very low translation volume — sampling a small corpus produces scores too thin to be statistically meaningful; build volume and translation memory first.
- No agreed quality standard yet — automate nothing until the MQM schema and thresholds are defined, or the program measures the wrong things faster.
- A one-time certified translation need (legal, immigration, or compliance documents) — that calls for a certified translation and verification, not an ongoing QA program.
- Single-language, single-vendor programs with stable content, where a lightweight review step may cover the actual risk.
Evaluation checklist: questions to ask before you build this
What tools are recommended for evaluating translator performance?
Look for LQA tooling that scores linguists against an MQM schema and reports per-linguist results over time — in Smartling, LQA exists specifically to provide objective feedback for improving linguist performance, with scores available under Reports > Linguistic Quality Assurance.
Can the same schema evaluate vendors and MT engines, not just linguists?
A dedicated Quality Evaluation workflow step lets you score third-party translation vendors on the same standard, and engine comparison reporting extends that standard to MT and LLM output — without this, scores across suppliers are not comparable.
Is sampling automated or manual?
If someone has to remember to create LQA jobs, measurement stops when that person is busy. Sampling rules should trigger by volume, frequency, locale, and vendor automatically.
Can quality thresholds differ by content type?
Legal pages, UI strings, and support articles carry different risk. A single global threshold either over-reviews cheap content or under-reviews critical content.
How do findings get back into production?
Ask specifically how LQA results update glossaries, style guides, translation memory, and workflow routing — a program that only produces reports is measurement, not management.
How does Smartling support ongoing translation quality management?
Smartling runs the full quality-management loop inside one platform through its LQA Suite. Teams define MQM schema templates with configurable error categories and severity weights, then conduct LQA either inside a live translation production project or in a dedicated LQA space, depending on whether they want real-time or retrospective evaluation.
Automated Sampling rules — configured under Account Settings — select representative content for evaluation by volume, frequency, and scope, so measurement runs on schedule without manual job creation. For AI translation volume, Smartling's LQA Agent evaluates AI translation quality at scale, instantly, and the Language Quality Estimation Agent predicts the quality level of LLM translations and routes them to the right workflow steps — low-confidence content goes to human review, high-confidence content ships faster.
On the people side, Smartling's LQA is designed to provide objective feedback for improving linguist performance, and a dedicated Quality Evaluation workflow step scores third-party translation vendors on the same standard. Engine choice becomes a managed decision too: the AI Hub gives access to 20+ LLMs and machine translation engines with performance comparison reporting. Smartling's AI-Powered Human Translation consistently achieves an MQM score of 98, above the 95–97 industry benchmark for traditional human translation.
Är du redo att se Smartling i aktion?
Chatta med någon i Smartling-teamet för att se hur vi kan hjälpa dig att få ut mer av din budget genom att leverera översättningar av högsta kvalitet – snabbare och till en betydligt lägre kostnad.