Skip to content

Translation quality checks that refuse to ship.

Everyone else flags a bad translation and ships it anyway. Four checks run on every row before a human sees it, a row that fails is repaired once, and a row that fails again is held rather than published. Open the report below and read what it caught.

GxGlossa Flux and Glossa Deep
one repair pass, then refusal

Everything checked before a customer can read it.

Everyone else flags. These repair, once, and then refuse rather than ship a broken row. Placeholders, plurals, glossary and length are checked deterministically before any model is asked for an opinion, so the cheap failures never cost a token.
self heal
off · suggest · apply

Catches it. Repairs it. Then refuses.

One corrective pass, and a row that still fails never ships.

de-DE length 67 > 48 budget
heal re-drafting inside the budget
de-DE length 41 pass
ja-JP glossary forbidden term
heal one retry · still failing · refused
the four gates
free, every row

Checked before a model is asked anything.

  • placeholders
  • plurals
  • glossary
  • length
the semantic pass
5

MQM dimensions, scored

Accuracy, terminology, fluency, style, locale convention. A scorecard, and the run out as CSV.

source drift

The English moved. The row knows.

A changed source marks every locale stale and re-drafts it, then sends it back to review.

memory
$0

A sentence you settled costs nothing again.

Exact matches are free. Semantic memory, scoped to your org, from Team.

terminology
warn, or refuse

Decide which words are law.

Lock a term, forbid a term, or set a preferred target per language. Then choose whether breaking one is a warning or a refusal.

  • do not translate
  • forbidden
  • preferred per locale
  • TBX and CSV import
term mining

It can write your glossary for you.

Point it at the strings you have already approved and it proposes the terms that keep recurring. You approve them; nothing is added behind your back.

self heal
your call

How much it is allowed to fix on its own.

Off, suggest, or apply. On apply it repairs a failing row once and re-checks; on suggest it shows you what it would have written and waits.

  • off
  • suggest
  • apply
the report
csv out

Take the whole run to someone who does not have a login.

Every finding, every severity and every language in one scorecard, exportable as CSV. It is the artefact a language lead asks for and the one an auditor accepts.

Six stages between your English and a language you cannot check.

It advances as you scroll, or pick a stage. Nothing here is a summary: each one lists what it actually does, in the order it does it.

01

The key arrives with its world attached

A string is never just text. The engine parses the format, extracts placeholders and plural skeletons, and attaches the rendered node it ships in where a screenshot exists.

specification
  • 1.1
    Parse the format, extract placeholders and plural categories
  • 1.2
    Attach the rendered node and its length budget
  • 1.3
    Pull the key history and any past repairs

Automated QA that repairs what it can, and refuses what it cannot.

live: choose how much the engine may fix on its own, then run the queue. Every row below is broken a different way, and one is still broken after its repair, because this plate does not pretend every repair works
acme / checkout · agents
self heal
keysourcetargetgatesstate
checkout.ctaDEPay {amount} nowJetzt bezahlenqueued
files.countRU{count, plural, one {# file} other {# files}}{count, plural, one {# файл} other {# файлов}}queued
account.titleDEYour accountIhr Kontoqueued
cta.upgradeDEUpgrade nowJetzt auf einen höheren Plan wechselnqueued
home.heroJATranslate everythingすべてを翻訳するqueued
footer.legalFRAll rights reservedTous droits réservésqueued
empty.cartDEYour cart is emptyDein Warenkorb ist leerqueued
queue
nothing run yet_
set self heal above, then run the queue. The gates print here
scorecard
0
passed
0
repaired
0
suggested
0
refused
0
stale
0
re-drafted
MQM by dimensionillustrative
accuracy·
terminology·
fluency·
style·
locale·
Scores appear once a row clears. The figures are illustrative, written into the fixture, and not a benchmark.
off refuses and holdssuggest waits for youapply writes it in, oncea repair that still fails is heldsource drift re-drafts a stale row
A simulation, in your browser. Seven rows on a fixture, gated by the shipped rules; the repairs and the MQM figures are illustrative. Nothing is sent anywhere and nothing is downloaded.
in the product

The gate report, as it ships.

Languages down the side, checks across the top, and the cell you open lists the strings behind it with a repair on each row.
The QA screen: three languages down the side, eight checks across the top, and the French placeholder cell opened onto its three failing strings.
four gates, then the review
a demo workspace, real checks
the contract

Safe automation is a shared contract.

Two lists. Neither one works without the other.

The engine holds

  • okThe four gates run on every row: code strings, connector content and campaigns alike
  • okYour content is never used to train a model, and the shared exact-match memory pool is anonymised and can be switched off per project
  • okOne automatic repair pass, then a refusal with the evidence attached
  • okHash chained audit trail, and an SLA target of 99.9% on Business and 99.95% on Enterprise
  • okThe exit: the pull endpoint, the CLI, the CDN bundle and file storage sync

You hold

  • okThe glossary: which words are law, and what they become per locale
  • okLength budgets, and the design time hints that set them
  • okEscalation rules: which keys always reach a human
  • okReviewers and locale owners, per market
  • okThe merge button. A red gate blocks, and you decide what happens next

How we test the thing that tests everything.

Four problems that changed the engine, written up as working notes rather than as claims. None of them has a number attached, because we have not published a study.

Ranking memory by provenance

Why the best segment is not the nearest embedding, and how who approved it beats how similar it looks.

Six plural forms are a search space

Arabic broke an assumption baked into the plural gate. Fixing it made every language safer.

Drafting against the rendered node

What changes when the draft can see the component instead of guessing at a character budget.

Why refusing beats flagging

The highest leverage behaviour in the whole engine is the one that declines to ship.

What the checks change, in shape rather than in numbers.

The chart below is an illustration of the shape we design for, not a published benchmark. We do not run a head to head against named competitors, and we will not print a number we cannot source.

First pass quality

illustrative
GlossaGeneral AIOff the shelf MT

Shape only. No measured study stands behind this.

Placeholder and plural accuracy

illustrative
GlossaGeneral AIOff the shelf MT

The gates are deterministic, which is why this bar is the tallest.

Reviewer time saved

illustrative
GlossaGeneral AIOff the shelf MT

Depends entirely on your glossary and your locales.

Illustrative only. No benchmark study is published; there is no measured figure behind these bars.

FAQs

What people ask before they let a machine write in a language they cannot read.

How is this different from wrapping a general purpose model?
A wrapper sends your string to a general model and hopes. This is a pipeline: retrieval binds your memory and your locks first, the draft is written inside hard constraints, four deterministic gates verify the result against the rendered component, and a repair pass rewrites or refuses. The model is one stage of six.
Plural category completeness, placeholder integrity, maximum length against the component budget, and glossary terminology. They are deterministic checks in the QA runner, not model opinions, and they run on every row of every run on every plan.
Two passes over the same pipeline. Flux is the fast lane and runs on every plan. Deep is the careful lane for the projects and locales where nuance is money, and is included from the Team plan. Neither is a versioned model you have to pick.
On the MQM typology: five dimensions, accuracy, fluency, terminology, style and locale convention, with three severities, minor, major and critical. That is the frame reviewers work in, and it is the frame the reports use.
No. Your strings are sent to the subprocessors named in the Trust Center to be translated and nothing else, and nothing you send is used to train a model. The shared exact-match memory pool is anonymised and can be switched off per project on every plan. The Trust Center is the authoritative list of who can process your content and in which region.
Not yet. There is no published evaluation study behind a win rate or a latency figure, so we do not print one. The chart above is labelled illustrative for exactly that reason.
Yes. Lock terms, set length budgets and write escalation rules so that a key set always reaches a human. The engine holds the mechanics and you hold the policy.
works with what you already run

41 connectors, already built.

your turn

Get the pipeline that knows when to say no.

Run it on your own strings and read the gate report before you decide anything.

Four gates on every plan. Glossa Deep from the Team plan.