This website uses Google Analytics cookies to enhance your user experience and analyse our performance.
We’re SEO nerds, after all! If you’d like to help us out a little, please click accept so we can track your visit.

Translation Memory (TM): How It Works in the AI Era (2026)

A translation memory (TM) is a database that stores previously translated text segments – sentences, phrases, paragraphs – paired with their source text. When similar or identical content appears in a new project, the software retrieves the stored translation and offers it as a suggestion. The translator accepts, edits, or replaces it, and moves on.

Used well, a TM reduces time spent on repeated content, enforces terminology consistency across documents and markets, and lowers costs on high-volume projects where repetition is high. It is one of the most widely used technologies in professional translation: according to CSA Research, only 12% of translators who use CAT tools never use a translation memory.

In this post:

What is a translation memory?

A translation memory (TM) is a bilingual database of translation units: each unit pairs a source-language segment with its approved target-language equivalent. As a translator works through a project, every completed segment is stored. The next time that sentence – or something close to it – appears in any future project, the TM surfaces the stored translation automatically.

Used well, a TM reduces time spent on repeated content, enforces terminology consistency across documents and markets, and lowers costs on high-volume projects where repetition is high. It is one of the most widely used technologies in professional translation: according to CSA Research, only 12% of translators who use CAT tools never use a translation memory.

A brief history of translation memory

The idea predates the technology by decades. In 1980, linguist Martin Kay at Xerox PARC published a paper called The Proper Place of Men and Machines in Language Translation, arguing for cooperative human–machine translation systems. The concept of reusing approved translations was central to his proposal.

Commercial TM software arrived in the 1990s, with Trados becoming the first widely adopted tool in the industry. Through the 2000s, as internet commerce expanded and the volume of content requiring translation grew sharply, TM became standard infrastructure for translation agencies and large enterprise localisation teams.

The 2010s brought cloud-based and server-side TM, which allows multiple translators to contribute to and draw from the same database in real time. A centralised server-based TM can increase match rates by 30–60% compared with desktop-only TM systems, because every translation produced by any member of the team immediately becomes available to everyone else.

Today, TM sits at the core of most CAT tools and translation management systems, and it increasingly works in combination with machine translation and AI-assisted quality assurance.

How does a translation memory work?

Translation memory columns
Source: ProZ

Segmentation

Before a translator opens a file, the CAT tool or TMS splits the source text into segments – typically sentence by sentence, though this depends on the tool’s segmentation rules and can be configured per project. Each segment becomes a translation unit (TU): a source-language segment paired with its target-language equivalent once translated.

Matching

As the translator works, each source segment is compared against the TM database in real time. The result is one of four match types:

Match typeWhat it meansTypical translator action
Context match (101%)Exact text + identical surrounding segmentsAuto-inserted and locked
Exact match (100%)Word-for-word identical, different contextInserted and reviewed in context
Fuzzy match (70–99%)Similar but not identicalTranslator edits the suggestion
No match (below threshold)No useful overlapTranslated from scratch

Most tools set a minimum fuzzy match threshold of around 70–75%. Below that, the editing effort is comparable to starting fresh, so no suggestion is surfaced.

Concordance search

Alongside the match system, most CAT tools include a concordance search function. Rather than waiting for a full-segment match, the translator can search the TM for any stored segment containing a specific word or phrase. This is particularly useful for verifying how a term was previously translated in context, even when no automatic suggestion has been triggered.

Concordance search in translation memory
Source: Wordfast

Updating the TM

Once a project is complete, all newly translated and approved segments are added to the TM database, expanding it for future use. The TM grows with every project. In collaborative environments, a centralised TM means each translator’s work immediately benefits everyone else on the team.

What are the benefits of using translation memories?

There are many benefits of using translation memories, both for individual translators and localization businesses. Let’s take a look at some of the main ones:

Increased efficiency and productivity

Translation memories help to speed up the translation process by suggesting translations for repeated or similar segments, which means that translators can work more quickly and efficiently. Productivity gains of 30-40% are not uncommon when using TM tools.

A word of caution if you’re a translation buyer: Beware of language service providers (LSPs) or translation agencies offering discounts based on “high match rates” from translation memory. While tempting, agencies tend to pass on the cost of these discounts to translators in the form of lower rates, which can erode quality.

Improved consistency and quality

One of the main advantages of using translation memories is that they promote consistency by allowing translators to re-use approved translations and ensuring that repeated segments remain consistent. This can be a real boon for large or long-running projects, where maintaining consistent terminology and style is essential.

In software localisation projects, for example, it’s often necessary to translate user interface (UI) elements such as menus, dialog boxes, and buttons. Using a translation memory can help to ensure that these UI elements are translated in the same way each time, which is crucial for a good user experience.

Faster turnaround times

Because translation memories help to speed up the translation process, they can also lead to shorter project turnaround times. This is especially true for long-running or ongoing projects with regular updates – like agile localisation projects – where being able to quickly re-use approved translations can make a big difference.

For projects where timing is critical, working with a language service provider that uses translation memories can give you a real competitive advantage.

Lower costs

Because translation memories help to improve efficiency and consistency, they can also lead to cost savings in the long run. These savings don’t only concern the cost of translation itself, but also the cost of project management, quality assurance, and other related activities.

It’s important to note that while using a translation memory can save money in the long run, there is an initial investment involved in setting up and maintaining a TM. This investment can be significant for large projects, so it’s important to weigh up the costs and benefits before deciding whether or not to use translation memory technology.

Translation memory vs glossary: what is the difference?

These two tools are complementary and almost always used together, but they work at different levels.

Translation memoryGlossary / termbase
What it storesFull translated segments (sentences, phrases)Individual approved terms
How it worksRetrieves suggestions when matching segments appearEnforces consistent use of specific terms throughout
LevelSentence levelWord / term level
BehaviourSuggestsTypically enforces (raises a warning if term is not used)
Best forRepetitive content across projectsBrand names, technical vocabulary, regulated terminology

A glossary or termbase lists approved translations for specific terms – product names, acronyms, technical jargon, regulated language – along with definitions and usage rules. A TM handles full sentences; a termbase handles the vocabulary inside them. For any project with strict terminology requirements, both are necessary.

Translation memory vs machine translation: what is the difference?

The two technologies are often mentioned together and frequently confused. They serve different purposes and work differently.

Translation memoryMachine translation
What it doesRetrieves stored human-approved translationsGenerates new translations algorithmically
SourceYour own approved contentTrained on publicly available data
OutputPreviously approved text, exact or near-exactNew generated text, quality varies
Context sensitivityLimited to segment-level matchesCan model broader context with modern NMT
Best forRepeated, known content you trustNew content with no TM match

In most modern workflows, TM is checked first. Segments with no match above the threshold may receive a machine translation suggestion as a fallback. The translator then uses whichever input is more useful.

Research on TM and MT integration shows that combining both technologies improves post-editing efficiency and translator productivity compared with either tool in isolation.

The key practical difference for clients: a TM is built from your own content and encodes your specific terminology, regulatory requirements, and brand voice. A machine translation engine has never seen your content and cannot replicate that specificity, however fluent its output.

TM analysis: understanding leverage before a project starts

One of the most useful but least-discussed features of TM is the pre-project analysis – sometimes called a TM leverage report or word count analysis.

Before translation begins, the CAT tool scans the source file against the TM and produces a breakdown of:

  • How many words are exact or context matches
  • How many words fall within fuzzy match ranges (typically shown in bands: 95–99%, 85–94%, 75–84%, 70–74%)
  • How many words have no match at all (new content)
  • Internal repetitions within the current document itself

This gives an immediate picture of how much work is genuinely new versus how much can be handled from the TM. For project managers and clients, it is the basis for realistic time and cost estimates before work begins.

A file with 40% exact matches and 20% high-fuzzy matches looks very different from one where everything is new content, and the TM analysis is what surfaces that difference upfront.

What types of content benefit most from translation memory?

TM delivers the greatest value where content is repetitive, high-volume, updated frequently, or subject to strict terminology requirements. The clearest use cases:

Software and product localisation

UI strings – buttons, labels, error messages, onboarding copy – repeat heavily across platforms and product versions. Strings like “Save changes”, “Cancel”, “Create account” appear in web, iOS, Android, and desktop builds. With a TM in place, minor product updates do not trigger a full re-translation; translators work only on what has actually changed.

According to Smartling’s 2024 State of Translation report, translation volumes are up 30% year-over-year as digital product teams scale content across more markets. A TM is what makes this volume manageable without proportional cost increases.

Technical documentation and product manuals

Warning messages, safety instructions, and procedural steps appear in near-identical form across product generations and product lines. A 10,000-word manual for a new product version may share 60–70% of its content with the previous version – all of that is potentially covered by TM matches.

Standard contract clauses, terms and conditions, privacy notices, and regulatory disclosures must be worded consistently across documents, markets, and time. A TM locks approved phrasing in place and flags if a translator departs from it, which is exactly what legal consistency requirements demand.

E-commerce product content

Product descriptions for large catalogues tend to follow templates: “available in X colours”, “free returns within X days”, warranty language, shipping conditions. These templates produce high TM match rates, and the resulting savings across a catalogue of thousands of SKUs can be substantial.

Ongoing client relationships

The longer a translator or agency works with a specific client, the more the TM reflects that client’s preferences, approved terminology, and house style. Each new project benefits from the accumulated decisions of every previous one. This is one of the practical arguments for translation continuity rather than re-tendering work each time.

Where translation memory has real limits

TM is built for repetition. Applied to content that should not repeat, it creates friction rather than value.

Marketing and brand copy

Good marketing copy deliberately avoids repetition. Campaign-critical content – headlines, taglines, CTAs, emotionally driven body copy – is written to create specific effects, and applying a TM too aggressively produces language that is technically consistent but creatively flat. Our marketing translation work uses TM selectively, primarily for product names and recurring informational sections, not for the copy that needs to land in a specific way.

SEO translation

SEO translation requires the flexibility to phrase the same concept differently across pages to target different keyword variants. TM’s consistency enforcement works against this. Title tags and meta descriptions need to be written fresh for each page, calibrated to local search intent and character limits. As our SEO translation guide explains, CAT tools should support SEO translation workflows, not drive them.

Literary and creative translation

A phrase recurring in a novel may need a different translation on its second appearance because context, character, or tone has shifted. TM works at the segment level and cannot account for this kind of narrative development.

Short or one-off projects

If a document is unique and unlikely to generate similar future work, the setup overhead outweighs any benefit. TM needs volume and continuity before it delivers a meaningful return.

How to build a translation memory from scratch

Starting fresh

If you have no prior translations to draw from, the TM starts empty. The first project builds the initial database. On its own, an empty TM provides no matches at the start – the value accumulates with every job completed.

For clients starting new multilingual programmes, this means the first large project is the most labour-intensive and expensive. Subsequent projects in the same domain and language pair will see increasing TM leverage, and costs typically fall over time.

Populating a TM from existing translations (alignment)

If you have previous translations – documents translated in the past, perhaps before a TM was in place – these can be imported into the TM using a process called alignment. Alignment tools match source and target documents segment by segment, creating translation units that can be loaded into the TM database.

The quality of the resulting TM depends on the quality of the original translations. Importing unreliable or poorly edited legacy material will produce poor suggestions, which waste translator time and may introduce errors if accepted uncritically. It is usually worth reviewing aligned content before committing it to the live TM.

Sometimes, duplicate efforts are inevitable. For example, team members from in-country offices may work on translations independently without using a CAT tool or TM, a translation vendor might fail to deliver a promised TM, or quality assurance (QA) checks may identify errors in a TM that need correcting.

In such cases, you may need to align your TM with new or updated translations. This process involves taking a new translation and matching it up with the corresponding source text in your TM. Once they’re aligned, you can then update your TM with the new translation.

There are various software tools available that can help with TM alignment, or you can align TMs manually. However, it’s generally best to leave this task to a professional if possible – they charge this work by the hour and the investment will save you a lot of time and effort in the long run.

The process generally involves uploading a set of source files and target files – in the same format – into the tool, linking them by file name, and running an automatic alignment. Once the alignment is complete, you can then download the updated TM and use it in your CAT tool as normal.

TMX format

TMs are exported and imported in TMX (Translation Memory eXchange) format, an open XML-based standard supported by all major CAT tools and TMS platforms. This makes TMs portable: a TM built in Trados can be imported into memoQ, Phrase, Lokalise, or any other platform. Regardless of which tool or provider manages your TM, you should always be able to export it in TMX format.

A TM built from your content is a proprietary asset. It encodes years of approved terminology, regulatory language, and editorial decisions specific to your organisation. It should be yours to take with you if you change provider.

Building a TM that stays useful: quality and maintenance

A TM is only as good as the content in it. Poor inputs produce poor suggestions, and a poorly maintained TM becomes a liability.

Start with clean source content

If the source text uses multiple phrasings for the same concept, the TM stores multiple translations and has no basis for choosing between them. Consistent, well-edited source copy – clear sentences, stable terminology – produces a more useful database.

Separate TMs by content type

UI strings, legal clauses, marketing copy, help documentation, and technical manuals behave very differently. A suggestion pulled from a regulatory document and appearing in a product description may be technically accurate but completely wrong in register. Keeping TMs segmented by content domain prevents this bleed.

Use QA before committing to the TM

Most professional workflows include automated quality assurance (QA) checks before new translations are committed to the database. These checks catch formatting errors, missing terms, numerical inconsistencies, and tag errors. Committing unreviewed or poorly QA’d translations to the TM permanently embeds those errors into future suggestions.

Maintain and clean regularly

Over time, TMs accumulate outdated segments: obsolete product names, discontinued features, superseded regulatory language, old brand copy from campaigns no longer active. These produce suggestions that translators consistently reject, which slows workflows and erodes trust in the database. Scheduling periodic clean-ups – archiving obsolete segments, resolving duplicates, removing low-quality entries – keeps the TM accurate and fast. Lokalise recommends making this a routine part of any localisation programme, not a one-off task.

Keep TM and MT separate in quality terms

If machine translation output is fed into a TM without human review, the database gets populated with unreviewed content. Future TM matches then surface MT-generated text as if it were approved human translation. Most platforms allow you to flag MT-derived segments separately or require review before they are committed – use these features.

Translation memory and AI: what is changing

AI has changed the role of TM in specific ways without displacing it.

Large language models produce more contextually fluent MT output than the statistical MT systems of a decade ago, which narrows the quality gap between a good fuzzy TM match and a fresh AI-generated suggestion. For content types where quality requirements are modest and repetition is low, this makes TM less decisive than it once was.

But for clients who most need TM – in legal, pharmaceutical, regulatory, and technical fields – the picture is different. These workflows require human-approved language that reflects the client’s specific regulatory context, approved terminology, and established phrasing. An AI model trained on general internet data cannot replicate the specificity stored in a well-maintained client TM. The TM’s function as a repository of approved, domain-specific, client-vetted translations remains as relevant as it was.

AI is also being applied to TM maintenance itself: some platforms now use AI to identify obsolete segments, flag inconsistencies, suggest TM clean-up priorities, and predict which new content is likely to produce high TM leverage. This is a genuinely useful development that reduces the manual overhead of keeping a large TM accurate.

According to Smartcat research, using translation memory leads to a 10–60% increase in translator productivity – a wide range that reflects how much the benefit depends on content type and match rates.

A note on TM discounts

When a project contains a high proportion of TM matches, many translation agencies apply a discount structure: exact matches charged at a reduced rate, fuzzy matches on a sliding scale, new content at the full rate. The logic is that pre-matched segments require less translator effort.

The practice is widespread but genuinely contested. A match is a suggestion, not a finished translation. The translator still reads every matched segment, confirms it makes sense in context, checks that changes elsewhere in the document have not affected meaning, and edits where necessary. A 95% fuzzy match might need one word changed; it might need substantial reworking if the surrounding content has shifted. Neither the review work nor the judgement involved is captured in the match percentage.

Pricing structures that price matched segments at zero or near-zero tend to undervalue this work in ways that affect quality over time: translators under time pressure are incentivised to accept suggestions quickly rather than review them carefully.

At Crisol, we do not price by TM match rates. Where TM genuinely reduces workload on a large ongoing project, we structure volume arrangements instead. Our translation pricing guide covers this in more detail.

Work with a team that uses TM properly

If you produce content regularly for international markets – product documentation, software strings, compliance materials, ongoing website copy – a well-managed TM is worth establishing from the start. Each approved translation is one fewer decision to make again, and the database compounds in value with every project.

The difference between a TM that saves money and one that creates problems is almost always in the setup, the quality controls, and the maintenance discipline. A TM that is populated from clean content, segmented by content type, maintained regularly, and managed by translators who know when to follow a suggestion and when to override it will deliver real returns. One that has been built up without those habits will eventually slow projects down.

If you would like to discuss how a TM workflow could work for your content, get in touch.

Frequently asked questions

What is a translation memory (TM)? A translation memory is a database that stores pairs of source and target language text segments. When similar or identical content appears in a new project, the software retrieves the stored translation and offers it as a suggestion. The translator accepts, edits, or replaces it.

How does a translation memory work? The CAT tool or TMS splits the source text into segments and compares each one against the TM database. Identical segments produce exact matches; similar segments produce fuzzy matches with a percentage score. New content with no match is translated from scratch and added to the TM on completion.

What is a fuzzy match? A fuzzy match is a stored segment that is similar but not identical to the current source segment, expressed as a percentage. A 90% match shares approximately 90% of its content with the stored segment. The translator edits the suggestion to handle the difference. Most tools only surface matches above a threshold of around 70–75%.

What is the difference between a context match and an exact match? An exact match (100%) is word-for-word identical to a stored segment, though the surrounding content may differ. A context match (101%) is an exact match where the surrounding segments are also identical, giving the system high confidence the suggestion is correct in context. Context matches are typically auto-inserted.

What is the difference between a translation memory and a glossary? A TM stores complete translated segments and works at the sentence level. A glossary or termbase stores individual approved terms and enforces consistent use across all content. Both are used together: the TM handles familiar phrasing; the termbase handles critical vocabulary.

Is translation memory the same as machine translation? No. A TM retrieves stored human-approved translations. Machine translation generates new translations algorithmically. In modern workflows, TM is checked first; MT is often used as a fallback for segments with no useful match.

Does translation memory reduce costs? It can, particularly on high-repetition projects: technical manuals, software localisation, ongoing compliance documentation, large e-commerce catalogues. The saving depends on how much content genuinely repeats, whether a relevant TM already exists, and how TM matches are priced by the provider.

What is TMX format? TMX (Translation Memory eXchange) is the standard file format for importing and exporting TMs, supported by all major CAT tools and TMS platforms. It makes TMs portable between tools and providers.

How do I create a translation memory? A TM starts empty and grows with each completed project. If you have legacy translations, these can be imported using an alignment tool, which matches source and target documents segment by segment. The quality of the resulting TM depends on the quality of the source material.

Who owns a translation memory? In most professional arrangements, the client owns the TM built from their content; the agency manages and maintains it. Confirm this explicitly and ensure you can export your TM in TMX format at any time. A TM built from your content is a proprietary asset – it should travel with you if you change provider.

Does TM work well for marketing content? Generally not for primary marketing copy, which is written to vary and achieve specific creative effects. TM is most useful for repetitive, consistency-sensitive content. Applying it too aggressively to brand or campaign copy tends to produce language that is technically consistent but creatively flat.

What is TM leverage or TM analysis? A TM analysis (also called a leverage report) is a pre-project scan of the source file against the TM database. It shows the breakdown of exact matches, fuzzy matches by band, and new content, giving project managers and clients a realistic picture of how much new translation work the project actually requires before work begins.

Related article: What Everybody Knows, Many Call Out, and Very Few Do Something About: The Ugly Reality of the Translation Industry


Maria Scheibengraf

Author: Maria Scheibengraf

Maria Scheibengraf is an award-winning marketer, multilingual SEO specialist, and English-to-Spanish translator specialised in software, including SaaS, martech, and fintech. She is the co-founder and Operations Manager of Crisol Translation Services and the author of The SEO Translation Bible.

Crisol Translation Services SaaS Marketing & SEO Translation Services - Team members hero 2

Hello!

We’re Crisol. We’re award-winning marketing and SEO translation specialists for the SaaS, hospitality, food, education, and wellness sectors.

We help big brands and small businesses to produce locally persuasive and SEO-friendly marketing content across markets, to increase conversions, reduce churn, and engage global customers like never before.

And we’re ready to help you, too! Hire one of us as a freelancer or all of us as a boutique agency.

Take a look at our latest blogs:

Scroll to Top