AI profile evaluation is available for plans that include custom AI profiles. This feature is currently in beta.
This article describes a feature in Lokalise Expert. This feature is not currently available in Lokalise Vantage.
Measure how closely AI translations match your past translations and prove the value of a custom AI profile using your own content.
AI profile evaluation compares three translation methods: your custom AI profile, the base profile (Pro AI), and standard AI or machine translation. Lokalise re-translates source text from your selected past translation examples, then compares each generated variant with your approved translations using similarity-based metrics.
The results provide both quantitative and qualitative proof. The dashboard shows how closely each method matches your reference translations, while side-by-side examples let you review the translations and judge their quality yourself.
AI profile evaluation supports custom AI profiles that use existing translations as context, including:
Profiles based on tagged keys or reviewed translations
TM-based AI profiles
It is not available for base AI profiles.
What AI profile evaluation is
AI profiles in Lokalise can use your past translations as context to help Pro AI follow your preferred terminology, wording, and style. Custom AI profiles can use context from tagged or reviewed translations, or from translation memories assigned to a project. The base AI profile does not use your past translations as RAG context.
AI profile evaluation provides proof of value by comparing newly generated AI translations with your approved past translations. It gives you a structured way to compare:
Your selected custom AI profile
Base AI profile — Pro AI without your past translations as context
Google Translate
This helps you determine whether your custom AI profile produces translations that are closer to your historical translation patterns than the base AI profile or standard machine translation.
Why use AI profile evaluation
Use AI profile evaluation to:
Check whether your custom AI profile improves translation consistency.
Compare its output with the base AI profile and Google Translate.
Prove the value of a custom AI profile using your own content.
Assess whether the context and reference translations are relevant and high quality.
Identify weaker translations by reviewing side-by-side examples.
Re-evaluate results after updating the profile, its context, or the reference data.
Evaluation helps you understand not only whether a custom AI profile performs better overall, but also which parts of your translation setup may need improvement.
How AI profile evaluation works
AI profile evaluation uses two sets of past translations:
The context set provides examples that the custom AI profile can retrieve through RAG when generating translations.
The reference set contains approved translations that Lokalise uses to evaluate the generated variants.
How these sets are created depends on the type of custom AI profile.
Profiles based on tagged or reviewed translations
When creating the AI profile, you select a source language, one or more target languages, and a data source such as tags or reviewed translations.
These profiles require at least 500 high-quality translation examples per language. Lokalise divides the selected examples into:
A context set used by the custom AI profile
A reference set of up to 200 translations used for evaluation
Because the profile can use the same tagged or reviewed translations across projects, the evaluation is not tied to a specific project.
TM-based AI profiles
For a TM-based profile, the evaluation is specific to a project.
The translation memories assigned to the selected project provide the context set. You then select high-quality translations from that project to use as the reference set. For optimal results, provide at least 200 reference translations per language.
The results apply to the selected project and may also be relevant to projects that use the same translation memories in the same priority order. Run a separate evaluation for projects with a different TM setup.
Translation and comparison
After the context and reference sets are prepared, Lokalise:
Takes the source text from the reference set.
Re-translates it to generate three variants:
The selected custom AI profile
The base AI profile
Google Translate
Compares each generated variant with the corresponding approved reference translation.
Calculates similarity-based metrics and displays side-by-side examples.
The results show how closely each translation method aligns with your past translations.
Before you begin
You need to be at least team admin to access this feature.
Before running an evaluation, make sure you have created a custom AI profile and prepared the required data. The requirements depend on how the profile uses past translations.
Profiles based on tagged or reviewed translations
Make sure that:
The profile uses tagged keys or reviewed translations as context.
The profile includes at least 500 high-quality translation examples for each language you want to evaluate.
The examples come from a source you trust and accurately represent your preferred terminology, wording, and style.
The target languages are configured in the profile.
Only languages that meet the minimum number of examples are available for evaluation.
TM-based AI profiles
Make sure that:
The profile uses translation memories as context.
The translation memories assigned to the project contain at least 500 high-quality entries for the relevant language pairs.
The project contains at least 200 high-quality translations per target language that can be used as the reference set.
The target languages you want to evaluate are available in the selected project.
The correct translation memories are assigned to the project in the intended priority order.
Evaluation results are specific to the selected project and the translation memories assigned to it. Run a separate evaluation for projects that use a different set or priority order of translation memories.
How to run an evaluation
Open AI profiles from the Lokalise side menu.
On the profiles page, custom AI profiles can show different states:
Evaluate — the profile has not been evaluated yet
Metrics — evaluation results are already available
The setup process depends on the type of profile.
Evaluate a profile based on tagged or reviewed translations
To run an evaluation:
Find the profile you want to evaluate.
Click Evaluate.
Review the evaluation overview.
Select one or more available target languages. Languages that do not have enough examples cannot be selected.
Start the evaluation.
During setup, Lokalise shows:
Which target languages are configured for the profile
Which languages have enough examples for evaluation
Which languages do not meet the minimum requirement of 500 examples
Lokalise evaluates each selected language separately.
Evaluate a TM-based AI profile
TM-based evaluation is specific to a project because the custom profile uses the translation memories assigned to that project as context.
To run an evaluation:
Find the TM-based profile you want to evaluate.
Click Evaluate.
Review the evaluation overview.
Select the project you want to evaluate.
Review the translation memories assigned to the selected project. These translation memories provide the context for the custom AI profile.
Select high-quality translations from the project to use as the reference set. You can select reviewed translations or translations associated with an available tag.
Select one or more target languages that have enough reference translations.
Start the evaluation.
During setup, Lokalise shows:
The translation memories assigned to the selected project
The selected source of reference translations
The number of available reference translations for each target language
Which languages meet the minimum requirement of 200 reference translations
The results apply to the selected project and its assigned translation memories. Run a separate evaluation for projects that use a different set or priority order of translation memories.
Evaluation progress
Once the evaluation starts, you will see its progress. The evaluation usually takes a few minutes, and you will receive an email notification when the results are ready.
After the evaluation finishes, the Evaluate button changes to Metrics.
Understanding the results
Click See metrics to view the evaluation results:
The overview compares three translation variants for the selected languages:
The Selected AI profile
The Base AI profile
Google Translate
All three variants are generated from the same source text and compared with the same reference translations. This lets you see which method produces translations that most closely match your approved past translations.
Review both:
Aggregated metrics, which provide a quantitative comparison across all evaluated translations
Examples, which show the generated variants side by side so you can assess their wording, terminology, and overall quality
The page also includes information about the evaluation run, such as:
The number of evaluated segments
The evaluated language
The source of the reference translations
The evaluation run
For TM-based evaluations, the run details also identify the selected project and the translation memories used as context.
Lokalise samples eligible translations for each evaluation, so the exact reference set may vary between runs. Use the dropdown at the top of the page to switch between evaluation runs and compare their results.
Sampling helps prevent the results from depending on a single fixed set of translations. Reviewing several runs can provide a more representative view of how the profile performs.
Metrics explained
AI profile evaluation currently uses similarity-based metrics. These metrics show how closely AI-generated translations match your approved past translations (the reference).
Perfect match
Perfect match shows the share of translations that match the reference exactly.
Higher is better. This is a strict metric, so even small differences in punctuation or wording reduce the score.
Translation edit rate (TER)
Translation edit rate shows how many edits are needed to turn the AI-generated translation into the reference, relative to the reference length.
Lower is better. Lower scores mean the translations are more similar. A score of 0 means the translations are identical.
BLEU
BLEU measures wording overlap between the AI-generated translation and the reference.
Higher is better. Higher values generally mean stronger overlap, while very low values often indicate substantial rewording.
ChrF
ChrF measures similarity between the AI-generated translation and the reference by comparing short character sequences.
Higher is better. Higher scores indicate greater similarity to the reference, while lower scores indicate more differences in wording and character sequences.
Important note about similarity-based evaluation
These metrics measure similarity to your past translations, not absolute translation quality. A higher score means that a generated translation is closer to the selected reference, but it does not necessarily mean that it is objectively better.
The reliability of the results depends heavily on the quality and relevance of the reference translations. The performance of the custom AI profile also depends on the translations provided to it as context.
For example:
If your reference translations are accurate and consistent, the evaluation can provide a useful signal.
If they contain errors, outdated wording, untranslated segments, or inconsistent terminology, the results may be misleading.
If the profile context contains irrelevant or low-quality translations, the custom profile may perform worse even when the reference set is reliable.
A lower similarity score does not automatically mean that the AI-generated translation is worse. It may also mean that:
The reference translations need to be reviewed or updated.
The selected tags or reviewed translations do not represent the desired style.
The translation memories assigned to the project contain outdated or irrelevant content.
The selected project does not provide a representative reference set.
Review the aggregated metrics together with the individual translations in the Examples view before drawing conclusions.
Act on evaluation results
Evaluation results help you understand whether a custom AI profile produces translations that are closer to your approved past translations and where your setup may need improvement.
Review both the aggregated metrics and the individual translations in the Examples view. Similarity scores provide a useful signal, but the examples help you determine whether the generated translations are actually suitable for your content.
The custom AI profile clearly performs best
The custom AI profile consistently achieves better similarity scores than the base AI profile and Google Translate, and the individual examples confirm that its translations better reflect your preferred terminology, wording, and style.
This indicates that the profile is providing measurable value for the evaluated content.
Recommended next steps:
Use the profile for the evaluated content.
Evaluate additional target languages or similar content types.
For profiles based on tagged or reviewed translations, consider using the profile with other projects that rely on the same context dataset.
For TM-based profiles, run separate evaluations for other projects, especially when they use different translation memories.
The custom AI profile and base AI profile perform similarly
The results are close, with no clear advantage for either profile.
This may mean that the base AI profile already performs well for this type of content. It may also indicate that the custom profile does not have sufficiently relevant or consistent context to produce a noticeable improvement.
To investigate and improve the results:
Review the individual translations to identify where the outputs differ.
Make sure the reference translations accurately represent the terminology and style you want to follow.
For profiles based on tagged or reviewed translations, refine the selected dataset by removing outdated, inconsistent, or irrelevant examples.
For TM-based profiles, review the translation memories assigned to the project and remove or update unsuitable content where possible.
Update the profile or its underlying data, then run another evaluation.
The custom AI profile performs well overall but struggles with specific content
A custom AI profile may achieve the strongest overall scores while still producing weaker translations for particular segments.
Use the Examples view to look for recurring issues, such as:
Missing or incorrect glossary terms
Outdated, inconsistent, or low-quality context translations
Reference translations that do not represent the desired output
Short or ambiguous source strings without enough context
Missing screenshots, key descriptions, or other contextual information
Improve the relevant terminology, context, or reference data, then run the evaluation again to see whether the changes produce better results.
The dashboard provides an overall comparison, while the individual examples help you identify where the profile works well and which parts of the translation setup need attention.
How and when to re-evaluate a profile
To start a new evaluation, click the More icon next to the profile name, then select Reevaluate.
Re-evaluation is useful when:
You updated the translations used as context or references.
You changed the data source, for example from tagged keys to reviewed translations.
You added, removed, or refined the tags used by the profile.
You cleaned up outdated, inconsistent, or low-quality translation examples.
You changed the translation memories assigned to a project.
You want to check whether changes to the profile or its underlying data improved the results over time.
For profiles based on tagged or reviewed translations, re-evaluate after changing the dataset used as context.
For TM-based profiles, re-evaluate after updating the project’s translation memories, changing their priority order, or selecting a different set of high-quality reference translations.
Each evaluation creates a separate run. You can switch between runs on the results page to compare how the profile performs before and after your changes.
Best practices
To get more useful evaluation results:
Use accurate, consistent, and up-to-date translations as references.
Select data that represents a clear content type, domain, terminology, and tone.
Avoid untranslated, outdated, inconsistent, or low-quality translations.
Make sure each target language has enough representative examples.
Review the metrics separately for each language, as profile performance may vary between languages.
Check individual translations in the Examples view instead of relying on aggregated metrics alone.
Re-run the evaluation after improving the profile context or reference data.
For profiles based on tagged or reviewed translations, carefully select the tags, projects, or reviewed content used by the profile.
For TM-based profiles, review both the high-quality translations selected as references and the translation memories assigned to the project. Outdated or irrelevant TM content can reduce the effectiveness of the custom profile even when the reference translations are reliable.
Useful evaluation results depend on relevant, representative, and high-quality context and reference data.
Known limitations
AI profile evaluation is available only for custom AI profiles that use past translations as context.
You cannot start an evaluation for the base AI profile itself. The base profile is included automatically as a comparison variant when you evaluate a custom AI profile.
Profiles based on tagged or reviewed translations require at least 500 examples per language.
TM-based evaluations require at least 200 high-quality reference translations per language for optimal results.
Only languages with enough eligible translations appear in the evaluation flow.
TM-based evaluation results are specific to the selected project and its assigned translation memories.
Results are based on sampled translations, so the evaluated set and scores may vary between runs.
Current metrics measure similarity to reference translations, not absolute translation quality.









