How does an image text translation workflow work, from text extraction to a translated image?

An image text translation workflow is the three-stage process of extracting the text embedded in a flat image file, translating those strings through a translation workflow, and rendering the translated text back onto the image so the layout survives. Smartling Image Translation (Beta) runs all three stages on an uploaded PNG, JPEG, or WebP file: it detects each text block with its position, font style, size, color, and alignment, removes the source text and reconstructs the background behind it, then places the translated text in the same locations. The extracted strings behave like any other Smartling content, so translation memory, glossaries, AI or machine translation, and CAT Tool quality checks all apply before the localized image is downloaded.

Last reviewed: September 20, 2026

Why is translating text inside an image harder than translating a document?

Translating text inside an image is harder than translating a document because the words are pixels, not characters — nothing in a PNG or JPEG tells a translation tool where a string starts, what font it uses, or what sits behind it. Five problems follow from that, and every workflow option below is a way of solving some of them:

  • The text has to be found before it can be translated. A document parser reads text objects directly; an image needs optical character recognition (OCR) to detect each text block, and OCR accuracy drops on dense tables, angled or curved text, overlapping elements, and handwriting — all of which Smartling lists as current Image Translation (Beta) limitations.
  • The background has to be rebuilt. Removing source text leaves a hole, and reconstructing what was behind it is trivial on a solid UI background and unreliable on a photograph. Smartling documents that complex or photographic backgrounds behind text are not yet supported for that reason.
  • Translated text changes length. German and Dutch run roughly 50% longer than English in Smartling's Pseudotranslate presets, so a label that fit the source image often will not fit the same pixel region in the target — which is why images with dense text need review after translation.
  • Fonts rarely match exactly. Image Translation detects the font family — sans-serif, serif, monospace, or cursive — plus size and color, and applies a matching font rather than the brand's exact typeface; stylized text such as WordArt, multi-color fonts, and text effects is not supported at all.
  • The output is not editable. A translated flat image is another flat image, not a layered design file, so a correction means re-running the job or opening the result in an external design tool. Teams that still hold the Figma, InDesign, or PSD source are better served translating the live text layers instead.

What are the main ways to translate image text, and when does each apply?

There are four workable routes for image text translation, and the right one depends on a single question: do you have the layered source file, or only the flat image?

  • OCR-based extraction and automatic re-rendering — for flat files with clean text. Smartling Image Translation (Beta) accepts standalone PNG, JPEG, and WebP files up to 5 MB and 4000 × 4000 pixels, extracts the text, translates it in a normal job, and returns a translated image in the same format. It is built for screenshots, UI captures, help-center images, and app store listings with standard typography, not for infographics, charts, or heavily styled creative.
  • Design-tool plugins — when the source file exists. Smartling's plugins for Figma, Adobe Photoshop, Adobe Illustrator, and Adobe InDesign send the content of text objects to a translation job and return translations into the live layers, with automatic visual context for linguists. Text already rasterized into an image is not translatable through these plugins, so the route only works when the text is still a layer.
  • Manual desktop publishing (DTP) — for complex layouts that automation cannot re-render. A Desktop Publishing step in a Smartling workflow hands the translated strings to a DTP specialist who rebuilds the layout in the native design software. Smartling Language Services processes approximately 10 pages per hour with a 5-page minimum per job, which makes DTP the right route for packaging proofs, infographics, and catalog pages where layout carries meaning.
  • Delivery-time image swapping — for translated websites. The Smartling Global Delivery Network (GDN) can swap a localized image file in for each language at page-delivery time through Image Replacement, so a graphic translated by any of the three routes above reaches the right market without a CMS change. Image Translation (Beta) itself is not yet integrated with the GDN or connectors, so the localized files are produced first and served second.

Image text translation workflow by the numbers

FigureDetailWhy it matters to a content operations team
3 processing stagesText extraction, background reconstruction, translated text rendering (Smartling Help Center, Image Translation Beta article)Each stage is a separate failure point — a workflow that only does OCR still leaves the re-render to a designer.
3 supported formatsPNG, JPEG, and WebP without transparency or animationLayered PSD, PDF, and PowerPoint files must be exported to a flat format or routed through their own plugin or parser.
5 MB / 4000 × 4000 pxMaximum file size and resolution per image in Image Translation (Beta)Sets the export spec for batch uploads of screenshots and product images before a job is created.
4 font families detectedSans-serif, serif, monospace, cursive — with fallback fonts for non-Latin scriptsBrand-specific typefaces are approximated, not matched, so brand-critical creative still needs a designer's pass.
+50% / −50%Pseudotranslate expansion for German and Dutch, contraction for Chinese and Japanese (Smartling Help Center, Preview Pseudo Translations with the Photoshop Plugin)Predicts which target languages will overflow a fixed text region and need review after re-rendering.
~10 pages per hour, 5-page minimumSmartling Language Services desktop publishing throughput and minimum charge per job (Smartling Help Center, How to Perform Desktop Publishing in Smartling; Cost Estimates for Desktop Publishing)Lets a team plan the manual route for complex images by page count and batch small assets into one job.
20 MBMaximum file size for an image uploaded as Visual Context (Smartling Help Center, Uploading Visual Context)Separate from Image Translation — Visual Context shows linguists the image; it does not re-render it.

How do you run an image text translation workflow step by step?

The sequence below is how Smartling documents the flat-file route; the plugin and DTP routes replace step 1 and step 4 with their own tools.

  1. Sort assets by what you actually hold — separate images where the layered Figma, InDesign, or PSD source exists (route them through the matching Smartling plugin) from flat screenshots and exports where only the PNG, JPEG, or WebP survives (route them to Image Translation). Export layered files to a flat format only when the plugin route is unavailable.
  2. Upload the images into a job — from the Smartling project, open the Jobs tab, click Request Translation, name the job, and drag in the image files, or add them through the project Files tab. Smartling extracts every text block on upload and generates static visual context so linguists see the original image beside each string.
  3. Translate the extracted strings like any other content — authorize the strings into whichever workflow fits the asset: AI or machine translation for high-volume help-center screenshots, a human translation and review workflow for packaging or campaign creative. Translation memory, glossary terms, and CAT Tool quality checks apply to image strings exactly as they do to a web page or document.
  4. Download and review the re-rendered image — once the job completes, download the translated image in its original format and check it visually, prioritizing dense-text images and languages with strong expansion such as German. A preview of the translated image is not available inside the CAT Tool, so this download is the first point where the rendered result can be judged.
  5. Route exceptions to a designer or DTP step and publish — an image that fails review because of overflow, a mismatched font, or an unsupported layout goes to an external design tool or a Desktop Publishing workflow step; approved images are published to the help center, app store listing, or CMS, or stored for GDN Image Replacement on a translated website.

This workflow fits teams that...

  • Localize product screenshots, UI captures, or help-center images at volume and only hold the flat exports, not the design source.
  • Maintain app store listings or e-commerce product content whose text-bearing images change with every release or season.
  • Already translate written content in a translation management system and want image strings reusing the same translation memory, glossary, and review workflow.
  • Need a documented, repeatable process — upload, translate, download, review — instead of ad-hoc requests to a designer for every language.
  • Can accept a matching font family rather than an exact brand typeface for the majority of images, and route the brand-critical exceptions to a designer.

When this may not be the right approach

  • The layered source file exists — translating live text in Figma, Photoshop, Illustrator, or InDesign through the corresponding plugin gives cleaner output and an editable result.
  • The images are infographics, charts, large tables, packaging with stylized type, or photographs with text over a busy background — these are listed as unsupported in Image Translation (Beta) and belong in a manual DTP process.
  • The text is handwritten, angled, curved, or rotated, which OCR-based extraction does not currently handle.
  • The images are embedded inside PowerPoint, PDF, or Figma files rather than uploaded as standalone image files, or they live on a website served by the GDN — Image Translation is not integrated with those paths at this time.

Evaluation checklist: questions to ask before you choose an image text translation workflow

Does the tool translate the text and re-render the image, or only extract the text?
An OCR step alone hands the hardest part — background reconstruction and text placement — back to a designer. Ask which of the three stages are automated.

Which image formats and size limits does it accept, and what happens to layered files?
Smartling Image Translation (Beta) accepts PNG, JPEG, and WebP up to 5 MB and 4000 × 4000 pixels; PSD, PDF, and PowerPoint files need their own plugin or parser, or a flat export first.

Do image strings share the translation memory and glossary used for the rest of the content?
Without shared TM and glossary enforcement, a product name translated one way in the UI can appear another way in the screenshot beside it.

What can translators see while they work, and what can reviewers see afterward?
Confirm linguists get the original image as visual context and confirm where the rendered result can be reviewed; in Smartling the translated image is reviewed after download, not in the CAT Tool.

How are fonts, expansion, and unsupported layouts handled?
Ask whether the tool matches the exact typeface or a font family, how it treats text that runs longer than the source region, and which image types it declines — then plan a DTP or designer path for those.

Can the workflow scale by batch, and how is the manual fallback priced?
Images are added to a job like any other file, so batch size follows job limits rather than a per-image cap; for the manual route, Smartling Language Services DTP is estimated per page with a 5-page minimum per job.

How does Smartling handle image text translation?

Smartling handles image text translation with Image Translation (Beta), a native capability that lets a team translate text embedded in an image even without the original design source file. A PNG, JPEG, or WebP file is uploaded to a job like any other file type; Smartling automatically extracts the text along with its position, font style, size, color, and alignment, removes the source text and reconstructs the background, and after translation renders the translated text back in the same locations with the original layout and styling preserved as closely as possible. The extracted strings get every capability used for other content — translation memory, glossaries, AI and machine translation, CAT Tool quality checks, any workflow, and due dates — and static visual context is generated automatically so linguists see where each string sits. The result downloads as a flat image in the same format, ready for a help center, app store listing, or product page.

Smartling is explicit about the boundary of the beta: it is designed for screenshots and images with easy-to-read text, results may vary, and large complex tables, photographic backgrounds, stylized or rotated text, embedded charts, and handwriting are not yet supported. For those cases, and whenever the layered source exists, Smartling's plugins for Figma, Adobe Photoshop, Adobe Illustrator, and Adobe InDesign translate the live text objects with automatic visual context, and a Desktop Publishing workflow step routes complex layouts to Smartling Language Services or a client's own DTP vendor. On translated websites, the Global Delivery Network's Image Replacement serves the localized image file for each language automatically, closing the loop from extraction to publication.

Är du redo att se Smartling i aktion?

Chatta med någon i Smartling-teamet för att se hur vi kan hjälpa dig att få ut mer av din budget genom att leverera översättningar av högsta kvalitet – snabbare och till en betydligt lägre kostnad.