Skip to content

Blog

How-to16 min read4,071 views

How to Translate a Scanned PDF Online (OCR + Keep Layout)

Translate a scanned PDF by confirming there is no text layer, running OCR, then using a layout-aware PDF translator. 300 DPI checklist, Google limits, and when scans still fail.

How to Translate a Scanned PDF Online (OCR + Keep Layout) - PDFTranslations guide
Layout-focused PDF translator

How do you translate a scanned PDF without losing the layout?

You translate a scanned PDF in two stages: optical character recognition to create a text layer, then document translation that tries to keep headings, tables, and reading order. A translator cannot translate pixels it cannot read.

Quick test: open the file and try to select a sentence. If nothing highlights, you have an image-only page. If text selects but looks like gibberish, you may have a bad existing OCR layer and should re-run OCR.

PDFTranslations is built for layout-preserving translation on files that already have a text layer. Complimentary first page is the honesty check. Complex scans, handwriting, and low-resolution photos can still degrade - that limit is real, which is why this guide exists.

Why scanned PDFs break Google Translate and chat tools

Google Translate document upload expects text. Image-only pages often come back blank, untranslated, or with the original photograph untouched. Chat tools can describe a screenshot, but they do not rebuild a paginated PDF with stable tables.

DeepL-class document translation is strongest when the OCR layer is already clean. A noisy scan becomes a noisy translation. Details: does DeepL keep PDF formatting? and why Google Translate destroys PDF formatting.

Scan quality checklist (do this before any OCR)

Resolution: 300 DPI minimum for body text, 400-600 DPI for small footnotes. A 150 DPI phone photo will hallucinate characters.

Geometry: page straight, all four corners visible, no thumb over the margin, even light, no flash glare on glossy paper.

Content: one page per image when you can. Avoid bound-book curves. Color vs grayscale matters less than sharpness.

If you do not control the scan (a vendor sent a fax PDF), say so in your review notes. You will spend more time checking names and numbers.

Three OCR paths

Desktop: Adobe Acrobat Recognize Text, ABBYY FineReader, or another local OCR tool. Best for confidential files your policy keeps off random websites.

Drive trick for non-sensitive files: upload to Google Drive, open with Google Docs, let Google OCR the page, then export a cleaner PDF or copy. Do not use this on privileged contracts.

All-in-one online translators that advertise built-in OCR. Convenient. Still verify page one. OCR errors become translation errors and are harder to see once the language changed.

After OCR: translate as a real PDF

  1. 1

    Confirm text is selectable on the page you care about.

  2. 2

    Set source language by hand if OCR language was wrong (common on mixed Latin plus CJK or Arabic).

  3. 3

    Review names, numbers, stamps, and table totals on page one.

  4. 4

    Re-run OCR if page one shows garbage characters. Do not buy a pack hoping later pages are cleaner.

  5. 5

    Keep the original scan. You will need it to spot-check figures.

Once you can select text, upload the file to a layout-aware translator. On PDFTranslations: upload, pick the target language, review complimentary first page, then unlock remaining pages on pricing only if cells and columns hold.

When OCR plus translation still fails

Handwriting, stamps, seals, low-contrast copies, and ornate certificates. Machine output is a draft. It is not a certified extract of an identity document or a court filing.

For those files, hire a human who can see the paper. Use the machine path only to understand what the document is about, if policy allows the upload at all (safety checklist).

Layout after OCR is still a separate problem

OCR gives you characters. It does not give you a perfect reading order. Two-column scans can OCR left-to-right across both columns. Fix that with a layout-preserving translator, or accept a text dump and rebuild.

See layout-preserving vs text-only.

Start with the selectable-text test

If your PDF already has a text layer, skip OCR and go straight to the upload card. If it does not, OCR first, then upload. Do not pay for a pack on a photograph of a page.

Frequently asked questions

Can I translate a scanned PDF for free?

You can evaluate page one on PDFTranslations after the file has a text layer. Pure photos still need OCR first. Free is not unlimited pages.

Can Google Translate translate a scanned PDF?

Usually not until OCR creates text. Blank or unchanged pages are the typical result.

Does PDFTranslations OCR every scan automatically?

The product is strongest on text-based PDFs. Prefer OCR on image-only pages before upload, then use the free first page as the quality gate.

What DPI should I scan at?

300 DPI is the practical minimum for body text. Small print needs more.

Is a translated scan legally valid?

No. Machine output is not certified, sworn, or notarized translation.

Translate a PDF now

Keep the layout. 133 languages. Pay per file or pick a monthly plan when you need more.

Updated 2026-08-29

See pricing

Keep reading