Processed locally · No upload

Local PDF OCR

Recognize Chinese, English, and mixed text in scanned PDFs, PNGs, JPGs, and WebP images locally in your browser. Export a searchable PDF that keeps the page visuals, or download plain text. Neither files nor recognized text are uploaded.

Processed locally, never uploaded. The file is passed only to the OCR workflow in this browser.

Make a searchable PDF in three steps

  1. 1Choose a scanAdd a PDF, PNG, JPG, or WebP. It opens only on your current device.
  2. 2Recognize locallyRun mixed Chinese-English OCR page by page to recover text, positions, and confidence.
  3. 3Download and reviewExport a searchable PDF or TXT, then review important text for accuracy.

The model downloads; your document does not upload

OCR first loads model and runtime files hosted by this site; these are generic recognition assets. Your PDF, images, page pixels, recognized text, and exports remain in browser memory. Once the model is loaded, pages do not need to be sent to a remote OCR API.

Searchable does not mean error-free

The OCR text layer is useful for search, copying, and indexing, but it is not an authoritative transcription without review. Blur, skew, low contrast, handwriting, stamps, and tables can produce wrong characters or reading order. Search for key terms after export and compare important content with the original page.

Frequently asked questions

Are my PDF, images, or recognized text uploaded?

No. File parsing, page rendering, recognition, and export all happen in your browser. On first use, the browser only downloads OCR model and runtime files from this site and caches them when possible; your document content is never sent with those requests.

How is a searchable PDF different from a normal scanned PDF?

A normal scanned PDF usually contains only page images. A searchable PDF keeps those visuals and adds an invisible text layer at the recognized positions, making text searchable, selectable, and copyable.

Which languages and files are supported?

The first version targets Chinese, English, and mixed Chinese-English content in scanned PDFs, PNGs, JPGs, and WebP images. Handwriting, decorative fonts, vertical text, and complex tables may be less accurate.

Is OCR always accurate?

No. Resolution, skew, shadows, compression noise, stamps, and complex layouts all affect accuracy. Search for a few key terms and review copied text after export; important contracts, identity documents, and legal records still need human verification.

Why can the first OCR run take longer?

The OCR model and browser runtime are larger than ordinary page assets, so the first run must download them from this site. Later visits can usually reuse the browser cache unless site data has been cleared.

What happens with a long PDF?

OCR processes pages sequentially to limit memory use and applies page and pixel caps to protect phones and lower-memory devices. Split long documents first and keep the tab open while recognition is running.