Are my PDF, images, or recognized text uploaded?
No. File parsing, page rendering, recognition, and export all happen in your browser. On first use, the browser only downloads OCR model and runtime files from this site and caches them when possible; your document content is never sent with those requests.
How is a searchable PDF different from a normal scanned PDF?
A normal scanned PDF usually contains only page images. A searchable PDF keeps those visuals and adds an invisible text layer at the recognized positions, making text searchable, selectable, and copyable.
Which languages and files are supported?
The first version targets Chinese, English, and mixed Chinese-English content in scanned PDFs, PNGs, JPGs, and WebP images. Handwriting, decorative fonts, vertical text, and complex tables may be less accurate.
Is OCR always accurate?
No. Resolution, skew, shadows, compression noise, stamps, and complex layouts all affect accuracy. Search for a few key terms and review copied text after export; important contracts, identity documents, and legal records still need human verification.
Why can the first OCR run take longer?
The OCR model and browser runtime are larger than ordinary page assets, so the first run must download them from this site. Later visits can usually reuse the browser cache unless site data has been cleared.
What happens with a long PDF?
OCR processes pages sequentially to limit memory use and applies page and pixel caps to protect phones and lower-memory devices. Split long documents first and keep the tab open while recognition is running.