Documents & PDF tool

PDF & Image to Text

Choose PDF pages or a JPEG/PNG, extract their text privately on your device, compare it with the original and copy or download a plain-text file.

Turn a document or photo into text

Copy text from PDFs and read printed English from scans, JPEGs and PNGs. Your files are processed on this device, not uploaded.

Check every result. Handwriting, small print, blur and complex columns can produce mistakes. Automatic mode may miss text inside images on a PDF page that also has selectable text; choose “Scan the whole page” in that case. This creates plain text, not a searchable PDF.

Original page

Your page preview will appear here.

English OCR engine files download from this site on first use; documents and results stay in memory and are not saved when you leave. Keep this tab open while processing.

Processed privately in your browser — your entries are not uploaded.

Good to know

How this tool works

Choose a PDF, JPEG or PNG. For a PDF, enter the pages to extract. Automatic mode copies existing page text where available and scans pages without text. Use the whole-page OCR option for incomplete text layers. Review each completed page, then copy or download the text.

Frequently asked questions

What is PDF & Image to Text useful for?

Extract PDF text or read printed English from scans and photos. Open the downloaded document and check every page, especially page order, clipping and readable text.

Is my information stored?

Your text is processed locally in a browser worker. It is not uploaded by this tool or saved in browser storage. Clear releases the working input and result. Downloads remain in your normal download folder.

How accurate is OCR?

OCR recognises printed English, but is not guaranteed to be correct. Check names, numbers, punctuation and text order against the original. Handwriting, blur, sideways text, tables and multiple columns are harder to read. Use a sharp, upright scan. This version does not offer automatic rotation or translation.

Why is some text missing or in the wrong order?

Automatic mode uses the existing text layer when a PDF page has one. It can miss words inside images on that same page, and the text layer may be incomplete or out of reading order. Select whole-page OCR to scan the visible page instead. Plain text does not preserve fonts, tables or page layout. Comments, form values and hidden text may not match the visible page; check results carefully.

Are files uploaded or saved?

No document content is sent to a server. PDF rendering and OCR run on your device. The English OCR engine downloads from this site, without sending it your files. Results are held in memory; download them before leaving. Password-protected PDFs are not supported: choose an unprotected copy you are authorised to use.

Can I create a searchable PDF?

Not with this version. It exports a plain-text file and leaves your original document unchanged. PDFs can contain up to 200 pages; process up to 20 selected pages per run. Pages and images are rasterised at a bounded resolution for browser memory, which can affect very small text.

OCR powered by Tesseract.js; PDF rendering by PDF.js.