Skip to content
Toolbase

Image to Text (OCR)

Extract text from images and screenshots in English, Urdu, Arabic, Hindi and Chinese with on-device OCR. Copy or export TXT. Free and private.

Your images never leave your browser
BulkOn-device AINo upload
Loading tool…

How to use the Image to Text (OCR)

  1. 1Drop one image or a batch of screenshots, photos or scans.
  2. 2Choose the language: English, Urdu, Arabic, Hindi or Chinese (or more than one).
  3. 3Click Extract and wait for the on-device OCR to finish.
  4. 4Review the text in the editable panel.
  5. 5Copy it or export a TXT file.

Extract text from images and screenshots, on your device

Text arrives as pictures all day: a supplier's price list photographed in a warehouse, a WhatsApp screenshot of a customer's address, a competitor listing you want to quote, an Urdu invoice or an Arabic product label. Retyping is slow and error-prone. Image to Text (OCR) reads the text out of the image with Tesseract running entirely in your browser and lets you copy it or save a TXT file. Nothing is sent to a cloud OCR service.

How it works

  1. Drop one image or a batch. Screenshots, phone photos and scans all work; PNG and JPG are best.
  2. Choose the language of the text: English, Urdu, Arabic, Hindi or Chinese. Pick two if the image mixes scripts, for example English and Urdu on a Pakistani invoice.
  3. The first time you use a language, its recognition pack (a few megabytes) is downloaded from this site and cached by your browser. After that it works offline.
  4. Click Extract. A progress bar shows each image being processed in a Web Worker, so the page stays usable.
  5. Review the text in the editable panel next to the image, then Copy or Export TXT. Batches export one TXT per image or one combined file.

Why it matters

  • OCR accuracy depends on resolution. Tesseract works best when letters are at least 20 px tall; on a phone photo that means shooting close enough that a line of text spans a good part of the frame. A 1080p screenshot of a listing is usually ideal.
  • Right-to-left scripts are handled by the Urdu and Arabic packs, which recognise printed text in common fonts. Handwriting is not supported by any pack.
  • Sellers use it to move data out of images and into spreadsheets: 200 SKU codes from a photographed catalogue, product descriptions from a manufacturer's PDF page, or addresses from order screenshots, without retyping.
  • Because it runs on-device, customer data in order screenshots never leaves your machine, which matters under the buyer-information rules on Amazon and other marketplaces.
SourceExpected accuracyWhat helps
Clean screenshot98% or higherNothing, just run it
Flat scan, 300 DPI95% or higherBlack and white filter
Phone photo of paper85 to 95%Straighten first, good light
Photo of a product label70 to 90%Crop to the label, increase size

Tips

  • Run document photos through the Document Scanner first. A straightened, black and white page can double OCR accuracy compared with a raw photo.
  • Crop to the text you need. Backgrounds, logos and product photos confuse layout detection and add junk to the output.
  • If a PDF holds the text, split it into 300 DPI page images with Image to PDF / PDF to Images and OCR those.
  • Check digits carefully. OCR confuses 0 and O, 1 and l, 5 and S, especially in SKUs; a quick find and replace usually fixes a whole batch.
  • An extracted URL or message can go straight into the QR Code Generator for packaging inserts.

Privacy

Your images never leave your browser and OCR runs on-device. The language packs are served from this site, the recognition happens in a Web Worker on your computer or phone, and the text you extract is never transmitted anywhere.

Frequently asked questions

Which languages does the OCR support?

English, Urdu, Arabic, Hindi and Chinese, using Tesseract language packs served from this site. You can select more than one for mixed-script images, such as English and Urdu on an invoice. Each pack downloads once, a few megabytes, and is cached for offline use.

Does Urdu OCR work with Nastaliq fonts?

Printed Urdu in common newspaper and invoice fonts is recognised, though accuracy is lower than for English because of the connected script. Use a high-resolution, straightened image and the black and white filter from the Document Scanner for the best results. Handwritten Urdu is not supported.

How accurate is it?

Clean screenshots and 300 DPI scans typically reach 95 to 99% for English. Phone photos of paper land around 85 to 95%, and photos of curved product labels lower still. Resolution, straightness and contrast matter most; letters should be at least 20 px tall.

Can I extract text from a PDF?

Yes, indirectly. Convert the PDF into 300 DPI page images with Image to PDF / PDF to Images, then drop those images here. This also works for image-only PDFs with no selectable text, which is common with scanned invoices and supplier catalogues.

Is my image sent to a server for recognition?

No. Tesseract runs inside your browser in a Web Worker. The only downloads are the language packs from this website. Your images and the extracted text stay on your device, which is why it is safe to use on customer order screenshots and identity documents.