Blog Upload Shorten Tools Gallery FAQ Contact Sign in
📢 永久無廣告,自動清除EXIF,保護隱私內容
Tools

📝 Image to Text (OCR)

Pull the text out of a screenshot

Picking the right language matters a lot. Chinese training data is around 12MB and English around 4MB; the first run downloads it and the browser caches it afterwards.
Greyscale is usually enough. Thresholding helps a lot with dark text on a light background but destroys anything photographic.
Text being too small is the most common cause of failure. Below roughly 20 pixels of character height, upscaling first usually helps noticeably.

ℹ️ The recognition engine (Tesseract) and its language data load from a CDN, so the first run takes a moment. Your image is never uploaded — the engine runs in your browser, and only the program and model files travel over the network.

🔒 Everything runs inside your browser — your files are never uploaded to our servers and we never see them. Close the tab and nothing is left behind.

Image to Text (OCR): before and after example
Original on the left, the actual output of this tool on the right — not a mock-up. Any file sizes or dimensions shown are real measured values.

Step-by-step

  1. Pick an image with text Screenshots, scans, photographed documents, signage — all work. The sharper, straighter, and larger the text, the better the result, and that matters more than any setting here.
  2. Choose the right language This is the setting with the biggest effect on accuracy. Choosing the wrong language drops recognition to nearly unusable, because the engine matches against that language's character models.
  3. Decide on pre-processing Greyscale suits most cases. Black and white thresholding is markedly better for dark text on a light background (scans, printed pages) but destroys anything photographic. When unsure, start with greyscale.
  4. Press start The first run downloads the engine and language data (around 12MB for Chinese), so expect a wait. The browser caches it afterwards and subsequent runs on the same device are much faster. Progress appears in the status line.
  5. Copy it or download the .txt The result appears below with an average confidence figure. Below about 70% confidence, proofread carefully. Copy it directly or download it as a plain text file.

When you would use this

Getting text out of a screenshot

Somebody sends a screenshot and you need to quote a line from it, but retyping is the only option. This is OCR's most everyday use — error messages, conversations, text on a page you cannot select.

Digitising paper

Photograph receipts, business cards, handouts, or book pages and turn them into text you can search, edit, and paste elsewhere. For scans, the black-and-white pre-processing noticeably improves accuracy.

Signs and menus in a language you cannot type

Recognise the text into something copyable first, then paste it into a translator. Far quicker than typing character by character from a photo, especially for Japanese or Korean.

Extracting text from scanned PDF pages

A scanned PDF contains images with no text layer, so nothing is searchable. Convert the page to PNG at 300 DPI with PDF to Images first, then recognise it here.

Four things that determine accuracy

How well OCR performs is mostly about what you feed it, not how you configure it. In order of impact:

  1. The right language — the engine matches against that language's character models. Running Chinese through the English model returns gibberish.
  2. Large enough text — at least 20 pixels of character height, 30 to be comfortable. This is the most common cause of failure, and phone screenshots easily fall below it after scaling.
  3. A straight image — a few degrees of skew measurably hurts accuracy. Shoot documents square on, or straighten them afterwards.
  4. Strong contrast — dark text on a light background is ideal. Coloured backgrounds, gradients, and shadows all interfere.

Greyscale or thresholding

Greyscale simply removes colour while keeping every level of light and shade. Safe and general — use it when unsure.

Black-and-white thresholding forces every pixel to pure black or white, using the image's mean brightness as the cut-off. For printed pages this is a powerful clean-up: background noise, paper texture, and light shadows all disappear, leaving only the text. But if the image contains photography, a gradient background, or text that is itself pale, thresholding destroys the information outright.

How this differs from online OCR services

Most online OCR uploads your image to their servers, recognises it there, and sends the text back. That means your receipts, business cards, and documents have passed through somebody else's machine.

This page uses the WebAssembly build of Tesseract, so the entire recognition happens in your browser. Only the engine and language models travel over the network (downloaded once from a CDN), and they are unrelated to your image content. The trade-offs are a first-run download and accuracy below the large cloud models commercial services use — an explicit exchange that is usually worth it when the documents are sensitive.


Frequently asked questions

Why does the first run take so long to download?
Because OCR needs a recognition engine plus the character models for your language. The Tesseract core is around 2MB, English data about 4MB, and Chinese about 12MB because of the sheer number of characters. It all runs in your browser, so it all has to come down. The browser caches it, so the second run on the same device is much faster.
The results are full of errors. What now?
Check in this order. One: is the language right? (biggest factor). Two: is the text large enough? Character heights under 20 pixels are almost never accurate — keep "upscale small images" on, or run it through Upscale Image first. Three: is the image straight? Accuracy falls off quickly past a few degrees of skew; straighten it with Rotate Image. Four: try black-and-white thresholding, which helps scans a lot.
Can it read handwriting?
Essentially no. Tesseract is trained on printed type, and handwriting recognition rates are very low — neat block capitals might get partway. Handwriting needs a different class of model entirely, outside what this engine does.
Is my image uploaded?
No. The engine is WebAssembly running inside your browser, and the image never leaves your device. Only the engine code and language models travel over the network (downloaded from a CDN), and they have nothing to do with your image content. That is a fundamental difference from online OCR services that send your image to a server.

What you might need next

Image work rarely ends in one step. These pair up with Image to Text (OCR) most often:

Want to share the result? Upload turns images, photos, or videos into a short link with optional password, expiry, and view limits. Or head back to all 12 tools.


More tools