How Chinese OCR works
Drop a photo of a Chinese document, a book page, a menu, a sign or a screenshot, and the characters are recognised and saved as Unicode text for Word, WeChat, Google Docs or a translator. Simplified Chinese is selected; for text from Taiwan, Hong Kong or Macau, or older books, choose Chinese (traditional), and for traditional text printed in columns from top to bottom, Chinese (traditional, vertical).
A Chinese OCR model has to tell apart thousands of characters, not dozens of letters, and many differ by a single stroke, such as 己, 已 and 巳, or 土 and 士. Simplified and traditional characters have separate models because they are, in effect, separate character sets: 书 and 書, or 语 and 語. Tesseract's models recognise each character as a whole, with a neural network reading the line, and output standard Unicode, including full-width punctuation such as , and 。.
Chinese does not put spaces between words, but Tesseract tends to put spaces between characters; they are removed here, while spaces around English words and numbers are kept. Small print, where strokes blur together, and decorative or brush fonts are the main sources of errors. English words and numbers inside the text are often read correctly too, but check them.
The Chinese model, about 1.7 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image. For an editable Word document instead, use Image to Word and choose Chinese (simplified).
Chinese OCR options explained
| Language | Chinese (simplified), already selected; traditional and vertical traditional are in the same menu. |
|---|---|
| Script | Chinese characters (hanzi), simplified or traditional, horizontal or vertical. |
| Model size | About 1.7 MB, downloaded once and cached. |
| Output | Plain UTF-8 text, one file per image. |
| Layout | Not kept; the words come out in reading order. |
When to use it
Use Chinese OCR to copy text from a scanned contract, certificate or invoice, to translate a menu, label or sign by pasting its text into a translator, to get lines out of a screenshot of a chat or web page, or to look up characters you cannot type.
Characters need more pixels than letters: make sure each character is at least 20 to 30 pixels tall in the image, and crop away other text so only the lines you want are read.
About the OCR engine
Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.
Chinese OCR troubleshooting
Characters are swapped for similar ones
Use a sharper, larger image; the strokes that separate similar characters are thin.
Traditional characters come out wrong
Choose Chinese (traditional); the simplified model does not know most traditional forms.
Vertical text is jumbled
Choose Chinese (traditional, vertical) for columns that run from top to bottom.
How to use Chinese OCR, step by step
- Press "Choose images" or drag files onto the box.
- Set the options if you need to; the defaults suit most uses.
- Press "Extract text". The work happens on your device.
- Save the result with its download button.
Is it safe to do this online?
With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.
You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.
Chinese OCR FAQ
Does it read traditional Chinese?
Yes. Choose Chinese (traditional), or its vertical version for columns.
Can it read written Cantonese?
Common characters are read with Chinese (traditional); characters used only in Cantonese may not be.
Can it read handwriting or calligraphy?
No. Only printed characters read reliably.
Can I get a Word document?
Yes. Use Image to Word and choose the Chinese version you need.
Is it free?
Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.
Are my files uploaded?
No. Everything happens inside your browser tab, on your own device.