TapToConvert All conversions

Home / Tools / Korean OCR

Korean OCR

Extract Korean text from photos, scans and screenshots with OCR that reads Hangul syllables and keeps word spacing. Get Unicode Korean you can copy and translate. The image stays on your device.

Drop files hereChoose images Nothing is uploaded. Any size your device can handle.

Add a file from a URL

The file is fetched by your browser straight from that website, not through our servers. It works when the site allows other websites to download its files.

Settings

Add files to begin.

    How Korean OCR works

    Drop a photo of a Korean document, a menu, a sign, a book page or a screenshot from KakaoTalk, and the text is recognised and saved as Unicode Korean for Word, Hangul (HWP), Google Docs or a translator.

    Hangul packs letters into square syllable blocks: 가 is ㄱ and ㅏ, and 한 is ㅎ, ㅏ and ㄴ stacked together. There are more than 11,000 possible blocks, of which a few thousand are common, and Tesseract's Korean model recognises each block as one character, which is how Unicode stores it. Unlike Chinese and Japanese, Korean puts spaces between words, and the model keeps them, although the spacing in the output can differ slightly from the page.

    Blocks that differ by one small stroke, such as 이 and 어 or 은 and 온, are the usual errors in small or blurred print. Older documents mix in Chinese characters (hanja), which the Korean model does not read reliably. For the vertical text of some old books and signs, choose Korean (vertical).

    The Korean model, about 1.6 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. It is trained on Hangul, so if the text mixes in English words, tick Also read English words and the English model reads alongside it, a little more slowly. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image. For an editable Word document instead, use Image to Word and choose Korean.

    Korean OCR options explained

    LanguageKorean, already selected.
    ScriptHangul, in syllable blocks, left to right with spaces between words; a vertical model is available.
    Model sizeAbout 1.6 MB, downloaded once and cached.
    OutputPlain UTF-8 text, one file per image.
    LayoutNot kept; the words come out in reading order.

    When to use it

    Use Korean OCR to copy text from a scanned document or contract, to translate a menu, label or sign by pasting its text into a translator, to get lines out of a screenshot of a chat or webtoon, or to search photographed notes.

    Webtoon and game screenshots read best when you crop to one speech bubble or text box at a time, away from the artwork.

    About the OCR engine

    Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.

    Korean OCR troubleshooting

    Similar syllables are mixed up

    Use a sharper, larger image; the strokes that separate them are small.

    Hanja come out as wrong characters

    The Korean model reads Hangul. For text that is mostly Chinese characters, choose Chinese (traditional).

    Vertical text is garbled

    Choose Korean (vertical) for text that runs from top to bottom.

    How to use Korean OCR, step by step

    1. Press "Choose images" or drag files onto the box.
    2. Set the options if you need to; the defaults suit most uses.
    3. Press "Extract text". The work happens on your device.
    4. Save the result with its download button.

    Is it safe to do this online?

    With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.

    You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.

    Korean OCR FAQ

    Does it keep the spaces between words?

    Yes, mostly; the spacing can differ slightly from the original.

    Can it read handwritten Korean?

    Not reliably; it is trained on printed text.

    Can I translate the result?

    Yes. Copy the text into any translator.

    Can I get a Word document?

    Yes. Use Image to Word and choose Korean.

    Is it free?

    Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.

    Are my files uploaded?

    No. Everything happens inside your browser tab, on your own device.

    Try MP4 to MP3, HEIC, merge PDF or compress video.