TapToConvert All conversions

Home / Tools / Urdu OCR

Urdu OCR

Extract Urdu text from images with OCR for the Urdu alphabet, written right to left. Photos, scans and screenshots become Unicode Urdu; Naskh-style print reads best. Nothing is uploaded.

Drop files hereChoose images Nothing is uploaded. Any size your device can handle.

Add a file from a URL

The file is fetched by your browser straight from that website, not through our servers. It works when the site allows other websites to download its files.

Settings

Add files to begin.

    How Urdu OCR works

    Drop a photo of a printed Urdu page, a screenshot or a scanned document, and the text is recognised and saved as Unicode Urdu that you can paste into Word, Google Docs, a search box or a translator. The text is stored right to left in logical order, as every modern app expects. Many Urdu books and newspapers were set in InPage, whose files often cannot be copied as Unicode; reading a picture of the page here is a way around that.

    Urdu uses the Arabic script with extra letters for its own sounds, such as ٹ, ڈ, ڑ, ں and ے, and Tesseract's Urdu model is trained to output them. The difficulty is the style of writing. Most Urdu books and newspapers are set in Nastaliq, a flowing style in which each word slopes down from right to left and letters stack on top of one another. Naskh-style type, whose letters sit on a flat line and which is common on websites, signs and forms, is much easier for OCR.

    On Nastaliq pages, expect errors where letters overlap and in small print, and check the dots that separate letters such as ب, پ, ت and ٹ. Short vowel marks (aerab) are often dropped. Numbers may be printed in Urdu digits (۰ to ۹) or as 0 to 9, and are read as printed.

    The Urdu model, about 1.0 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. It is trained on the Urdu alphabet, so if the text mixes in English words, tick Also read English words and the English model reads alongside it, a little more slowly. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image, stored in logical order so it displays right to left. For an editable Word document instead, use Image to Word and choose Urdu.

    Urdu OCR options explained

    LanguageUrdu, already selected.
    ScriptUrdu (Arabic script), right to left, printed in Nastaliq or Naskh style.
    Model sizeAbout 1.0 MB, downloaded once and cached.
    OutputPlain UTF-8 text, one file per image.
    LayoutNot kept; the words come out in reading order.

    When to use it

    Use Urdu OCR to copy text from a screenshot or scanned letter, to get the text of a notice, form or certificate into a document, to search photographed pages of a book, or to paste Urdu from an image into a translator.

    For Nastaliq pages, use the largest, sharpest image you can: a close photo of half a page reads better than a distant photo of the whole page.

    About the OCR engine

    Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.

    Urdu OCR troubleshooting

    Many errors on a book or newspaper page

    It is probably set in Nastaliq, which is harder to read than Naskh. Use a closer, sharper photo and proofread the result.

    Letters with dots are confused

    The dots were lost to blur or compression. Use a sharper image at full size.

    The words run the wrong way in my editor

    Set the paragraph direction to right to left; the text itself is stored correctly.

    How to use Urdu OCR, step by step

    1. Press "Choose images" or drag files onto the box.
    2. Set the options if you need to; the defaults suit most uses.
    3. Press "Extract text". The work happens on your device.
    4. Save the result with its download button.

    Is it safe to do this online?

    With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.

    You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.

    Urdu OCR FAQ

    Does it read Nastaliq?

    Yes, but with more errors than Naskh, particularly in small print. Proofread the result.

    Is the text Unicode?

    Yes, standard Unicode Urdu, unlike text copied from InPage files.

    Does it work for Punjabi in Shahmukhi?

    Shahmukhi uses the Urdu alphabet, so choosing Urdu usually works. For Gurmukhi, use Punjabi OCR.

    Can I get a Word document?

    Yes. Use Image to Word with Urdu selected; paragraphs are set right to left.

    Is it free?

    Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.

    Are my files uploaded?

    No. Everything happens inside your browser tab, on your own device.

    Try MP4 to MP3, HEIC, merge PDF or compress video.