TapToConvert All conversions

Home / Tools / Persian OCR

Persian OCR

Extract Persian (Farsi) text from photos, scans and screenshots with OCR that writes the Persian forms of ی and ک. Get Unicode Farsi text you can edit. The image is never uploaded.

Drop files hereChoose images Nothing is uploaded. Any size your device can handle.

Add a file from a URL

The file is fetched by your browser straight from that website, not through our servers. It works when the site allows other websites to download its files.

Settings

Add files to begin.

    How Persian OCR works

    Drop a photo of a Persian book page, a scanned document, a screenshot or a sign, and the text is recognised and saved as Unicode Persian for Word, Google Docs, Telegram or a translator. Lines are stored right to left in logical order, the way every modern app handles Farsi.

    Persian uses the Arabic script with four extra letters, پ, چ, ژ and گ, and writes ی and ک in its own forms, which are different characters in Unicode from the Arabic ي and ك even where they look alike. The Persian model is trained to output the Persian forms, so searches and spell-checkers work on the result. It reads whole lines with a neural network, and at about 0.4 MB it is the smallest language download on this site.

    The model writes the half-space (a zero-width non-joiner) that keeps the parts of a word such as زبان‌ها apart, but where the print leaves no visible gap it can come out as a full space or with the parts joined; search for those words if exact typing matters. Blurred dots turn letters into their neighbours, such as ب into پ. Numbers printed in Persian digits (۰ to ۹) are read as Persian digits. Nastaliq calligraphy, used in poetry books and titles, reads poorly; ordinary Naskh-style type reads well.

    The Persian model, about 0.4 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. It is trained on the Persian alphabet, so if the text mixes in English words, tick Also read English words and the English model reads alongside it, a little more slowly. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image, stored in logical order so it displays right to left. For an editable Word document instead, use Image to Word and choose Persian (Farsi).

    Persian OCR options explained

    LanguagePersian (Farsi), already selected.
    ScriptPersian (Arabic script with پ چ ژ گ), written right to left.
    Model sizeAbout 0.4 MB, downloaded once and cached.
    OutputPlain UTF-8 text, one file per image.
    LayoutNot kept; the words come out in reading order.

    When to use it

    Use Persian OCR to copy text from a scanned book or article, to get the wording of a certificate, contract or letter into a document, to search photographed pages, or to paste Farsi from an image into a translator when you do not have a Persian keyboard.

    Photograph pages flat and close, so the dots and small letters such as ژ stay distinct, and crop away pictures and margins.

    About the OCR engine

    Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.

    Persian OCR troubleshooting

    Arabic ي and ك appear instead of ی and ک

    The language was set to Arabic. Choose Persian (Farsi) and convert again.

    Words are split or joined oddly

    That is usually the half-space. Correct those words, or search for the prefix می and the suffix ها.

    Poetry titles are unreadable

    Nastaliq and decorative calligraphy are beyond this engine; body text in ordinary type reads well.

    How to use Persian OCR, step by step

    1. Press "Choose images" or drag files onto the box.
    2. Set the options if you need to; the defaults suit most uses.
    3. Press "Extract text". The work happens on your device.
    4. Save the result with its download button.

    Is it safe to do this online?

    With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.

    You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.

    Persian OCR FAQ

    Is Farsi the same as Persian here?

    Yes. Farsi is the Persian name for the language; choose Persian (Farsi).

    Does it work for Dari or Tajik?

    Dari uses the same script and usually reads well with Persian selected. Tajik is written in Cyrillic, so this model does not read it.

    Is the result right to left?

    Yes, saved in logical order as Unicode, which apps display right to left.

    Can I get a Word document?

    Yes. Use Image to Word with Persian selected.

    Is it free?

    Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.

    Are my files uploaded?

    No. Everything happens inside your browser tab, on your own device.

    Try MP4 to MP3, HEIC, merge PDF or compress video.