How Arabic OCR works
Drop a photo of a printed Arabic page, a scanned document, a book cover or a screenshot, and the text is recognised and saved as Unicode Arabic that you can paste into Word, a translator or a search box. Lines are written out right to left in logical order, the way Arabic is stored in every modern app, so the text displays and edits correctly.
Arabic is written in a joined script: most letters take a different shape at the start, middle and end of a word or when standing alone, and many letters differ only by their dots, as with ب ت ث and ج ح خ. Tesseract's Arabic model reads whole lines with a neural network trained on printed Arabic, so it follows the joins and ligatures such as لا instead of trying to cut words into separate letters, which is where older Arabic OCR failed.
Accuracy depends on the dots surviving. In a blurry photo or a small, compressed screenshot, the dots under ب or over ق can merge or vanish, and one letter turns into another. Short vowel marks (harakat) such as fatha and kasra, used in the Quran, children's books and poetry, are often dropped or misplaced, so fully vowelled text needs checking. Calligraphic styles such as Kufi, Diwani and Thuluth are beyond it: it reads the everyday Naskh-style type of books, newspapers and screens. English names, email addresses and product names inside Arabic text are garbled by the Arabic model alone; tick Also read English words for those.
The Arabic model, about 1.7 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image, stored in logical order so it displays right to left. For an editable Word document instead, use Image to Word and choose Arabic.
Arabic OCR options explained
| Language | Arabic, already selected. |
|---|---|
| Script | Arabic, right to left, in joined letters; digits may be Western (0 to 9) or Eastern Arabic (٠ to ٩). |
| Model size | About 1.7 MB, downloaded once and cached. |
| Output | Plain UTF-8 text, one file per image. |
| Layout | Not kept; the words come out in reading order. |
When to use it
Use Arabic OCR to copy a paragraph from a scanned book, to get the text of an official letter, contract or certificate into a document, to search photographed articles, or to paste text from an image into a translator instead of typing it on a keyboard you do not have.
Photograph the page flat and in even light so the dots stay sharp, and crop to the text. For Persian or Urdu, which use extra letters, choose that language instead of Arabic.
About the OCR engine
Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.
Arabic OCR troubleshooting
Letters come out wrong
Usually the dots were lost. Use a sharper, larger image; screenshots should be taken at full size, not scaled down.
Words appear in the wrong order when pasted
The text is stored right to left correctly, but an editor set to left to right can show it oddly. Set the paragraph direction to right to left.
Vowel marks are missing
Tesseract often skips harakat. The consonants are usually right; add the marks by hand where they matter.
How to use Arabic OCR, step by step
- Press "Choose images" or drag files onto the box.
- Set the options if you need to; the defaults suit most uses.
- Press "Extract text". The work happens on your device.
- Save the result with its download button.
Is it safe to do this online?
With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.
You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.
Arabic OCR FAQ
Is the text right to left?
Yes. It is saved in logical order as Unicode Arabic, which every modern app displays right to left.
Does it read the Quran or other vowelled text?
It reads the letters, but short vowel marks are often dropped or misplaced, so check the result.
Can it read handwriting or calligraphy?
No. Only printed, Naskh-style text reads reliably.
Can I get a Word document?
Yes. Use Image to Word with Arabic selected; its paragraphs are set right to left.
Is it free?
Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.
Are my files uploaded?
No. Everything happens inside your browser tab, on your own device.