How Punjabi OCR works
Drop a photo of a Punjabi newspaper, a scanned document, a gutka or book page, or a screenshot, and the Gurmukhi text is recognised and saved as Unicode Punjabi for Word, Google Docs, WhatsApp or a translator. The output is Unicode, not the AnmolLipi or Asees legacy fonts many older Punjabi documents were typed in, so it displays correctly everywhere.
Gurmukhi hangs its letters from a headline, like Devanagari, with vowel signs above, below and beside the consonants, and a set of marks that matter for meaning: tippi (ੰ) and bindi (ਂ) for nasal sounds, and addak (ੱ), which doubles the following consonant. Subscript forms of ਰ, ਹ and ਵ appear under letters in words such as ਪ੍ਰੇਮ. Tesseract's Punjabi model reads whole lines with a neural network trained on printed Gurmukhi.
The small marks are the usual errors: tippi and bindi confused or dropped, a missing addak, and subscript letters lost in small print. Gurmukhi digits (੦ to ੯) are read when printed. Punjabi written in Shahmukhi, the Urdu-based script used in Pakistan, is a different alphabet: choose Urdu for it.
The Punjabi model, about 1.4 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. It is trained on Gurmukhi, so if the text mixes in English words, tick Also read English words and the English model reads alongside it, a little more slowly. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image. For an editable Word document instead, use Image to Word and choose Punjabi.
Punjabi OCR options explained
| Language | Punjabi, already selected. |
|---|---|
| Script | Gurmukhi, written left to right, with a headline; Shahmukhi Punjabi needs the Urdu model. |
| Model size | About 1.4 MB, downloaded once and cached. |
| Output | Plain UTF-8 text, one file per image. |
| Layout | Not kept; the words come out in reading order. |
When to use it
Use Punjabi OCR to copy text from a scanned book, newspaper or gutka, to type up a photographed notice, form or letter, to search old documents, or to paste Punjabi from an image into a translator.
Photograph the page close and straight, and check by zooming in that tippi, bindi and addak are clear marks, not specks, before converting.
About the OCR engine
Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.
Punjabi OCR troubleshooting
Tippi, bindi or addak are missing
Use a sharper, larger image; these marks are small.
An AnmolLipi document copies as English letters
Legacy fonts store Gurmukhi as Latin letters. Convert the page to an image and read it here to get Unicode.
Shahmukhi text is unreadable
Shahmukhi uses the Urdu alphabet. Choose Urdu and convert again.
How to use Punjabi OCR, step by step
- Press "Choose images" or drag files onto the box.
- Set the options if you need to; the defaults suit most uses.
- Press "Extract text". The work happens on your device.
- Save the result with its download button.
Is it safe to do this online?
With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.
You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.
Punjabi OCR FAQ
Is the output Unicode Gurmukhi?
Yes, standard Unicode that works in every app.
Does it read Shahmukhi?
Choose Urdu for Shahmukhi; the Punjabi model reads Gurmukhi.
Can it read handwritten Punjabi?
Not reliably; it is trained on printed text.
Can I get a Word document?
Yes. Use Image to Word and choose Punjabi.
Is it free?
Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.
Are my files uploaded?
No. Everything happens inside your browser tab, on your own device.