How Sanskrit OCR works
Drop a photo of a printed Sanskrit text, a stotra, a page from the Gita, a textbook or a scanned printed edition, and the Devanagari is recognised and saved as Unicode Sanskrit that you can copy, search and edit, or paste into a dictionary or transliteration tool.
Sanskrit needs its own model even though it shares Devanagari with Hindi and Marathi. It uses conjuncts that modern languages rarely do, sometimes of three or four consonants, as in the ङ्क्ष of आकाङ्क्षा; marks such as the visarga (ः), the avagraha (ऽ) and the virama at the end of words; and sandhi, which joins words into very long unbroken strings. The Sanskrit model is trained on printed Sanskrit, and at about 5.7 MB it is the largest language download on this site.
Vedic accent marks (svaras) above and below the letters are not read, and text printed with them loses the marks and can gain errors around them. Rare conjuncts drawn as stacked ligatures in older type are the other main source of errors. Sanskrit in Roman transliteration (IAST, with ā, ṛ and ś) is not Devanagari, so this model does not apply to it.
The Sanskrit model, about 5.7 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. It is trained on Sanskrit in Devanagari, so if the text mixes in English words, tick Also read English words and the English model reads alongside it, a little more slowly. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image. For an editable Word document instead, use Image to Word and choose Sanskrit.
Sanskrit OCR options explained
| Language | Sanskrit, already selected. |
|---|---|
| Script | Devanagari, with conjuncts, visarga (ः) and avagraha (ऽ). |
| Model size | About 5.7 MB, downloaded once and cached. |
| Output | Plain UTF-8 text, one file per image. |
| Layout | Not kept; the words come out in reading order. |
When to use it
Use Sanskrit OCR to copy verses from a printed book into your notes, to make a searchable text of a stotra or commentary, to check a quotation against a scanned edition, or to paste a passage into a dictionary or transliteration tool.
When a page mixes Sanskrit with a Hindi or English translation, crop the parts apart and read each with its own language.
About the OCR engine
Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.
Sanskrit OCR troubleshooting
Complex conjuncts come out wrong
Use a larger, sharper image; stacked conjuncts need detail. Correct rare ones by hand.
Vedic accents are missing or cause errors
Accent marks are not read. Remove stray characters from the result, or add the accents by hand.
The first run takes a while
The Sanskrit model is about 5.7 MB. It downloads once and is cached.
How to use Sanskrit OCR, step by step
- Press "Choose images" or drag files onto the box.
- Set the options if you need to; the defaults suit most uses.
- Press "Extract text". The work happens on your device.
- Save the result with its download button.
Is it safe to do this online?
With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.
You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.
Sanskrit OCR FAQ
Is Sanskrit OCR different from Hindi OCR?
Yes. The Sanskrit model is trained on Sanskrit, with its conjuncts, visarga and avagraha, so it reads Sanskrit texts more accurately than the Hindi model.
Does it read Vedic accents?
No. Accent marks are dropped.
Can it output IAST transliteration?
No, it outputs Devanagari. Paste the result into a transliteration tool to convert it.
Can I get a Word document?
Yes. Use Image to Word and choose Sanskrit.
Is it free?
Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.
Are my files uploaded?
No. Everything happens inside your browser tab, on your own device.