How Tamil OCR works
Drop a photo of a Tamil book, a newspaper cutting, a scanned certificate or a screenshot, and the text is recognised and saved as Unicode Tamil for Word, Google Docs, WhatsApp or a translator. The output is standard Unicode, not TAM, TAB or another legacy font encoding, so it displays correctly anywhere.
Tamil has 12 vowels and 18 consonants, which combine into 216 vowel-consonant letters (uyirmei), plus the aytham ஃ. Most combinations are written by adding a vowel sign to the left, the right or both sides of the consonant, as in கொ, and a dot above, the pulli, marks a consonant with no vowel, as in க். Tesseract's Tamil model reads whole lines with a neural network trained on printed Tamil, so it recognises these combined shapes as a whole rather than as loose strokes.
The pulli is the mark most often lost, turning க் into க, especially in small or blurred print. Books printed before the 1978 script reform use older shapes for letters such as ணை, லை and றா; the model knows modern printing best, so old books need more checking. The Grantha letters used for Sanskrit and English sounds, ஜ, ஷ, ஸ and ஹ, are recognised like any other letter. Tesseract's Tamil model also follows every pulli with an invisible zero-width character that stops words matching in a search; it is removed here, so the text is stored the way you would type it.
The Tamil model, about 1.4 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. It is trained on Tamil script, so if the text mixes in English words, tick Also read English words and the English model reads alongside it, a little more slowly. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image. For an editable Word document instead, use Image to Word and choose Tamil.
Tamil OCR options explained
| Language | Tamil, already selected. |
|---|---|
| Script | Tamil, written left to right; vowel signs can sit on both sides of a consonant. |
| Model size | About 1.4 MB, downloaded once and cached. |
| Output | Plain UTF-8 text, one file per image. |
| Layout | Not kept; the words come out in reading order. |
When to use it
Use Tamil OCR to copy text from a scanned book or magazine, to type up a photographed land record, notice or form, to search old family documents, or to paste Tamil text from an image into a translator or an email.
Make sure the pulli dots are visible when you zoom in on the image. If they are specks of a pixel or two, take a closer or sharper photo before converting.
About the OCR engine
Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.
Tamil OCR troubleshooting
Pulli dots are missing
The image is too small or blurred. Use a sharper photo at a larger size.
Old book text has many errors
Pre-reform letter shapes are read less reliably. Correct the result by hand, or search it for the most common mistakes.
The output is English letters
Choose Tamil in the language menu, then convert again.
How to use Tamil OCR, step by step
- Press "Choose images" or drag files onto the box.
- Set the options if you need to; the defaults suit most uses.
- Press "Extract text". The work happens on your device.
- Save the result with its download button.
Is it safe to do this online?
With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.
You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.
Tamil OCR FAQ
Is the text Unicode Tamil?
Yes, standard Unicode that works in every app and font.
Can it read handwritten Tamil?
Not reliably; it is trained on printed text.
Does it read Tamil numerals?
Printed documents almost always use 0 to 9, which it reads. The traditional Tamil numerals are rare and may not be recognised.
Can I get a Word file?
Yes. Use Image to Word and choose Tamil.
Is it free?
Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.
Are my files uploaded?
No. Everything happens inside your browser tab, on your own device.