How Sinhala OCR works
Drop a photo of a Sinhala newspaper, a scanned document, a school book page or a screenshot, and the text is recognised and saved as Unicode Sinhala for Word, Google Docs, Facebook or a translator. Unlike text typed in the legacy FM fonts common in older Sri Lankan documents, the result is standard Unicode that displays correctly everywhere.
Sinhala letters are rounded, built from loops and curls, and many differ only by a small stroke. Vowel signs (pili) attach before, after, above or below a consonant, and some, such as ො, wrap around it on both sides. The al-lakuna (්) removes the inherent vowel, and conjuncts and touching letters (bandi akuru) appear in formal text. Tesseract's Sinhala model reads whole lines with a neural network rather than cutting out single letters, which suits a script this connected.
Sinhala is one of the harder scripts for OCR, and results are best on clean, modern print. Expect errors where letters that look alike, such as බ and ඛ or ද and ඳ, are printed small, and where the fine strokes of the pili are blurred. Old books with worn type or unusual fonts need careful checking.
The Sinhala model, about 1.1 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. It is trained on Sinhala script, so if the text mixes in English words, tick Also read English words and the English model reads alongside it, a little more slowly. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image. For an editable Word document instead, use Image to Word and choose Sinhala.
Sinhala OCR options explained
| Language | Sinhala, already selected. |
|---|---|
| Script | Sinhala, written left to right, with rounded letters and vowel signs on every side. |
| Model size | About 1.1 MB, downloaded once and cached. |
| Output | Plain UTF-8 text, one file per image. |
| Layout | Not kept; the words come out in reading order. |
When to use it
Use Sinhala OCR to copy text from a scanned certificate, deed or letter, to type up a photographed newspaper article, to search old documents, or to paste Sinhala text into a translator or an email without typing it.
Photograph the page closely enough that the small strokes of the vowel signs are sharp. Clean black text on white paper gives much better results than colour, glossy or patterned backgrounds.
About the OCR engine
Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.
Sinhala OCR troubleshooting
Similar letters are mixed up
The print is too small or blurred for the details that separate them. Use a larger, sharper image.
Text from an FM-font PDF copies as English letters
Legacy FM fonts store Sinhala as Latin letters. Convert the page to an image and read it here to get Unicode.
Nothing is recognised
Check that Sinhala is selected and that the text is dark on a light background.
How to use Sinhala OCR, step by step
- Press "Choose images" or drag files onto the box.
- Set the options if you need to; the defaults suit most uses.
- Press "Extract text". The work happens on your device.
- Save the result with its download button.
Is it safe to do this online?
With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.
You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.
Sinhala OCR FAQ
Is the output Unicode Sinhala?
Yes, standard Unicode that works in every app.
How accurate is it?
Good on clean, modern print; weaker on old books, small type and photos taken at an angle. Always proofread.
Can it read handwritten Sinhala?
No, it is trained on printed text.
Can I get a Word document?
Yes. Use Image to Word and choose Sinhala.
Is it free?
Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.
Are my files uploaded?
No. Everything happens inside your browser tab, on your own device.