TapToConvert All conversions

Home / Tools / Marathi OCR

Marathi OCR

Extract Marathi text from images with OCR trained on Marathi Devanagari, including ळ and the eyelash ra. Photos, scans and screenshots become Unicode Marathi. The image is not uploaded.

Drop files hereChoose images Nothing is uploaded. Any size your device can handle.

Add a file from a URL

The file is fetched by your browser straight from that website, not through our servers. It works when the site allows other websites to download its files.

Settings

Add files to begin.

    How Marathi OCR works

    Drop a photo of a Marathi newspaper, a 7/12 land record extract, a school textbook page or a screenshot, and the text is recognised and saved as Unicode Marathi you can paste into Word, Google Docs or a translator. It is real Unicode text, not a legacy font encoding such as Shree Lipi or Kruti Dev, so it works in every app.

    Marathi is written in Devanagari, like Hindi, but it has its own letters and habits: ळ is common, the eyelash ra (ऱ्) appears in words such as दुसऱ्या, and spelling follows Marathi rules. That is why it has a separate model, trained on printed Marathi, which expects Marathi words and letters instead of guessing Hindi ones. Like the Hindi model, it reads whole lines with a neural network, which handles the headline, the matras and conjuncts such as क्ष and ज्ञ.

    Errors cluster around the smallest marks: anusvara dots, the short and long i and u matras, and the half-letters of conjuncts in small print. Numbers may be printed in Devanagari digits (० to ९) or as 0 to 9, and both are read as printed. Choosing Hindi for a Marathi page works in part but loses accuracy on Marathi words, so always pick Marathi.

    The Marathi model, about 2.0 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. It is trained on Marathi in Devanagari, so if the text mixes in English words, tick Also read English words and the English model reads alongside it, a little more slowly. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image. For an editable Word document instead, use Image to Word and choose Marathi.

    Marathi OCR options explained

    LanguageMarathi, already selected.
    ScriptDevanagari, left to right, with ळ and the eyelash ra (ऱ्) that Marathi uses.
    Model sizeAbout 2.0 MB, downloaded once and cached.
    OutputPlain UTF-8 text, one file per image.
    LayoutNot kept; the words come out in reading order.

    When to use it

    Use Marathi OCR to copy text from a scanned land record or government circular, to type up photographed notes from a Marathi textbook, to search old issues of a newspaper or magazine, or to paste a notice into a translator.

    Scan official records at 300 dpi or photograph them in daylight, and crop out stamps and signatures that overlap the text where you can, because they confuse the recognition.

    About the OCR engine

    Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.

    Marathi OCR troubleshooting

    ळ comes out as ल or another letter

    Check that Marathi, not Hindi, is selected, and use a sharper image.

    Legacy-font PDFs copy as nonsense

    Text typed in Shree Lipi or similar fonts does not copy as Unicode. Convert the page to an image and read it here.

    Numbers in a table are wrong

    Stamps, ruled lines and table borders near numbers confuse the reader. Crop tightly around the text.

    How to use Marathi OCR, step by step

    1. Press "Choose images" or drag files onto the box.
    2. Set the options if you need to; the defaults suit most uses.
    3. Press "Extract text". The work happens on your device.
    4. Save the result with its download button.

    Is it safe to do this online?

    With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.

    You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.

    Marathi OCR FAQ

    Is Marathi OCR different from Hindi OCR?

    Yes. Both use Devanagari, but the Marathi model is trained on Marathi text, so it reads Marathi letters and words more accurately.

    Does it give Unicode text?

    Yes, standard Unicode Marathi that works in every app.

    Can it read handwriting?

    Not reliably; it reads printed text.

    Can I read several pages at once?

    Yes. Add all the images; each gets its own text file.

    Is it free?

    Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.

    Are my files uploaded?

    No. Everything happens inside your browser tab, on your own device.

    Try MP4 to MP3, HEIC, merge PDF or compress video.