TapToConvert All conversions

Home / Tools / Amharic OCR

Amharic OCR

Extract Amharic text from photos, scans and screenshots with OCR trained on the Ge'ez (Ethiopic) script. Get Unicode Amharic you can copy and edit. Nothing is uploaded.

Drop files hereChoose images Nothing is uploaded. Any size your device can handle.

Add a file from a URL

The file is fetched by your browser straight from that website, not through our servers. It works when the site allows other websites to download its files.

Settings

Add files to begin.

    How Amharic OCR works

    Drop a photo of an Amharic book page, a newspaper, a document or a screenshot, and the Fidel characters are recognised and saved as Unicode Amharic for Word, Google Docs, Telegram or a translator. Because the output is Unicode, it does not depend on the legacy Ethiopic fonts, such as older Power Geez versions, that some documents were typed in.

    Amharic is written in the Ge'ez script, Fidel, a syllabary in which each of more than 30 base consonants has seven main forms, one for each vowel, made by adding small strokes, rings and bends to the base letter: ለ, ሉ, ሊ, ላ, ሌ, ል, ሎ. That gives over 200 common characters, many of which differ only by one of those small additions. Tesseract's Amharic model reads whole lines with a neural network trained on printed Amharic and outputs each syllable as one Unicode character.

    Errors come from the small vowel markers, which blur in low-resolution photos, and from letters that sound alike but are written differently, such as ሀ, ሐ and ኀ, where the model can rely only on the shape. Ethiopic punctuation, the word separator ፡ and the full stop ።, is usually read, although many modern texts use spaces instead of ፡. Ethiopic numerals such as ፩, ፪ and ፫ are less reliable than 0 to 9.

    The Amharic model, about 2.4 MB, downloads from the jsDelivr CDN the first time you press the button and is cached, along with the OCR engine of about 4 MB, so later images start straight away. It is trained on the Ge'ez script, so if the text mixes in English words, tick Also read English words and the English model reads alongside it, a little more slowly. The image itself is read inside this tab and never uploaded. The result is a UTF-8 text file named after the image. For an editable Word document instead, use Image to Word and choose Amharic.

    Amharic OCR options explained

    LanguageAmharic, already selected.
    ScriptGe'ez (Ethiopic Fidel), a syllabary written left to right.
    Model sizeAbout 2.4 MB, downloaded once and cached.
    OutputPlain UTF-8 text, one file per image.
    LayoutNot kept; the words come out in reading order.

    When to use it

    Use Amharic OCR to copy text from a scanned book, letter or certificate, to type up a photographed newspaper article or notice, to search old documents, or to paste Amharic from an image into a translator without an Ethiopic keyboard.

    Use a sharp, close image so the vowel strokes on each character are clearly visible; they are what tell ለ from ሉ and ሊ.

    About the OCR engine

    Text recognition uses Tesseract, the open-source OCR engine first developed at HP, then for many years at Google, and now maintained by its open-source community. Its current generation reads whole lines of text with a neural network (an LSTM) rather than matching letters one at a time, which is what lets it handle joined scripts such as Arabic and Devanagari. Tesseract.js compiles it to WebAssembly so it runs in this tab. Each language has its own trained model, downloaded from the jsDelivr CDN only when that language is chosen and then cached, so reading English never downloads Hindi, and the other way round. The models used here are the integer versions of Tesseract's most accurate models, which keep nearly all of their accuracy at a fraction of the size. Tesseract is built for printed text: it reads books, letters, forms, signs and screenshots well, handwriting poorly, and it keeps the words of a page in reading order but not its layout.

    Amharic OCR troubleshooting

    Characters of the same consonant are confused

    The vowel strokes are too small. Use a sharper, larger image.

    A legacy-font document copies as nonsense

    Older Ethiopic fonts do not copy as Unicode. Convert the page to an image and read it here.

    The result is in English letters

    Choose Amharic in the language menu and convert again.

    How to use Amharic OCR, step by step

    1. Press "Choose images" or drag files onto the box.
    2. Set the options if you need to; the defaults suit most uses.
    3. Press "Extract text". The work happens on your device.
    4. Save the result with its download button.

    Is it safe to do this online?

    With most online tools, "online" means your file is uploaded to a company's server, processed there and kept for a while before it is deleted. Here it is not. The page downloads the tool's code to your browser, and your file is read and processed inside the tab on your own device. It is never sent to TapToConvert or anyone else.

    You can check this yourself: once the page and its engine have loaded, turn off Wi-Fi and the tool still works. That also means there is no queue, no daily limit and no file size cap set by a server; the only limit is the memory your browser gives a single tab.

    Amharic OCR FAQ

    Does it work for Tigrinya?

    Tigrinya has its own model. Choose Tigrinya in the language list.

    Is the output Unicode?

    Yes, standard Unicode Ethiopic.

    Can it read handwritten Amharic?

    Not reliably; it is trained on printed text.

    Can I get a Word document?

    Yes. Use Image to Word and choose Amharic.

    Is it free?

    Yes. No sign-up, no watermark and no limit on use. TapToConvert is supported by advertising.

    Are my files uploaded?

    No. Everything happens inside your browser tab, on your own device.

    Try MP4 to MP3, HEIC, merge PDF or compress video.