Tesseract OCR (running on your own computer)
Turns a scanned page or a photo into searchable text, entirely on your own computer.
Safe for sensitive material
Tesseract reads the text out of a scanned page or a photograph and hands back text you can search, copy and edit.
What you get depends more on the scan than on the software. A straight, high-contrast page comes out well. A photo taken at an angle in poor light, or handwriting, comes out as something a person has to sit down and correct. Budget that correction time before promising anyone a searchable archive.
It has no window and no buttons: it is a command you run, on one page or one folder at a time. Someone in the team has to be willing to work that way.
What happens to your data
Where this data goes
- your own computer
Trains on what you type?
No. Your input is not used to train models.
How long it is kept
Nothing is sent anywhere. Tesseract is a program you run yourself on a scanned document or photo; the text it produces stays on your machine.
Account required
No account is needed.
Getting it, and using it in Iraq
Access in Iraq
Downloaded and installed with no account and no geographic restriction. Works with no internet connection once installed.
Language quality
No one has published a measurement for this.
Arabic (Modern Standard)
Tesseract has published Arabic-language trained data since version 3 (https://github.com/tesseract-ocr/tessdata), and community reports describe known problems specifically with Arabic and other right-to-left scripts, but no official per-language accuracy figure is published and no independent benchmark was run for this entry. Untested: expect to need to review and correct its output on scanned Arabic text, especially handwriting or low-quality scans.
Iraqi Arabic
OCR reads printed or handwritten script rather than spoken dialect, so Iraqi-dialect text written in standard Arabic script is covered by the same trained data as MSA, at the same unverified accuracy. Untested.
Kurdish (Sorani)
Community-contributed Kurdish trained data exists for Tesseract but is not part of the officially maintained language packs checked for this entry, and no benchmark was run. Untested.
Sources
- tesseract-ocr/tesseractGitHub
- tesseract-ocr/tessdataGitHub
Last checked: 2026-08-31