ئەم وەرگێڕانە هێشتا لەلایەن مرۆڤەوە پشکنینی بۆ نەکراوە
بە یارمەتی وەرگێڕانی ئۆتۆماتیکی دروستکراوە. بۆ هەر شتێک کە کاری لەسەر دەکەیت، وەشانە ئینگلیزییەکە بخوێنەوە.
ئەو بڕگەیە بدۆزەرەوە کە پێویستتە لە کۆمەڵێک بەڵگەنامەی یاسایی عەرەبیی سکانکراودا، بەبێ متمانەکردن بە وەڵامی ئامێرەکە
کۆمەڵێک ڕۆژنامەی فەرمی و بڕیاری سکانکراو بە سفر دۆلار بکە بە شتێکی گەڕانپێکراو، پاشان هەموو وەڵامێک لەو لاپەڕەیەدا بپشکنە کە لێیەوە هاتووە.
- کاتی پێویست
- 120 خولەک · 26 هەنگاو
- چی دەخوازێت
- پێویستی بە دامەزراندنی نەرمامێرە
- تێچوو
- ڕێگایەکی بەخۆڕایی هەیە
- دوایین پشکنین
- 2026-08-31
- کاری یاسایی و لێپرسینەوە
- تۆمارکردن و چاودێری
- توێژەر یان شیکەر
- بەدواداچوونی یاسا و دادگا و ڕۆژنامەی فەرمی
- چاودێری هەواڵ و سۆشیال میدیا بە شێوەیەکی فراوان
- تێگەیشتن لە کۆمەڵێکی گەورەی دۆکیومێنت
- توێژینەوە
- پشکنینی وێنە و ڤیدیۆ و بانگەشەکان
- پاراستنی زانیارییە هەستیارەکان
ئەمانە ڕێنمایین، نەک لێکۆڵینەوەی حاڵەت
ڕەچەتەیەک پێت دەڵێت چی بکەیت. ڕاپۆرتی شتێک نییە کە پێشتر لە شوێنێکی تردا ڕوویدابێت. هیچ شتێک لێرەدا نایڵێت کە ڕێکخراوێک ئەمەی کردووە؛ بۆ ئەوە نووسراوە کە تۆ ئێستا جێبەجێی بکەیت، و هەموو هەنگاوێک پێت دەڵێت چۆن دڵنیا ببیتەوە کە کاری کردووە.
ڕۆژنامەیەکی فەرمی، بڕیارێکی دادگا، کۆمەڵێک فەرمانی سکانکراو، و هیچ ڕێگایەک نییە بۆ گەڕان لە ناویاندا، چونکە وێنەی دەقن نەک دەق. ئەم ڕەچەتەیە دەیانکات بە شتێکی گەڕانپێکراو، پاشان وات لێ دەکات وەڵامی ئامێرەکە لەو لاپەڕەیەدا بپشکنیت کە لێیەوە هاتووە. هەموو هەنگاوێک کە دوو ڕێگای هەبێت، ڕێگای ویندۆز یەکەم دەدات.
هێڵی سەلامەتی
هەنگاوی یەکەم پۆلێنکردنی کۆمەڵەکەیە، و ئەویش هەموو ڕەچەتەکەیە. سیفەتی «گشتی» بە بەڵگەنامەکەوە دەلکێت، نەک بە سەلامەتی ئەو کەسانەی ناویان تێیدا هاتووە: بڕیارێکی بڵاوکراوە کە ناوی دەستبەسەرکراوێک یان پەنابەرێکی هێشتا زیندووی تێدا بێت، بەڵگەنامەیەکی گشتییە کە شوێنی لە بوخچەی هەستیاردایە.
هەرگیز بەڵگەنامەیەک کە ناوی کەسێکی زیندوو، شوێنی، شایەتییەکەی، یان دۆخی پزیشکی و یاساییەکەی تێدا بێت مەخە سەر هیچ ئامرازێکی هەوریی ژیریی دەستکرد، بەخۆڕایی بێت یان پارەدراو. و ئەم ئەرشیفە گەڕانپێکراوە لەسەر لاپتۆپێک دروست مەکە کە دیسکەکەی شێواو نەکرابێت.
هیچ کورتکردنەوە یان ئاماژەیەکی ژیریی دەستکرد ئەنجامێکی یاسایی نییە. ئامرازێکی دۆزینەوەیە کە پێت دەڵێت کام لاپەڕە بکەیتەوە. پاشان مرۆڤێک ئەو لاپەڕەیە دەکاتەوە و دەیخوێنێتەوە پێش ئەوەی هیچ شتێک تۆمار بکرێت، بڵاو بکرێتەوە یان پشتی پێ ببەسترێت.
دەربارەی کوردی
ناسینەوەی وێنەیی بۆ سۆرانی چارەسەر نەکراوە، و هیچ ئامرازێکی بەخۆڕایی لەم ڕەچەتەیەدا بە دروستی مامەڵەی لەگەڵدا ناکات. هیچ مۆدێلێکی سۆرانیی چاودێریکراو بۆ TesseractTesseract OCR (running on your own computer)بۆ زانیاری هەستیار سەلامەتەسەرچاوە کراوەApache-2.0وەشانێکی خۆڕایی هەیە: Completely free, no usage limits.بینینی ئامرازەکە نییە. هەنگاوی حەڤدەیەم داوات لێ دەکات بە یەک فەرمان لەسەر ئامێری خۆت ئەوە بسەلمێنیت، لە جیاتی ئەوەی باوەڕمان پێ بکەیت.
وردەکارییە تەواوەکانی خوارەوە، هەنگاوەکان، فەرمانەکان و سەرچاوەکان، بە ئینگلیزی بڵاوکراونەتەوە، چونکە فەرمانەکان و ناوی دوگمەکان لە خودی نەرمامێرەکاندا بە ئینگلیزین، و ژمارەی وەشانێک کە لە نێوان سێ وەرگێڕاندا جیاواز بێت ژمارەیەکە کە کەس متمانەی پێ ناکات. پێشەکی و پوختەی سەلامەتی لە سەرەوەی پەڕەکە وەرگێڕدراون.
چی بەسەر زانیارییەکانتدا دێت
هەرگیز ئەمە بەکارمەهێنە بۆ
Never upload a document containing a living person's name, location, testimony, medical or legal status to any cloud AI tool, free tier or paid, no matter how convenient it is. Note that this rule and the PUBLIC/SENSITIVE sort in Step 1 can collide, and Step 1 resolves the collision explicitly: public status attaches to the DOCUMENT, not to the safety of the individuals named in it. A published judgment naming a still-living detainee or asylum seeker is a public document that still belongs in the SENSITIVE folder.
Never build this searchable archive on a laptop whose disk is not encrypted. See Step 2.
And never treat any AI summary or citation as a legal conclusion: it is a finding aid that tells you which page to open, and a human must open that page and read it before anything is filed, published or relied on.
چی لە ئامێرەکەت دەردەچێت
کام یاسا لەسەری جێبەجێ دەبێت · بژاردەی تەواو ناوخۆیی، بەبێ ئینتەرنێت
This depends entirely on which lane you put a document in, which is why Step 1 is sorting the pile.
OFFLINE LANE (Steps 11-24): the documents and your questions stay on your own disk. Tesseract and OCRmyPDF read the PDF on your own disk and write a new PDF and a text file on your own disk. You can physically disconnect the internet after installation and the whole lane still works. Verify this yourself rather than believing it: turn off Wi-Fi, then run the Step 18 command. It will finish normally.
Two honest qualifications to 'nothing leaves the device', both of which the previous version of this recipe got wrong:
(a) Your Desktop and Documents folders may not be on your device. On most Windows 10 and 11 machines signed into a Microsoft account, OneDrive Known Folder Move is on by default and the Desktop IS a cloud folder, copying a witness statement onto it uploads it to Microsoft, silently, with no error message and no prompt. The same applies to iCloud 'Desktop & Documents Folders' sync on macOS. The previous version of this recipe told readers to put their documents on the Desktop. That instruction caused the exact harm this recipe exists to prevent, and it is gone: Step 3 now creates a plain local folder outside the synced ones, and gives you a way to check.
(b) The optional tool in Step 23, AnythingLLM Desktop, sends anonymous telemetry unless you turn it off. Its marketing page says 'Your models, documents, and chat history stay on your machine. Nothing phones home.' Its own README says telemetry is something you 'opt out' of, by setting DISABLE_TELEMETRY or 'in-app by going to the sidebar > Privacy and disabling telemetry'. Both quotes verified 30 Aug 2026. What is transmitted is metadata, not content: installation type, the fact that 'a document is added or removed. No information about the document', the vector database type, the LLM provider and model tag, and the fact that 'a chat is sent'. That is not your documents, but for an audience where 86% fear information leaving the organisation, 'the app phones home unless you turn it off' has to be an instruction, not an omission. Step 23 makes turning it off a required sub-step before you add any document.
ONLINE LANE (Steps 5-10): the complete PDF is uploaded to Google. Every page, every stamp, every name printed on it, plus the questions you type about it. On the free tier Google's own pricing page states that content is 'used to improve our products'. That means a scanned page you upload can influence a future model. There is no undo.
AND, EVEN FOR A PUBLIC DOCUMENT: uploading it leaks nothing about the document, the world already has it, but it does reveal what you are investigating. Which gazette issue, which decree number, which case, on which date, from which account and which IP address. For a human rights organisation the pattern of interest is often more sensitive than any single document in it. Correct that framing in your own head before Step 5: the document is public, your research agenda is not.
A client file, a witness statement, an internal memo, a list of detainees, a complaint with a phone number in the margin. Those never go into the online lane, at any price tier.
کام یاسا لەسەری جێبەجێ دەبێت
Online lane: Google LLC, United States, subject to US law and US legal process. Data may be processed in any Google region. Google's Iraq availability does not mean Iraqi data protection law governs the processing.
Offline lane: no third-party jurisdiction. The data never becomes anyone else's to hold. Tesseract is Apache 2.0, OCRmyPDF is MPL-2.0, Ghostscript is available free under the AGPL, none of them phone home during OCR.
But 'no jurisdiction' is not the same as 'safe', and this is the gap the previous version of this recipe left open. For an Iraqi documenter or journalist the dominant threat is usually not a cloud leak. It is a laptop taken at a checkpoint, at a border, or in an office raid. Everything this recipe produces is plain readable text at rest: the text layer inside each searchable- PDF, the .txt sidecar files, and the storage folder of the optional tool in Step 23. A searchable, indexed, full-text archive of witness statements on an unencrypted laptop is a worse object to lose than the shoebox of paper it replaced, because it is instantly searchable by whoever takes it. That is why full-disk encryption is Step 2, above every software install, and why Step 24 tells you what to delete when a project ends.
One licensing caveat, for anyone who follows up on the Surya OCR system cited in the evidence section below: its code is Apache 2.0, but its model weights use a modified AI Pubs Open Rail-M licence described on its own project page as 'free for research, personal use, and startups under $5M funding/revenue'. A civil society organisation is comfortably inside that, but read it before building anything on top of it.
بژاردەی تەواو ناوخۆیی، بەبێ ئینتەرنێت
Yes, and it is the spine of this recipe rather than a footnote. Steps 11-24 are a complete, unlimited, no-account, offline pipeline: Tesseract adds an Arabic text layer, OCRmyPDF wraps that into searchable PDFs that still look exactly like the original scan and writes a plain .txt transcript beside each one, and Step 22 searches every one of those transcripts at once with a single command that needs no additional software on any platform. Add the optional AnythingLLM Desktop in Step 23, but only on a 16 GB machine, if you want to ask questions in sentences instead of searching for words.
The honest cost of staying offline: on a badly photocopied 1980s gazette page, free offline Arabic OCR is noticeably worse than Google's, and a small local language model writes weak Arabic. You trade output quality for the guarantee that nothing left the room. For documents containing a real person's name, that is the right trade.
تێچوو
پارەدان لە عێراقەوە
Availability from Iraq is genuinely fine for the tools used here, and was checked, not assumed: - Google's Gemini API / AI Studio available-regions page lists Iraq explicitly, immediately before Ireland (verified 30 Aug 2026). - Anthropic's supported-countries page lists Iraq in both the Claude.ai section and the API section (verified 30 Aug 2026). So you can create free accounts and use the free tiers from Iraq without a VPN.
Payment is the part that breaks. Iraqi-issued cards are frequently declined by international SaaS billing processors, and one of the 14 organisations surveyed named 'payment or access restrictions in our country' as its blocker. None of the free tiers above ask for a card, so none of them expose you to that. If you ever do need a paid tier, expect to need a foreign card or a colleague abroad, and plan for it rather than discovering it at the checkout page at 11pm before a filing deadline.
TWO ACCESS FRICTIONS THAT ARE NOT PAYMENT, AND THAT THE PREVIOUS VERSION DID NOT WARN ABOUT:
1. Phone verification. Creating a new Google account from Iraq frequently demands an SMS code, and a newly created or long-dormant account can be locked pending phone or identity re-verification. If you have never made a Google account, budget for this happening BEFORE Step 5, not during it. If you cannot pass it, the online lane is simply closed to you, and the offline lane does everything you actually need.
2. Download volume, on a metered or intermittent Erbil connection. State this before you start, because two of the fourteen organisations surveyed are blocked by internet or hardware. Windows: about 250 MB total (Python roughly 25 MB, the Tesseract installer roughly 50-70 MB, Ghostscript 10.07.1's 64-bit installer 64.9 MB as listed on its release page, and the OCRmyPDF Python packages). macOS: about 1.5 GB, because Homebrew's tesseract-lang package carries every language, measured at 654 MB installed on the machine used to write this recipe, plus Ghostscript at 258 MB, plus Homebrew itself and possibly Apple's command line developer tools. The optional tool in Step 23 adds a further multi-gigabyte model download, which is one more reason it is optional.
The offline route involves no payment, no account and no geographic check at all. After the downloads above, it works on a laptop in Erbil with the Wi-Fi switched off, and Step 18's check tells you to prove that to yourself.
ڕێگا بەخۆڕاییەکە
ئەگەر پارە بدەیت
The entire recipe can be done for USD 0, and the offline half needs no account at all.
OFFLINE, NO ACCOUNT, NO CARD, UNLIMITED: Tesseract OCR, OCRmyPDF and Ghostscript are open-source and free forever, with no page cap, no file cap and no sign-up. This is the route that never runs out. Verified 30 Aug 2026: ocrmypdf 17.11.0 is on PyPI (published 28 Aug 2026, needs Python 3.11 or newer); the Homebrew formulae ocrmypdf, tesseract-lang, unpaper and ghostscript all exist and install; Ghostscript 10.07.1 (released 19 May 2026) is downloadable free under the AGPL from ghostscript.com.
ONLINE FREE TIERS (public documents only): - Google AI Studio: 'Google AI Studio usage is free of charge in all available regions' (verified verbatim on ai.google.dev/gemini-api/docs/pricing, 30 Aug 2026). The Gemini Flash and Flash-Lite model families are listed as 'Free of charge' on the free tier. Hard trade-off, stated by Google on that same page as a row in its own table: free tier 'Content used to improve our products'; paid tier 'Content not used to improve our products'. - Gemini Notebook (the product formerly called NotebookLM) free plan, verified on Google's own support page 30 Aug 2026: 100 notebooks, up to 50 sources per notebook, 500,000 words per source, 50 chat queries per day, 3 audio generations per day. - Google Drive PDF-to-Docs OCR: free, but Google's help page states 'The file should be 2 MB or smaller', which rules out most gazette scans.
What costs money and is NOT needed here: any paid subscription to a cloud assistant. Nothing in this recipe asks for a card.
REMOVED SINCE THE LAST VERSION ON COST GROUNDS: Recoll, the desktop search tool. Its own Windows page says 'I recommend contributing a small fee for downloading and using the Windows binary version' (verified 30 Aug 2026). The author does offer it free on request by email, 'If the payment is a problem for you for any reason, please send me an email, no need for justification', but that means emailing a stranger from your work address before you can search your own archive, and Windows is the majority platform in this room. Step 22 now uses a method that needs no download at all.
ئەگەر پارە بدەیت
USD 0 for everything in this recipe. There is no paid step and no upgrade path you need.
No monthly figure for paid cloud assistants is printed here. The previous version of this recipe published a 'USD 20-22 per month' range while admitting in the same sentence that the vendors' pricing pages had returned HTTP 403 and could not actually be read. Publishing a price nobody opened is the exact failure this recipe warns readers about, so the number has been removed rather than softened. If your organisation later wants a paid tier, open the vendor's pricing page in your own browser on the day you decide, and read the figure yourself.
For an organisation where 29% spend nothing, the honest answer is that the offline route in Steps 11-24 is not a downgrade. For documents with a living person's name in them it is the correct route regardless of budget.
پێش ئەوەی دەست پێ بکەیت
- A laptop (Windows or Mac). The offline lane genuinely needs a laptop. It cannot be done on a phone. About 4 GB of free disk space for the software and your PDFs.
- A laptop whose disk you are willing and able to encrypt. This is Step 2 and it is not optional. If you cannot turn on BitLocker or FileVault on this machine, do not build a searchable archive of sensitive documents on it.
- Your documents as PDF or image files in one folder. If they are still paper, scan them at 300 dpi in black and white or greyscale, not colour, and not at 150 dpi. Scan quality sets the ceiling on everything that follows.
- Roughly 90-120 minutes the first time, most of it waiting for downloads. Steps 1-4 and Steps 25-26 are paper-and-thinking work and take about 20 minutes on their own. You can do those before you install anything. After the first time, running the pipeline on a new pile takes about 10 minutes of your attention.
- Internet for the installation downloads only, and enough of it: about 250 MB on Windows, about 1.5 GB on Mac. On a metered or intermittent connection, do the installs where the connection is good. After that the offline lane runs with the Wi-Fi switched off.
- A free Google account, ONLY if you intend to use the online lane for already-public documents. Not needed for the offline lane. Be warned that creating one from Iraq often requires SMS verification, see the Iraq payment and access note.
- One colleague who reads the language of the documents fluently and has agreed to do the verification in Step 25. This is a real prerequisite, not a nice-to-have. Without it, do not run this recipe.
لەسەر چ ئامێرێک کار دەکات
WINDOWS IS THE MAJORITY PLATFORM AND IS FULLY SUPPORTED. Every step in this recipe works on Windows. Windows instructions come first everywhere.
WHAT DIFFERS BY PLATFORM:
Windows. Four separate installs: Tesseract (Step 12), Python and Ghostscript and OCRmyPDF (Step 13). Ghostscript is NOT available through winget and must be downloaded by hand, an earlier version of this recipe told you to run `winget install -e --id ArtifexSoftware.GhostScript`, which was removed from the winget catalogue on 10 October 2023 and now fails with "No package found matching input criteria". Total download about 250 MB. Windows leaves out the `--clean` option because it needs `unpaper`, which is awkward to install on Windows; `--deskew` and `--rotate-pages` do most of the same work. Windows has one extra hazard the other platforms do not: OneDrive. On most Windows 10 and 11 machines signed into a Microsoft account, your Desktop and Documents folders are cloud folders that upload automatically. Step 3 puts your work in C:\ocr-work at the root of the C: drive precisely to stay out of OneDrive.
macOS. One package manager (Homebrew) then one install command, but a much bigger download, roughly 1.5 GB, and more if Apple's command line developer tools are not already present. macOS has the same cloud-folder hazard under a different name: if iCloud "Desktop & Documents Folders" sync is on, those two folders are cloud folders. Step 3 uses a folder in your home directory instead. macOS gets `--clean` (unpaper installs cleanly through Homebrew).
Linux. Not written out step by step here, because nobody in the target room reported using it. If you are on Linux, install `ocrmypdf tesseract-ocr-ara unpaper ghostscript` through your distribution's package manager and every command from Step 18 onward is identical to the macOS line.
PLATFORM-LIMITED ROUTES, STATED BEFORE YOU INSTALL ANYTHING: - The optional local question-answering tool in Step 23 needs 16 GB of RAM on any platform. On an 8 GB laptop, skip it. It is optional and the recipe is complete without it. - The desktop search tool Recoll, recommended in the previous version of this recipe, has been REMOVED. Its Windows build asks for a payment contribution, and on Apple Silicon Macs the download is unsigned and refuses to run without manually code-signing it from the Terminal. Neither is acceptable for this audience. Step 22 replaces it with a search method that needs no extra software on any platform. - The online lane (Steps 5-10) is browser-only and works identically everywhere, including on a phone.
هەنگاوەکان
هەنگاو 1 / 26هەموو سیستەمەکان
Before touching any software, sort your pile into two folders. Name them exactly as below. Everything else in this recipe branches off this one decision.
ئەمە بە دروستی کۆپی بکە
Folder 1 name: PUBLIC-documents Folder 2 name: SENSITIVE-documentsدڵنیا ببەرەوە کە کاری کردووە
Every document is now inside one of the two folders and none are left loose. Count them: the two folders together must add up to the number of documents you started with. If you find yourself with a third pile of 'not sure', that pile is SENSITIVE.
تێبینی
PUBLIC means the document is already published to the world: an official gazette issue, a published judgment, a bill, a ministry circular, a public tender notice. SENSITIVE means a real person can be identified: client files, witness statements, complaints, internal memos, anything with a name, phone number, address or case number belonging to a living individual.
THE HARD CASE, WHICH THE PREVIOUS VERSION OF THIS RECIPE GOT WRONG. Public status attaches to the DOCUMENT, not to the safety of the people named in it. Published judgments routinely name parties. Iraqi gazettes publish named appointments, dismissals, and in some periods citizenship and property decisions naming individuals. So: if the document is public but a named person could be harmed by your interest in them becoming visible, a detainee, an asylum seeker, a dismissed official, a witness, anyone whose situation is still live. It goes in SENSITIVE, even though anyone can buy a copy at a kiosk. Uploading it would not leak the document. It would leak that you are looking at that person.
If you hesitate for more than five seconds about which folder a document belongs in, it goes in SENSITIVE. 86% of the organisations surveyed named information leaving the organisation as their top fear. This folder is where you act on that fear instead of worrying about it.
هەنگاو 2 / 26هەموو سیستەمەکان
Turn on full-disk encryption BEFORE you install anything or copy any document. This recipe is about to turn a pile of unreadable scans into a fast, searchable, full-text archive. That is exactly as useful to whoever takes the laptop as it is to you.
ئەمە بە دروستی کۆپی بکە
WINDOWS: Settings > Privacy & security > Device encryption -> switch it On If there is no 'Device encryption' entry: press Start, type Manage BitLocker and turn on BitLocker for drive C: MAC: System Settings > Privacy & Security > FileVault -> Turn On Both: set a strong login password at the same time. Encryption does nothing if the machine is left logged in and unlocked.دڵنیا ببەرەوە کە کاری کردووە
Windows: the Device encryption switch reads On, OR the Manage BitLocker panel shows 'BitLocker on' for drive C:. Mac: the FileVault panel says FileVault is turned on for your disk. Encryption can take an hour or more to finish in the background after you switch it on. You can carry on with the recipe, but do not copy sensitive documents onto the machine until the panel stops saying it is still encrypting. If neither setting exists on your edition of Windows, stop and read the note before you put any sensitive document on this machine.
تێبینی
Why this is Step 2 and not an appendix. For Iraqi documenters and journalists the dominant threat is usually not a cloud leak. It is a laptop taken at a checkpoint, at a border, or in an office raid. Everything this recipe creates is plain readable text at rest: the text layer inside each searchable PDF, the .txt transcripts from Step 20, and the storage folder of the optional tool in Step 23. Without encryption you have made your archive more useful to an adversary, not less.
RECOVERY KEY WARNING, which matters for this audience. Windows backs up the BitLocker recovery key to your Microsoft account by default; macOS offers to escrow the FileVault key with iCloud. That means Microsoft or Apple can be compelled to hand over the key. When you are offered the choice, pick the local option: on Mac choose 'Create a local recovery key' rather than the iCloud option, and write it on paper stored somewhere other than the laptop bag. On Windows you can print or save the key to a file on a USB stick. If you lose the key and the password, the data is gone permanently. That is the whole point of it, so store the paper carefully.
If your edition of Windows offers neither Device encryption nor BitLocker, use VeraCrypt (free, open source, veracrypt.io, version 1.26.29, released 9 June 2026, verified 30 Aug 2026) to make an encrypted container file and keep the work folder from Step 3 inside it. That is more clicking, and it is better than nothing by a very large margin.
هەنگاو 3 / 26هەموو سیستەمەکان
Make a plain local work folder that is NOT inside OneDrive or iCloud, and put your documents there. Do not use the Desktop. Do not use Documents.
ئەمە بە دروستی کۆپی بکە
WINDOWS (File Explorer): Open This PC > Local Disk (C:) and create a new folder named: ocr-work Then inside it create another folder named: docs The address bar must read exactly: C:\ocr-work\docs MAC (Finder): Press Shift + Command + H to open your home folder, create a new folder named: ocr-work Then inside it create another folder named: docs The path must be: /Users/YOURNAME/ocr-work/docs Now copy the PDFs you want to process into that docs folder.دڵنیا ببەرەوە کە کاری کردووە
WINDOWS: click into the folder and read the address bar. It must say C:\ocr-work\docs. If it says anything containing OneDrive, you are inside a cloud folder, delete it and create it again at the root of the C: drive. Second check: right-click the ocr-work folder. If you see a green tick, a cloud icon, or a 'Free up space' / 'Always keep on this device' option in the menu, it is being synced to Microsoft. Move it.
MAC: the folder must sit directly in your home folder alongside Desktop, Documents and Downloads, not inside them. To confirm iCloud is not involved, open System Settings > [your name] > iCloud > Drive and look at 'Desktop & Documents Folders'. If that is on, those two folders are cloud folders; ~/ocr-work is not, which is why we use it.
تێبینی
This is the single most dangerous default in the whole exercise, and the previous version of this recipe walked straight into it by telling readers to put test.pdf and a docs folder on the Desktop.
On most Windows 10 and 11 machines signed into a Microsoft account, a feature called OneDrive Known Folder Move is on by default and your Desktop IS a OneDrive folder. Copying a witness statement onto it uploads that witness statement to Microsoft. There is no prompt, no warning and no error message; the file simply appears in OneDrive later. The same applies on macOS when iCloud 'Desktop & Documents Folders' sync is enabled.
So the recipe that was written to stop sensitive material leaving the organisation was, in its own instructions, causing sensitive material to leave the organisation. Making one folder in the right place costs thirty seconds and closes it.
هەنگاو 4 / 26هەموو سیستەمەکان
Find out what kind of PDFs you actually have. Open one PDF and try to select a line of text with your mouse, as if you were going to copy it. Then paste what you selected into Notepad (Windows) or TextEdit (Mac).
ئەمە بە دروستی کۆپی بکە
Open one PDF -> drag the mouse across a line of text -> Ctrl+C (Windows) or Command+C (Mac) -> paste into Notepad / TextEditدڵنیا ببەرەوە کە کاری کردووە
You are now in exactly one of three cases, and you need to write down which one each file is in:
(a) The text highlighted letter by letter and pasted correctly. This PDF already has real text. It needs no OCR at all, go straight to Step 22 to search it, or Step 9 if it is public and you want to ask questions about it.
(b) Nothing highlighted; your mouse just drew a blue rectangle over a picture. It is a scan, a photograph of words. No search tool on earth can find anything in it until a text layer is added. That is what Steps 18 to 21 do.
(c) Text highlighted, but what you pasted is garbage, random letters, boxes, or Latin characters where Arabic should be. This PDF has a BAD text layer from earlier, poor OCR. This is the case people miss, and Step 18 gives you a different command for it.
تێبینی
Most Iraqi official gazette PDFs are case (b). Old digitisation projects sometimes produce case (c), and case (c) is dangerous because a search over it silently returns nothing while the tool reports success.
If your documents are public and you want the better-quality online route, the online lane is Steps 5 to 10. If they are sensitive, skip Steps 5 to 10 entirely and go to Step 11.
هەنگاو 5 / 26هەموو سیستەمەکان
ONLINE LANE, PUBLIC DOCUMENTS ONLY. If your documents are in the SENSITIVE folder, skip Steps 5 to 10 completely and go to Step 11. If they are public, open Google AI Studio in your browser and sign in with a free Google account.
ئەمە بە دروستی کۆپی بکە
https://aistudio.google.comدڵنیا ببەرەوە کە کاری کردووە
You can see a chat box with a model name at the top of the page, and at no point were you asked for a credit card. If Google demanded an SMS verification code and you could not receive it, stop here: the online lane is closed to you, and the offline lane from Step 11 does everything you genuinely need. Do not try to work around a verification wall with a borrowed number, an account you do not control is a worse place for your research history than no account.
تێبینی
Verified 30 Aug 2026: Google's own available-regions page lists Iraq explicitly, immediately before Ireland, so this works from Erbil or Baghdad without a VPN, and Google's pricing page states verbatim 'Google AI Studio usage is free of charge in all available regions'. No credit card is requested, which sidesteps the Iraqi-card rejection problem entirely.
The price you pay instead is stated by Google on the same pricing page, as a row in its own comparison table: free tier, 'Content used to improve our products'; paid tier, 'Content not used to improve our products'. For a document that is already published, that costs you nothing.
But be clear about the second thing you are spending. Uploading a public document leaks nothing about the document, the world already has it. It does reveal what you are investigating: which gazette issue, which decree number, which case, on which date, tied to your account and your IP address. For a human rights organisation the pattern of interest is often more sensitive than any single document in it. Decide that you are comfortable with that before you upload, not after.
هەنگاو 6 / 26هەموو سیستەمەکان
Upload one public PDF using the paperclip or 'Insert' button, then paste this exact prompt and send it. Do one document first, not the whole pile.
ئەمە بە دروستی کۆپی بکە
Transcribe this document verbatim into Arabic text. Do not summarise, do not correct, do not modernise the wording. Preserve the original line and paragraph structure. Mark the start of every page as [صفحة 1], [صفحة 2] and so on. Where the scan is unreadable, write [غير مقروء] instead of guessing the word. Do not invent any text that you cannot actually see.دڵنیا ببەرەوە کە کاری کردووە
The reply is Arabic text that visibly corresponds to your page, it contains at least a [صفحة 1] marker, and it ends at the end of a sentence rather than stopping mid-word. If it stopped mid-word or mid-sentence, it ran out of room, go straight to Step 7, which you must do anyway. Save the result into a plain text file inside your ocr-work folder, next to the PDF.
تێبینی
The two instructions that matter most are the last two. A model asked to transcribe a smudged line will happily invent a plausible legal phrase to fill the gap, and you will never spot it because it reads perfectly. Forcing it to write [غير مقروء] converts an invisible fabrication into a visible hole you can go and check with your own eyes.
هەنگاو 7 / 26هەموو سیستەمەکان
Count the pages. Compare the number of page markers in the transcript against the number of pages in the actual PDF. Do this before you read a single word of the transcript.
ئەمە بە دروستی کۆپی بکە
1. Open the PDF and note its total page count, e.g. 40 pages. 2. In the transcript, find the LAST page marker: [صفحة 40] ? 3. The last marker number must equal the PDF's page count. 4. Also check that no marker numbers are missing in the middle: 1, 2, 3 ... with no gaps.دڵنیا ببەرەوە کە کاری کردووە
Last marker number equals the PDF page count, and there are no gaps in the sequence. If the PDF has 40 pages and the last marker is [صفحة 12], then 28 pages are missing and nothing told you. If markers jump from [صفحة 5] to [صفحة 9], four pages were skipped.
تێبینی
This step is new, and it exists because the previous version's spot-check could not catch the failure it most needed to catch.
Every model has a maximum reply length. Give it a forty-page gazette and it will transcribe until it hits that limit and then stop. The stopping point looks like a perfectly clean ending, the last paragraph is well-formed, the Arabic is fluent, nothing is marked. If your only check is to inspect 'the start, the middle and the end' of the transcript, you inspect the end of the transcript, which is fine, and never discover that it is not the end of the document.
A missing second half is a safety problem, not a quality problem: it is how an organisation concludes that a gazette contains no reference to something when in fact the model never read that part.
THE FIX when it truncates: do it in chunks. Re-send the Step 6 prompt with the extra line 'Transcribe pages 11 to 20 only.' and repeat until the page markers add up. Ten pages at a time is a reasonable starting size. Then join the chunks and run this count again on the joined file.
هەنگاو 8 / 26هەموو سیستەمەکان
Spot-check the transcription before you trust any of it. Pick three random places in the transcript, near the start, the middle and the end, and find the same passage on the actual scanned page with your own eyes.
دڵنیا ببەرەوە کە کاری کردووە
For each of the three passages you can point at the same words on the scanned image. If you cannot find one of them anywhere on the page it claims to come from, you have found a fabrication and the transcript is not usable for legal work.
تێبینی
You are checking for two different failures. Small character errors (hamza, taa marbuta) are normal and harmless for searching. Whole invented sentences, a paragraph that appears in the transcript but not on the page, or a page that quietly got skipped are disqualifying, if you find one, fall back to the offline lane and read the pages yourself. Do this check every time. It takes four minutes and it is the difference between a research tool and a liability.
هەنگاو 9 / 26هەموو سیستەمەکان
To work across a whole set of public documents rather than one, open Gemini Notebook, create a new notebook, and upload your public PDFs as sources.
ئەمە بە دروستی کۆپی بکە
https://notebook.google.comدڵنیا ببەرەوە کە کاری کردووە
The notebook lists your uploaded files as sources, and when you ask a question the answer carries clickable citations that jump into the source document. If a file uploaded but the notebook cannot quote anything at all from it, that file is an image-only scan and the notebook is not reading it, see the note.
تێبینی
Use the address exactly as written above. The previous version of this recipe had a misspelled domain in this field, which for people in this line of work is a security problem and not a typo: a mistyped address is how a lookalike site gets your Google password.
Name and address, verified 30 Aug 2026: the product Google used to call NotebookLM is now called Gemini Notebook in Google's own support documentation, and notebooklm.google.com issues a 301 redirect to notebook.google.com.
Verified free-plan limits from Google's support page, 30 Aug 2026: 100 notebooks, up to 50 sources per notebook, 500,000 words per source, 50 chat queries per day, 3 audio generations per day. Fifty sources is enough for a year of one gazette or one case file's worth of public filings; it is not enough for a whole archive, so split by year or by topic into separate notebooks.
Its advantage for legal work is specific and real: answers carry clickable citations back into the source document, so verification is one click rather than a hunt. Its limit is equally specific: it cannot read an image-only scan any better than anything else. If your PDF is a pure scan, either upload the transcript text file from Step 6, or OCR the file first using Steps 11 to 21 and upload the searchable- version.
هەنگاو 10 / 26هەموو سیستەمەکان
Ask your question, but require a citation in the question itself. Paste this template and fill in your own question where marked.
ئەمە بە دروستی کۆپی بکە
Answer this question using only the documents I have given you: [YOUR QUESTION HERE]. For every single statement in your answer, give me: the document name, the page number, and the exact sentence in the original Arabic that supports it. If the documents do not answer the question, say clearly "غير موجود في المستندات" and stop. Do not use any knowledge from outside these documents. Do not tell me what the law probably says.دڵنیا ببەرەوە کە کاری کردووە
Every sentence in the answer carries all three things: a document name, a page number, and a quoted Arabic sentence. If any sentence in the answer has no citation attached, that sentence is not usable, ask again, or treat it as absent. An answer where the model says غير موجود في المستندات is a good answer, not a failure.
تێبینی
'Do not tell me what the law probably says' is doing heavy lifting. Left to itself, a model will smoothly blend what it read in your gazette with what it half-remembers about Iraqi law generally, and the join is invisible. Demanding the exact original sentence for every claim gives you something checkable: a fabricated citation usually cannot produce a verbatim Arabic sentence that actually appears on the page you open. Step 25 is where you actually open it.
هەنگاو 11 / 26هەموو سیستەمەکان
OFFLINE LANE, START HERE FOR SENSITIVE DOCUMENTS. This lane is where you go when the documents name a living person. Read this routing box carefully; it is the one place people get lost.
ئەمە بە دروستی کۆپی بکە
IF YOU ARE ON WINDOWS -> your next step is Step 12. IF YOU ARE ON A MAC -> your next step is Step 15. Both platforms meet again at Step 18. Windows readers: do not run any command beginning with 'brew'. Those are Mac-only and will fail. Mac readers: do not run any command beginning with 'winget' or 'py'. Those are Windows-only and will fail.دڵنیا ببەرەوە کە کاری کردووە
You know which single number you are going to next, and it is 12 if you are on Windows. The previous version of this recipe sent Windows readers to the macOS Homebrew installer at this exact point, which meant pasting a Unix command into PowerShell and getting an incomprehensible error. If you find yourself looking at a command that starts with /bin/bash on a Windows machine, you are in the wrong section.
تێبینی
Download volume before you begin, so you can choose where to do it. Windows: about 250 MB in total. Mac: about 1.5 GB, and more if Apple's command line developer tools are not already installed. On a metered or intermittent connection, do the installs somewhere with a good connection and then take the laptop away. Everything after Step 18 works with the Wi-Fi switched off, and Step 18's check asks you to prove that.
هەنگاو 12 / 26ویندۆز
WINDOWS ONLY. Install Tesseract, the OCR engine, together with the Arabic language data. Open your browser, go to the address below, and download the 64-bit Windows installer (the file ending in .exe).
ئەمە بە دروستی کۆپی بکە
https://github.com/UB-Mannheim/tesseract/wiki The current file is named like: tesseract-ocr-w64-setup-5.5.3.20260724.exe (the version number will be higher by the time you read this, take the newest 64-bit one)دڵنیا ببەرەوە کە کاری کردووە
After the installer finishes, open File Explorer and confirm that the folder C:\Program Files\Tesseract-OCR exists and contains tesseract.exe. The full proof that Arabic installed comes in Step 14, do not skip it.
تێبینی
CRITICAL, and the single most common point of failure: during the installer, when you reach the screen listing components, expand the section called 'Additional language data (download)' and TICK 'Arabic'. If you click Next past this screen without ticking Arabic, Tesseract installs English-only, and every Arabic OCR command later will fail with a message about a missing language.
Also, on the installer screen that offers it, choose the option to add Tesseract to the system PATH (wording varies by version; it may appear as 'Add to PATH' or as an install-for-all-users choice). If you are not offered it, do not worry, Step 14 shows you how to check and how to fix it.
IF YOU ALREADY MADE THE ARABIC MISTAKE, do not reinstall. Download the Arabic model by hand from the link below and put the file into C:\Program Files\Tesseract-OCR\tessdata (Windows will ask for administrator permission to write there, which is normal):
https://github.com/tesseract-ocr/tessdata_best/raw/main/ara.traineddata
That file is 12.6 MB, verified present in the tessdata_best repository on 30 Aug 2026. The tessdata_best models are the slowest but the most accurate, which is what you want on damaged gazette scans.
هەنگاو 13 / 26ویندۆز
WINDOWS ONLY. Install Python, then Ghostscript by hand, then OCRmyPDF. Do the three parts in this order. Use a NORMAL PowerShell window, not 'Run as administrator'.
ئەمە بە دروستی کۆپی بکە
PART 1, Python. Click Start, type PowerShell , open 'Windows PowerShell' NORMALLY (do NOT run as administrator). Paste and press Enter: winget install -e --id Python.Python.3.12 PART 2, Ghostscript. This one is NOT available through winget. Download it in your browser: https://ghostscript.com/releases/gsdnld.html Choose the 'Ghostscript AGPL Release' for Windows (64 bit). The file is named like gs10071w64.exe (64.9 MB). Run it and accept the default install folder, which will look like C:\Program Files\gs\gs10.07.1 PART 3: OCRmyPDF. CLOSE every PowerShell window. Open a NEW normal PowerShell window. Then paste: py -m pip install ocrmypdfدڵنیا ببەرەوە کە کاری کردووە
Part 1 ends with 'Successfully installed'. Part 3 ends with a line beginning 'Successfully installed ocrmypdf-' followed by a version number (17.11.0 or higher). If Part 3 says 'py : The term "py" is not recognized', Python either did not install or you did not open a fresh window, close every PowerShell window, open exactly one new one, and run Part 3 again. Do not proceed to Step 14 until you have seen the 'Successfully installed ocrmypdf-' line.
تێبینی
WHY GHOSTSCRIPT IS DOWNLOADED BY HAND. The previous version of this recipe told you to run `winget install -e --id ArtifexSoftware.GhostScript` and attributed that command to OCRmyPDF's official documentation. Both halves were wrong. That package was removed from the winget catalogue on 10 October 2023 and the command now fails with 'No package found matching input criteria', verified 30 Aug 2026 against Microsoft's winget-pkgs repository, where the ArtifexSoftware folder now contains only 'mutool'. And OCRmyPDF's own installation page says the opposite of what was claimed: 'You will need to install Ghostscript manually, since it does not support automated installs anymore.' Ghostscript is a hard dependency, so on the old instructions every later OCR command failed too. Third-party winget mirrors still list the dead package, which is how the error survived a casual check.
WHY A NORMAL WINDOW, NOT ADMINISTRATOR. winget installs Python for your user account. An administrator PowerShell runs as a different account and may not see it, so `py -m pip` fails in a way that is very hard to diagnose. Use a normal window for all three parts.
WHY `py -m pip` AND NOT BARE `pip`. Bare `pip` is frequently not on PATH after a fresh Windows Python install. `py` is the Python launcher that Windows installs reliably, and `py -m pip` is the form OCRmyPDF's documentation recommends.
Verified 30 Aug 2026: ocrmypdf 17.11.0 published to PyPI on 28 Aug 2026, requires Python 3.11 or newer, so the 3.12 above is correct. Ghostscript 10.07.1 released 19 May 2026; the Windows 64-bit installer on its release page is gs10071w64.exe at 64.9 MB. The winget identifier Python.Python.3.12 was confirmed to still exist in the winget catalogue.
هەنگاو 14 / 26ویندۆز
WINDOWS ONLY. Prove that all three tools can actually be found before you try to OCR anything. Open a new normal PowerShell window and run these four commands one at a time.
ئەمە بە دروستی کۆپی بکە
tesseract --version tesseract --list-langs gswin64c --version ocrmypdf --versionدڵنیا ببەرەوە کە کاری کردووە
All four must print something, and one specific thing must appear:
1. `tesseract --version` prints a version line such as 'tesseract v5.5.3...'. 2. `tesseract --list-langs` prints a header line like: List of available languages in "C:\Program Files\Tesseract-OCR/tessdata/" (N): followed by one three-letter code per line. THE LINE `ara` MUST BE IN THAT LIST. If it is not, Arabic did not install, go back to Step 12 and use the manual ara.traineddata download. 3. `gswin64c --version` prints a number like 10.07.1. 4. `ocrmypdf --version` prints a number like 17.11.0.
If any of them says 'The term ... is not recognized as the name of a cmdlet', that program is installed but Windows cannot find it. See the note for the fix. Do not continue until all four print.
تێبینی
This step exists because a missing PATH entry is the single most common Windows problem with this toolchain, and it produces error messages later that mean nothing to a non-technical reader.
TO ADD SOMETHING TO PATH: Press Start, type environment variables , open 'Edit the system environment variables'. Click the 'Environment Variables...' button. In the lower box ('System variables'), select the row named 'Path' and click 'Edit...'. Click 'New' and type the folder that contains the missing program: for Ghostscript: C:\Program Files\gs\gs10.07.1\bin for Tesseract: C:\Program Files\Tesseract-OCR (If your Ghostscript version number is different, use the folder name you actually see in C:\Program Files\gs.) Click OK on all three windows. CLOSE PowerShell completely and open a new window, then run the four commands again.
The manual Ghostscript installer definitely does not add itself to PATH, so if exactly one of the four fails it is usually gswin64c.
هەنگاو 15 / 26ماک
MAC ONLY. Open the Terminal application: press Command and the space bar together, type the word Terminal, and press Enter. Then install Homebrew, the tool that installs the other tools. Copy this entire line, paste it, and press Enter. It will ask for your Mac login password, type it and press Enter. The password will NOT appear on screen as you type; that is normal, keep typing.
ئەمە بە دروستی کۆپی بکە
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"دڵنیا ببەرەوە کە کاری کردووە
When it finishes, type brew --version and press Enter. It must print something like 'Homebrew 4.x.x'. If it says 'command not found: brew', the installer printed two lines beginning with (echo near the end and told you to run them, scroll up, copy those two lines exactly as printed, paste them, press Enter, then close the Terminal window, open a fresh one, and try brew --version again.
تێبینی
If you have never opened this window before: it is a place where you type an instruction and press Enter, and the computer does it. It will print a lot of text while it works. That text is progress, not errors. You do not need to read it. When it stops and shows your username with a $ or % sign again, it has finished and is waiting for the next instruction.
Homebrew takes 5 to 15 minutes and downloads several hundred megabytes; it may also install Apple's command line developer tools, which is a further large download. Skip this step entirely if you already have Homebrew.
هەنگاو 16 / 26ماک
MAC ONLY. Install the OCR software and the Arabic language data. Paste this one line and press Enter, then wait.
ئەمە بە دروستی کۆپی بکە
brew install ocrmypdf tesseract-lang unpaper ghostscriptدڵنیا ببەرەوە کە کاری کردووە
The command returns you to the prompt without a line beginning 'Error:'. The real proof is Step 17, which you must do next.
تێبینی
This installs four things at once: OCRmyPDF (which adds a searchable text layer to a scanned PDF while leaving the scan looking identical), Tesseract's full language pack including Arabic, unpaper (which cleans up speckles and shadows on bad photocopies), and Ghostscript (which OCRmyPDF needs internally).
Verified 30 Aug 2026 against Homebrew's own formula listing: all four formulae exist and are current, ocrmypdf 17.11.0, tesseract-lang 4.1.0, unpaper 7.0.0, ghostscript 10.07.1. This matches OCRmyPDF's official macOS installation instructions, which give `brew install ocrmypdf` plus `brew install tesseract-lang` for non-English languages.
Expect 10-20 minutes. Measured on the machine used to write this recipe: tesseract-lang alone occupies 654 MB installed and Ghostscript 258 MB, so allow roughly 1 GB of disk on top of Homebrew itself. Everything installed here is free, open-source, and works with the internet switched off from now on.
هەنگاو 17 / 26ماک
MAC ONLY. Prove the tools work and that Arabic is really there. Run these two commands.
ئەمە بە دروستی کۆپی بکە
ocrmypdf --version tesseract --list-langsدڵنیا ببەرەوە کە کاری کردووە
`ocrmypdf --version` prints a version number such as 17.11.0. `tesseract --list-langs` prints a header line then a long list of three-letter language codes, and THE LINE `ara` MUST BE IN IT. On the machine used to write this recipe the header read: List of available languages in "/opt/homebrew/share/tessdata/" (163): and ara was present. If ara is missing, run brew install tesseract-lang again. If ocrmypdf prints a long red error about a library that could not be loaded rather than a version number, run brew reinstall ocrmypdf and try again.
تێبینی
While that list is on your screen, look for two other codes, because this is the Kurdish problem made visible in one command rather than argued about.
Search the list for `ckb` (Sorani, Central Kurdish). It is not there. Search for `kur`. It is not there either. Search for `kmr`. That one IS there, and it is Kurmanji, Northern Kurdish, written in Latin script, which is not what a Sorani document is written in.
That is a complete Tesseract language installation, 163 languages, with no Sorani model in it. On the machine used to write this recipe, on 30 Aug 2026, that is exactly what it printed. Read the Kurdish Sorani section of this recipe before you spend an afternoon on Sorani material.
هەنگاو 18 / 26هەموو سیستەمەکان
BOTH PLATFORMS MEET HERE. Test on ONE bad page before committing to the whole pile. Copy a single difficult PDF into your work folder, rename it test.pdf, then run the command that matches your platform AND the case you found in Step 4.
ئەمە بە دروستی کۆپی بکە
WINDOWS (PowerShell), for case (b), a page that is a pure picture: cd C:\ocr-work ocrmypdf -l ara --skip-text --deskew --rotate-pages --sidecar test.txt test.pdf test-searchable.pdf MAC (Terminal), same case (b): cd ~/ocr-work ocrmypdf -l ara --skip-text --deskew --clean --rotate-pages --sidecar test.txt test.pdf test-searchable.pdf BOTH PLATFORMS: for case (c), a PDF that already has a BAD, garbled text layer. Note that --deskew and --clean must be left OUT of this one: ocrmypdf -l ara --redo-ocr --sidecar test.txt test.pdf test-searchable.pdfدڵنیا ببەرەوە کە کاری کردووە
Three things must ALL be true. Check all three; the first one on its own is not enough.
1. The command finished with lines about 'Postprocessing' and an 'Output file is a PDF/A' message, and gave you the prompt back with no red error. 2. A file named test-searchable.pdf now exists in the folder. 3. A file named test.txt now exists AND IS NOT EMPTY. Open it in Notepad or TextEdit. You must see Arabic words in it that you can recognise from the page.
An empty or one-line test.txt after a run that reported success is the failure mode that matters most here: OCR found nothing, and nothing told you. It usually means the page is too degraded, the language code is wrong for what is on the page, or the page was blank. Try -l ara+eng, try a different page, and if five of your worst pages all come out empty, this archive needs better scans before it needs better software.
تێبینی
WHAT EACH PART DOES. -l ara, the language is Arabic. Use -l ara+eng if pages mix Arabic and English, which Iraqi official documents often do. --skip-text, OCR pages that are pictures, and leave completely alone any page that already has real text. --deskew, straightens a crooked scan. --clean, removes photocopy speckle. Mac only: it needs unpaper, which is awkward on Windows, so the Windows line leaves it out. --deskew and --rotate-pages still do most of the useful work. --rotate-pages, fixes pages scanned upside down or sideways. --sidecar test.txt, also writes the recognised text into a plain text file. This is what makes Step 22's search possible without installing anything else, and it is what makes the empty-output check above possible at all.
WHY --skip-text AND NOT --force-ocr. The previous version of this recipe used --force-ocr on everything. --force-ocr rasterises every page and throws away any existing text layer, replacing correct embedded text with OCR guesses, so a born-digital PDF that happened to be sitting in your folder came out worse than it went in, silently. --skip-text never touches a page that already has text. This was tested while writing this recipe: running the --skip-text command on a PDF that already had a text layer printed 'skipping all processing on this page' and left it alone.
WHY --redo-ocr HAS NO --deskew. If you try to combine them, OCRmyPDF refuses to start and prints: '--redo-ocr (or --mode redo) is not currently compatible with --deskew, --clean-final, and --remove-background'. That exact message was reproduced on OCRmyPDF 17.10.0 on 30 Aug 2026 while writing this recipe. OCRmyPDF's documentation explains why the trade is acceptable: redo mode 'will replace OCR without rasterizing', so it repairs a bad text layer without degrading the page image.
If --redo-ocr does not fix a case (c) file, --force-ocr is the last resort for that one file, named deliberately, used deliberately, never applied to a whole folder.
Pick your genuinely worst page for this test, not your best one. Ten minutes here tells you whether the whole approach will work on your archive.
هەنگاو 19 / 26هەموو سیستەمەکان
Confirm the text layer actually works, with your own eyes, before doing anything else. Open test-searchable.pdf. Look at the scanned image and pick a clear Arabic word you can see, ideally in the middle of a page. Now press Ctrl+F (Windows) or Command+F (Mac) and type that same word.
دڵنیا ببەرەوە کە کاری کردووە
The word is found and highlighted on the page where you saw it. If nothing is found: on Windows, check that Arabic really installed by re-running `tesseract --list-langs` from Step 14 and confirming `ara` is in the list; on Mac, re-run the same check from Step 17. If ara is present and the search still fails, the page was too degraded for OCR, see Step 18's note.
تێبینی
This is the verification step people skip, and skipping it is how organisations discover, after processing 400 documents, that none of them are searchable.
One important quirk that alarms people unnecessarily: Arabic text copied out of a PDF often looks scrambled or reversed when pasted into a program that does not handle right-to-left text properly. That is a display problem in the other program, not an OCR failure. Searching still works correctly.
هەنگاو 20 / 26هەموو سیستەمەکان
Process the whole folder at once. Your PDFs should already be in the docs folder you made in Step 3. Run the block for your platform.
ئەمە بە دروستی کۆپی بکە
WINDOWS (PowerShell), paste all four lines together: cd C:\ocr-work\docs $files = @(Get-ChildItem -File -Filter *.pdf | Where-Object { $_.Name -notlike "searchable-*" }) "Found $($files.Count) PDFs to process" foreach ($f in $files) { ocrmypdf -l ara --skip-text --deskew --rotate-pages --sidecar "$($f.BaseName).txt" $f.Name "searchable-$($f.Name)" } MAC (Terminal), paste both lines together: cd ~/ocr-work/docs for f in *.pdf; do case "$f" in searchable-*) continue;; esac; ocrmypdf -l ara --skip-text --deskew --clean --rotate-pages --sidecar "${f%.pdf}.txt" "$f" "searchable-$f"; doneدڵنیا ببەرەوە کە کاری کردووە
On Windows the second line prints 'Found N PDFs to process' before anything else happens, check that N matches the number of original PDFs you put in the folder. Then Step 21 does the real verification: do not skip it.
تێبینی
This creates two new files next to each original, searchable-something.pdf and something.txt, and never modifies your originals. If it goes wrong you have lost nothing.
Budget roughly 20-60 seconds per page on an ordinary laptop, so a 500-page gazette is a lunch break and a folder of 100 documents is something you start before going home. The laptop can be offline the entire time; if you want to prove that to yourself, turn off the Wi-Fi first.
If one file fails, the loop carries on to the next. Step 21 finds the ones that failed.
TWO THINGS THE PREVIOUS VERSION GOT WRONG HERE. First, it used --force-ocr, which destroys good embedded text (see Step 18's note). Second, its Windows loop streamed the file list, so files created during the run got picked up and processed again, producing searchable-searchable-something.pdf and doubling the runtime. The @( ) wrapper and the -notlike "searchable-*" filter above exist to stop exactly that.
HONESTY ABOUT TESTING: the Mac line was run start to finish on a real folder while writing this recipe, on 30 Aug 2026, and produced one searchable- PDF and one non-empty .txt per original. The PowerShell block could not be executed on the machine used to write this recipe, because that machine is a Mac. Have one Windows colleague run it on a real Windows laptop before you teach it to a room.
هەنگاو 21 / 26هەموو سیستەمەکان
Verify the batch. This is the step that catches the documents that silently produced nothing. Run the block for your platform, in the same docs folder.
ئەمە بە دروستی کۆپی بکە
WINDOWS (PowerShell): (Get-ChildItem -File -Filter *.pdf | Where-Object {$_.Name -notlike "searchable-*"}).Count (Get-ChildItem -File -Filter searchable-*.pdf).Count Get-ChildItem -File -Filter *.txt | Where-Object { $_.Length -lt 100 } | Select-Object Name, Length MAC (Terminal): ls *.pdf | grep -vc '^searchable-' ls searchable-*.pdf | wc -l find . -name "*.txt" -size -100c -printدڵنیا ببەرەوە کە کاری کردووە
Two conditions, both required.
1. The first two numbers must be EQUAL. If you had 40 originals you must have 40 searchable- files. If the second number is smaller, some files failed, find which original has no searchable- twin, run the Step 18 command on that one file alone, and read the error it prints.
2. The third command must print NOTHING AT ALL. Every file it prints has a text transcript under 100 characters, meaning OCR found essentially nothing on that document. Those documents are NOT searchable, and a search across the folder will silently skip them. Write their names down and either re-scan them at better quality or plan to read them by hand.
On the test run for this recipe both counts were 2 and the third command printed nothing.
تێبینی
This step is the answer to the worst failure mode in the whole recipe: a tool that reports success and produces an empty result. Everything upstream of here can look fine while a quarter of your archive is invisible to search. Ten seconds of counting closes that gap.
Do not tell a colleague 'the archive is searchable' until you have run this and both conditions held.
هەنگاو 22 / 26هەموو سیستەمەکان
Search across everything. There are two ways and you will use both: one document at a time inside the PDF, and all documents at once across the .txt transcripts.
ئەمە بە دروستی کۆپی بکە
ONE DOCUMENT AT A TIME, works everywhere, needs no extra software: Open any searchable-....pdf in any PDF reader and press Ctrl+F (Windows) or Command+F (Mac). ACROSS THE WHOLE FOLDER, search the .txt transcripts: WINDOWS (PowerShell), in C:\ocr-work\docs : Select-String -Encoding utf8 -Path *.txt -Pattern "4712" MAC (Terminal), in ~/ocr-work/docs : grep -in "4712" *.txt Replace 4712 with what you are actually looking for.دڵنیا ببەرەوە کە کاری کردووە
Test it with something you KNOW is there. Pick a decree number, a year or an article number that you can see with your own eyes on one of the scanned pages, and search for that. The command must print the filename and the matching line. On the test run for this recipe, searching for a number printed the filename, the line number and the whole line for each matching document.
If you get no output at all for a term you can see on the page: either the OCR missed it on that document (open that document's .txt file and look), or you searched a longer phrase than the OCR got right. Try a shorter word. Never conclude 'the archive does not contain X' from a single failed search.
تێبینی
ARABIC SEARCH ADVICE THAT WILL SAVE YOU AN HOUR. Search SHORT words. OCR gets long inflected phrases slightly wrong often enough that exact-phrase searching fails, while a short root word like قانون or مادة or قرار still hits. Better still, search for a distinctive NUMBER, an article number, a decree number, a year, because digits survive bad OCR far better than Arabic letters do.
Searching by number has a second benefit on Windows: typing Arabic into a PowerShell window can be awkward, and the console may render right-to-left text badly even when the match is correct. If you need to search an Arabic word, either paste it in from Notepad, or use Ctrl+F inside the searchable PDF, which handles right-to-left properly.
WHY THERE IS NO SEARCH APP TO INSTALL HERE. The previous version of this recipe recommended Recoll at this point. It has been removed. Its own Windows page asks for a payment contribution to download the Windows build (the author will send it free if you email him and ask, but that means emailing a stranger from your work address before you can search your own archive), and on Apple Silicon Macs, every Mac sold since 2020. Its download is unsigned, and its own macOS page states that 'The ARM-based Mac computers will absolutely refuse to run unsigned code', requiring you to code-sign the app yourself from the Terminal. Both verified 30 Aug 2026. For a room with no engineer and a majority on Windows, that is a hard stop dressed up as a one-liner. Searching the .txt files costs nothing and works identically on both platforms.
Before you finish the day, read Step 24. Those .txt files are readable copies of your documents.
هەنگاو 23 / 26هەموو سیستەمەکان
OPTIONAL, offline question-answering. Only do this if your laptop has 16 GB of RAM. Check first; this is a go/no-go, not a suggestion.
ئەمە بە دروستی کۆپی بکە
STEP A, check your RAM before downloading anything: WINDOWS: Settings > System > About -> read 'Installed RAM' MAC: Apple menu > About This Mac -> read 'Memory' If it says 8 GB: STOP. Skip this step entirely. Steps 20-22 already do the useful work. If it says 16 GB or more: continue. STEP B, download AnythingLLM Desktop: https://anythingllm.com/desktop STEP C, BEFORE you add a single document, open the app and go to: sidebar > Privacy -> turn telemetry OFFدڵنیا ببەرەوە کە کاری کردووە
The Privacy panel shows telemetry disabled, and you can see that BEFORE you have added any document. If you added documents first, remove them, turn telemetry off, and add them again. Then, to confirm it really is running locally: disconnect the Wi-Fi and ask it a question about a document you added. If it answers, the model is on your machine. If it errors, it is configured to use a cloud provider and you must fix that before putting anything sensitive into it.
تێبینی
WHY THE RAM CHECK IS A GATE. AnythingLLM's own system requirements page gives the recommended specification as 16 GB of RAM and an 8-core CPU. The previous version of this recipe stated that correctly and then added 'it will run on less but slowly', which buries the decision. For an audience where organisations are already blocked by hardware and are working on old 8 GB laptops, a 16 GB recommendation is closer to a prerequisite than a footnote. On 8 GB it will either swap continuously or fail, after a multi-gigabyte download you may have paid for by the megabyte.
WHY THE TELEMETRY STEP IS MANDATORY. AnythingLLM's download page says: 'Your models, documents, and chat history stay on your machine. Nothing phones home.' Its own project README says something different about metadata: telemetry is a thing you 'opt out' of, 'Set DISABLE_TELEMETRY in your server or docker .env settings to "true" to opt out of telemetry. You can also do this in-app by going to the sidebar > Privacy and disabling telemetry.' Both quotes verified 30 Aug 2026.
What is actually sent is metadata, not content: the installation type, the fact that 'a document is added or removed. No information about the document', the vector database type, the LLM provider and model tag, and the fact that 'a chat is sent'. The README states 'No IP or other identifying information is collected'. So your documents are not leaving. But 'nothing phones home' is not true out of the box, and for a room where 86% fear information leaving the organisation, the correct response is an instruction, not an omission.
WHAT YOU ACTUALLY GET. A small model running on a laptop writes considerably weaker Arabic than a cloud model and will miss nuance in legal language. Treat its answers strictly as a pointer to which document and page to open, which is genuinely useful, and never as a reading of the law. Everything in Step 25 still applies to every sentence it produces.
هەنگاو 24 / 26هەموو سیستەمەکان
Know what is now sitting on your disk in plain readable text, and delete it when the project ends. Do this before you close the laptop for the day.
ئەمە بە دروستی کۆپی بکە
THESE ARE ALL READABLE TEXT COPIES OF YOUR DOCUMENTS, on this laptop: WINDOWS: C:\ocr-work\docs\*.txt the transcripts from Step 20 C:\ocr-work\docs\searchable-*.pdf the text layer inside each PDF MAC: ~/ocr-work/docs/*.txt ~/ocr-work/docs/searchable-*.pdf Plus AnythingLLM's storage folder, if you did Step 23. WHEN A PROJECT ENDS, delete the transcripts: WINDOWS (PowerShell): Remove-Item C:\ocr-work\docs\*.txt MAC (Terminal): rm ~/ocr-work/docs/*.txtدڵنیا ببەرەوە کە کاری کردووە
After deleting, run the Step 22 folder-search command again for a word you know was in those documents. It must now return nothing at all, on Windows, Select-String will report that it cannot find the path because no .txt files remain; on Mac, grep will say 'No such file or directory'. That is the correct result. If it still prints matches, the files are still there.
If the project is completely finished and you only need the original scans, delete the searchable- PDFs too. They carry the same text inside them.
تێبینی
Deleting a file does not wipe it from the disk; it removes the pointer. On an encrypted disk that is acceptable, because the whole disk is unreadable without your password. On an unencrypted disk it is not. That is the whole reason full-disk encryption is Step 2 and not an afterthought.
Think about it this way: before this recipe, a seized laptop gave someone a folder of unreadable scans they would have to sit and read. After it, the same laptop gives them a searchable, indexed archive with a text file per document. The capability you built for yourself is the same capability you built for whoever takes the machine. Encrypt the disk, keep the work folder out of OneDrive and iCloud, and clear the transcripts when you are done.
هەنگاو 25 / 26هەموو سیستەمەکان
MANDATORY VERIFICATION PROTOCOL. Before any AI-derived statement leaves your organisation, in a filing, a report, a press release, or advice to a client, run every single point below. Print this and put it on the wall.
ئەمە بە دروستی کۆپی بکە
1. SOURCE: every claim carries a document name AND a page number. No page number means the claim does not exist yet. 2. OPEN IT: a human opens that exact page and reads the sentence with their own eyes. Not the transcript. The original scan. 3. THE INSTRUMENT IS REAL: confirm the gazette issue number, date, decree or article number against the printed gazette or the official register, never against the AI, and never against a second AI. 4. VERBATIM MATCH: the Arabic sentence the model quoted must appear on that page. Approximately is a failure. 5. COMPLETENESS: if the claim rests on a transcript, confirm the transcript covers the whole document, page markers counted against the PDF's page count, as in Step 7. An answer drawn from half a gazette is wrong in a way no citation check will reveal. 6. ASK TWICE: put the same question in a fresh, empty chat. If the two answers disagree on any citation, discard both. 7. LANGUAGE CHECK: a fluent Arabic or Kurdish reader confirms the meaning, not just the presence, of the passage. 8. SIGN IT: write in the file who verified each citation and on what date. 9. LOG THE MISSES: when this protocol catches a fabricated citation, decree number or quotation, write it down, date, tool, the exact prompt used, and what it invented. Keep that log. 10. THE RULE: nothing unopened gets filed. No exceptions, no deadlines, no "it looked right".دڵنیا ببەرەوە کە کاری کردووە
Take one real AI-produced claim from your own work this week and run all ten points on it, timing yourself. If it takes more than about ten minutes per claim, your prompts are not demanding enough citation detail, go back to the Step 10 template. If any point cannot be completed, the claim does not leave the organisation.
تێبینی
Why this is not optional. In Mata v. Avianca, a New York lawyer filed a brief containing case citations that ChatGPT had invented. On 22 June 2023 Judge P. Kevin Castel sanctioned the lawyers USD 5,000, finding they had acted with subjective bad faith. That is a verified, checkable case, and it is the one this recipe stakes its argument on.
A statement about later trends has been REMOVED from this recipe. The previous version cited an academic database of AI hallucination cases in court, attributed it to named institutions, and gave a count of over 1,300 instances. That database could not be opened during this revision, the site returned HTTP 403, so neither the attribution nor the figure could be confirmed, and an unverifiable citation has no place in the paragraph arguing that unverifiable citations destroy cases. If you want a current figure, find the source and read it yourself before you quote it to a room.
What does not need a citation is the mechanism. Every lawyer in every one of these cases believed the citation looked correct. It always looks correct. That is the entire failure mode. A fabricated Iraqi decree number is indistinguishable from a real one until someone opens the gazette. The model is a finding aid that tells you where to look. It is never the finding.
Point 9 is new. A caught fabrication is the single most useful artefact your organisation will produce for its own policy: it is local, specific, undeniable evidence for the colleague who thinks this caution is excessive. The previous version let those near-misses evaporate.
هەنگاو 26 / 26هەموو سیستەمەکان
Write your organisation's one-page rule and give it to everyone, including interns and volunteers. Four sentences is enough to start.
ئەمە بە دروستی کۆپی بکە
1. Documents naming a living person never go into any online AI tool, free or paid, ever, including a public document that names someone whose situation is still live. 2. No AI-produced citation, quote or legal conclusion leaves this organisation until a named human has opened the source page and signed for it. 3. Public official documents may be used with cloud tools; when in doubt, a document is not public. 4. Every time we catch the tool inventing something, we write it down, date, tool, prompt, what it invented, and we keep the list.دڵنیا ببەرەوە کە کاری کردووە
Print it, tape it above the scanner, and ask one colleague who was not in the room to read it and tell you which folder a specific real document from your pile belongs in. If they hesitate or get it wrong, the wording is not clear enough yet, fix the wording, not the colleague.
تێبینی
71% of the organisations surveyed have no written AI policy, and 43% have already put sensitive material into an AI tool, usually because nobody had ever said out loud where the line was. These four sentences, printed and taped above the scanner, close most of that gap in an afternoon. Add to it later; do not wait for a perfect policy.
This recipe contains that rule in full, inline, on purpose. Do not send anyone to a policy-builder page or an online template to get it, write these four sentences on paper today.
چۆن بزانیت هەموو کارەکە سەرکەوتووە
The recipe has eight checkpoints, and passing all eight is what "it worked" means. Anything less and you have a pile of files you cannot trust.
1. SORTING (Step 1). Every document is in PUBLIC-documents or SENSITIVE-documents, and the two folders add up to the number you started with.
2. ENCRYPTION (Step 2). Windows shows Device encryption On, or BitLocker on for C:. Mac shows FileVault turned on. Do this before any sensitive file touches the machine.
3. FOLDER LOCATION (Step 3). The address bar reads exactly C:\ocr-work\docs on Windows, or the folder sits directly in your home folder on Mac. No OneDrive in the path, no cloud icons on the folder, not the Desktop, not Documents.
4. TOOLS FOUND (Step 14 on Windows, Step 17 on Mac). All four commands print a version or a list. Critically, `tesseract --list-langs` includes a line that is exactly `ara`. On the test machine this printed 163 languages including ara, with no ckb and no kur.
5. ONE PAGE WORKS (Steps 18-19). test-searchable.pdf exists, test.txt exists AND contains recognisable Arabic words, and pressing Ctrl+F for a word you can see on the scan finds it. All three, not one.
6. THE BATCH IS COMPLETE (Step 21). The count of originals equals the count of searchable- files, and the empty-transcript command prints nothing at all. This is the check that catches silent-empty-success, which is the worst failure mode in this whole workflow.
7. SEARCH RETURNS SOMETHING KNOWN (Step 22). Search for a decree number or year you can read with your own eyes on a page, and the command prints the filename and the line.
8. VERIFICATION SURVIVES CONTACT (Step 25). Take one AI-produced claim and run all ten points. If any point cannot be completed, the claim does not leave the organisation.
If you used the online lane, add: the page-marker count in the transcript equals the PDF's page count (Step 7), and three random passages were found on the actual scan by eye (Step 8).
چی هەڵە دەبێت، و ئەو کاتە چی بکەیت
- SILENT EMPTY SUCCESS, the worst one. OCRmyPDF prints a normal 'Output file is a PDF/A' message and returns a searchable- PDF whose text layer is empty, because the page was too degraded. Nothing warns you. Catch it with the --sidecar text file: Step 18 checks one file, Step 21 checks the whole folder with a command that must print nothing.
- SILENT TRUNCATION IN THE ONLINE LANE, a model transcribes a 40-page gazette, hits its reply-length limit around page 12, and stops. The transcript ends with a well-formed paragraph and no marker of where it stopped. Spot-checking 'the end' inspects the end of the transcript, not the end of the document. Catch it with the page-marker count in Step 7.
- SENSITIVE DOCUMENTS UPLOADED TO MICROSOFT OR APPLE BY ACCIDENT, copying a witness statement onto the Windows Desktop uploads it to OneDrive on most Windows 10 and 11 machines signed into a Microsoft account, with no prompt and no error. Same for iCloud Desktop & Documents sync on macOS. Caused by the previous version of this recipe. Prevented by Step 3.
- BORN-DIGITAL TEXT DESTROYED BY OCR, running --force-ocr over a mixed folder rasterises every page and replaces correct embedded text with OCR guesses. The file gets worse and looks the same. Prevented by using --skip-text (Steps 18 and 20).
- THE WINDOWS LOOP EATING ITS OWN OUTPUT, a streaming PowerShell pipeline re-enumerates the files it just created, producing searchable-searchable-something.pdf and doubling the runtime. Prevented by the @( ) wrapper and the -notlike "searchable-*" filter in Step 20.
- A DEAD PACKAGE STOPPING THE WINDOWS INSTALL DEAD, `winget install -e --id ArtifexSoftware.GhostScript` returns 'No package found matching input criteria' because it was removed from the winget catalogue in October 2023. Ghostscript is a hard dependency, so every later OCR command fails too. Fixed by the manual download in Step 13.
- PATH NOT SET ON WINDOWS, Tesseract or Ghostscript is installed but Windows cannot find it, producing 'The term ... is not recognized' at a point where the reader has no idea what it means. This is the most common Windows problem with this toolchain. Caught by the four-command check in Step 14.
- PIP RUN IN AN ADMINISTRATOR WINDOW, an elevated PowerShell may not see a per-user winget Python install, and bare `pip` is often not on PATH. Use `py -m pip install ocrmypdf` in a normal window (Step 13).
- --redo-ocr COMBINED WITH --deskew, OCRmyPDF refuses to start and prints '--redo-ocr (or --mode redo) is not currently compatible with --deskew, --clean-final, and --remove-background'. Reproduced live on version 17.10.0 while writing this recipe. The Step 18 redo command deliberately omits those flags.
- ARABIC MISSING FROM TESSERACT ON WINDOWS, clicking Next past the 'Additional language data' screen installs English only, and every Arabic command fails later with a confusing message about a missing language file. Caught by `tesseract --list-langs` in Step 14.
- SEARCHING A PHRASE THAT OCR GOT SLIGHTLY WRONG, an exact-phrase search over Arabic OCR returns nothing while the page is sitting right there. The reader concludes the archive does not contain the thing. Search short root words or numbers instead (Step 22).
- SORANI RUN THROUGH THE ARABIC MODEL, ڕ ڵ ۆ ێ پ چ ژ گ come out as the nearest Arabic letter or as nothing. The output is searchable for some words and silently wrong everywhere else. There is no Sorani Tesseract model to switch to. Treat any Sorani OCR as a locator only.
- A FABRICATED CITATION THAT READS PERFECTLY, a model produces a plausible Iraqi decree number and a fluent Arabic sentence that is not on the page. Indistinguishable from a real one until a human opens the gazette. Step 25 is the only defence.
- AN ENCRYPTED-NOTHING LAPTOP, the recipe succeeds, and now a seized machine contains a fully searchable full-text archive of witness statements instead of a shoebox of unreadable scans. Prevented only by Step 2 and the clean-up in Step 24.
- GOOGLE ACCOUNT PHONE VERIFICATION, creating a new Google account from Iraq often demands SMS verification, and a new or dormant account can be locked pending re-verification. Discovered at Step 5, after the reader has already invested time. Named in advance in the Iraq access note.
چەندە بە زمانی تۆ کار دەکات
عەرەبیی ستاندارد · عەرەبیی عێراقی · کوردیی سۆرانی
عەرەبیی ستاندارد
This is the best-supported case, and it still has a sharp quality cliff.
Cleanly printed Modern Standard Arabic at 300 dpi, a recent PDF gazette, a laser-printed judgment, a ministry circular, comes out of free offline Tesseract well enough to search reliably. You will get occasional errors on hamza forms (أ إ آ ا), taa marbuta versus haa (ة / ه), and yaa versus alef maqsura (ي / ى), which matters because a search for the exact word may miss a page that Tesseract spelled slightly differently.
Old, photocopied, skewed, stamped, two-column or handwriting-annotated scans, which is exactly what most Iraqi official gazette archives look like, are where offline Tesseract degrades badly: merged words, dropped diacritics, whole lines lost under a stamp. On this material a cloud vision model is markedly better, which is precisely why the recipe splits into a public-document lane and a sensitive-document lane instead of pretending one tool serves both.
Practical consequence for searching: search short root words, not long inflected phrases. 'قانون' will find pages that 'قانون الأحوال الشخصية' misses. Better still, search a number, an article number, a decree number, a year. Digits survive bad OCR better than Arabic letters do, and they also spare you having to type Arabic into a PowerShell window.
عەرەبیی عێراقی
Mostly not applicable, and that is good news here. Official gazettes, statutes, judgments and ministry correspondence are written in Modern Standard Arabic, not Iraqi dialect, so the OCR step is unaffected by dialect entirely.
Dialect becomes relevant only in two places. First, if your source pile includes handwritten complaints, witness statements or transcribed interviews in Iraqi Arabic, OCR on Arabic handwriting is unreliable in every free tool considered here and should be treated as not working. Do not budget an afternoon on it. Second, when you ask a model to summarise in Iraqi Arabic: cloud models produce something readable, small local models produce noticeably stilted output. For legal work, ask for summaries in MSA or English and let a human colleague render them into dialect.
کوردیی سۆرانی
Be direct with your colleagues about this: Kurdish Sorani OCR is not solved, and no free tool in this recipe handles it properly.
There is no maintained Sorani (ckb) Tesseract model. You can prove this on your own laptop in one command, and Step 17 asks you to. On the machine used to write this recipe, `tesseract --list-langs` printed 163 languages from a complete Homebrew tesseract-lang installation. `ara` was present. `kmr`, Kurmanji, Northern Kurdish, written in Latin script, was present. `ckb` was not there. `kur` was not there. The same is true of the tessdata_best repository's file listing: ara.traineddata and kmr.traineddata exist, ckb.traineddata and kur.traineddata do not. Google Drive's OCR help page lists 'Arabic (Modern Standard)' among its supported languages and lists no Kurdish at all.
The workaround people try is running Sorani through the Arabic model, because Sorani uses Arabic script. It half-works and then fails on exactly the letters that make it Kurdish: ڕ ڵ ۆ ێ پ چ ژ گ are not in the Arabic model's alphabet, so they come out as the nearest Arabic letter or as nothing. A Sorani page OCR'd with -l ara is searchable for words that happen to use only shared letters, and silently wrong everywhere else. That silent wrongness is worse than a visible failure.
On cloud vision models and Sorani, the previous version of this recipe asserted that they read Sorani better than Tesseract does. That was not tested and no source was opened for it, so it is withdrawn. What can be said honestly is narrower: Tesseract has no Sorani model at all, so anything that produces plausible Sorani text is doing better than a tool that cannot attempt it, and none of it is verified accurate enough to quote from. Test it on your own pages before relying on it either way.
Honest recommendation, unchanged: for Sorani material, use any OCR output as a rough locator only, never as a text you quote from, and have a Kurdish-reading colleague type out any passage you intend to rely on. Budget human time for this rather than assuming a tool will absorb it.
ئەمە پشت بە چی دەبەستێت
KURDISH CLAIM, corrected since the last version, and the correction matters. The previous version stated that the Tesseract data-files page 'lists the newer alternative as Kurmanji, not Sorani'. That is not what the page says. The verbatim note on tesseract-ocr.github.io/tessdoc/Data-Files.html, read 30 Aug 2026, is: 'The kur data file was not updated from 3.04. For Fraktur, use the newer data files from the tessdata_fast or tessdata_best repositories.' It says FRAKTUR, an apparent copy-paste artefact in Tesseract's own documentation, and it does not mention Kurmanji anywhere. Neither 'ckb' nor 'kmr' appears on that page at all. The recipe had invented an attribution to a named, checkable page; the misquote is removed and the correct quote is above.
The Kurmanji-not-Sorani point is real, but it is sourced from somewhere else: the tessdata_best repository's own file listing, checked 30 Aug 2026, which contains ara.traineddata and kmr.traineddata but no ckb.traineddata and no kur.traineddata. It is also directly observable on your own machine, which is better than any citation: `tesseract --list-langs` after a full language install printed 163 codes including ara and kmr, and neither ckb nor kur, on 30 Aug 2026.
ARABIC-SUPPORTED CLAIM: the same Tesseract page lists 'ara' (Arabic) as an available traineddata file. Google's Drive OCR help page (support.google.com/drive/answer/176692, read 30 Aug 2026) lists 'Arabic (Modern Standard)' among supported OCR languages, note that it names MSA specifically, and lists no Kurdish variety.
ARABIC QUALITY-CLIFF CLAIM: partly measured, partly experience. The Surya OCR project's published multilingual benchmark (github.com/VikParuchuri/surya, read 30 Aug 2026) reports 'Arabic 72.7%' against an 'Overall pass rate: 87.2% across 91 languages'. I.e. Arabic sits well below the average even in a current, well-funded OCR system. Tesseract on degraded gazette scans performs worse than that figure, not better. That last sentence is an inference, not a measurement.
UNTESTED, AND YOU SHOULD KNOW IT: nobody ran a controlled accuracy test on actual Iraqi Official Gazette (الوقائع العراقية) scans for this recipe. The claim that old gazette scans OCR badly is an inference from how these tools behave on comparable degraded Arabic print, not a measurement on your specific archive. Run Step 18 on five of your own worst pages before you commit an afternoon to the whole pile. That ten-minute test is worth more than any benchmark quoted here.
ئەو نەرمامێرانەی لەم ڕەچەتەیەدا ناویان هاتووە
- Tesseract OCR (open source, Apache 2.0, the ara language model; note there is no ckb/Sorani model)
- OCRmyPDF 17.11.0 (open source, MPL-2.0, adds a searchable text layer to a scanned PDF)
- Ghostscript 10.07.1 (free under the AGPL, a hard dependency of OCRmyPDF; must be installed by hand on Windows)
- unpaper 7.0.0 (photocopy speckle removal; macOS/Linux only in this recipe)
- Homebrew (macOS package manager)
- Python 3.12 via winget (Windows only, to install OCRmyPDF)
- Windows PowerShell / macOS Terminal
- Google AI Studio (Gemini free tier, public documents only)
- Gemini Notebook, formerly NotebookLM, at notebook.google.com (free tier, public documents only)
- Google Drive OCR (free, 2 MB file limit, mentioned only to explain why it is unsuitable for gazette scans)
- AnythingLLM Desktop (MIT, optional, 16 GB RAM gate, telemetry must be turned off)
- BitLocker / Windows Device encryption, macOS FileVault (full-disk encryption, a required step)
- VeraCrypt 1.26.29 at veracrypt.io (named as the fallback where BitLocker is unavailable)
- REMOVED FROM THIS RECIPE: Recoll (Windows build asks for a payment contribution; unsigned on Apple Silicon). REMOVED: the winget package ArtifexSoftware.GhostScript (deleted from the catalogue in October 2023).
سەرچاوەکان
- https://ocrmypdf.readthedocs.io/en/latest/installation.html · Windows install: winget Python.Python.3.12 and UB-Mannheim.TesseractOCR only; 'You will need to install Ghostscript manually, since it does not support automated installs anymore.' Read 30 Aug 2026.
- https://ocrmypdf.readthedocs.io/en/latest/cookbook.html · --redo-ocr, --skip-text, --force-ocr and --sidecar behaviour; 'This method will replace OCR without rasterizing'. Read 30 Aug 2026.
- https://pypi.org/pypi/ocrmypdf/json · ocrmypdf 17.11.0, published 2026-08-28, requires Python >=3.11. Read 30 Aug 2026.
- https://github.com/ocrmypdf/OCRmyPDF (source of v17.11.0, src/ocrmypdf/cli.py), confirms --skip-text, --redo-ocr and --force-ocr are all still accepted flags. Read 30 Aug 2026.
- Live test on OCRmyPDF 17.10.0, 30 Aug 2026, the full Mac command with -l ara+eng --skip-text --deskew --clean --rotate-pages --sidecar ran to completion and produced a non-empty sidecar; re-running --skip-text on a PDF that already had text printed 'skipping all processing on this page'; combining --redo-ocr with --deskew produced the error '--redo-ocr (or --mode redo) is not currently compatible with --deskew, --clean-final, and --remove-background'.
- https://api.github.com/repos/microsoft/winget-pkgs/contents/manifests/a/ArtifexSoftware · the directory contains only 'mutool'; ArtifexSoftware.GhostScript is gone. Manifests for UB-Mannheim/TesseractOCR and Python/Python/3/12 confirmed present. Read 30 Aug 2026.
- https://ghostscript.com/releases/gsdnld.html and https://api.github.com/repos/ArtifexSoftware/ghostpdl-downloads/releases/latest, Ghostscript/GhostPDL 10.07.1, published 19 May 2026, Windows 64-bit installer gs10071w64.exe (64.9 MB), free AGPL release available. Read 30 Aug 2026.
- https://github.com/UB-Mannheim/tesseract/wiki · current Windows installer tesseract-ocr-w64-setup-5.5.3.20260724.exe. Read 30 Aug 2026.
- https://tesseract-ocr.github.io/tessdoc/Data-Files.html · verbatim: 'The kur data file was not updated from 3.04. For Fraktur, use the newer data files from the tessdata_fast or tessdata_best repositories.' Lists 'ara'. Neither 'ckb' nor 'kmr' appears on the page. Read 30 Aug 2026.
- https://github.com/tesseract-ocr/tessdata_best · ara.traineddata (12.6 MB) and kmr.traineddata present; ckb.traineddata and kur.traineddata absent. Read 30 Aug 2026.
- Live observation, 30 Aug 2026, `tesseract --list-langs` on a complete Homebrew tesseract-lang installation printed 163 languages including ara and kmr, and neither ckb nor kur.
- https://formulae.brew.sh/api/formula/{ocrmypdf,tesseract-lang,unpaper,ghostscript}.json · all four formulae exist; ocrmypdf 17.11.0, ghostscript 10.07.1, tesseract-lang 4.1.0, unpaper 7.0.0. recoll returns 404 as both formula and cask. Read 30 Aug 2026.
- https://ai.google.dev/gemini-api/docs/pricing · verbatim: 'Google AI Studio usage is free of charge in all available regions'; free tier 'Content used to improve our products' versus paid tier 'Content not used to improve our products'; Flash and Flash-Lite families listed as free of charge. Read 30 Aug 2026.
- https://ai.google.dev/gemini-api/docs/available-regions · Iraq listed, immediately before Ireland. Read 30 Aug 2026.
- https://www.anthropic.com/supported-countries · Iraq listed for both Claude.ai and the API, between Indonesia and Ireland. Read 30 Aug 2026.
- https://support.google.com/notebooklm/answer/16269187 · product is called Gemini Notebook; free plan limits: 100 notebooks, up to 50 sources each, 500,000 words each, 50 chat queries per day, 3 audio generations per day. Read 30 Aug 2026.
- Live check, 30 Aug 2026, https://notebooklm.google.com/ returns HTTP 301 to https://notebook.google.com/.
- https://support.google.com/drive/answer/176692 · 'The file should be 2 MB or smaller'; 'Text should be at least 10 pixels high'; 'Arabic (Modern Standard)' listed among supported OCR languages; no Kurdish variety listed. Read 30 Aug 2026.
- https://github.com/VikParuchuri/surya · 'Arabic 72.7%' against 'Overall pass rate: 87.2% across 91 languages'; model weights under a modified AI Pubs Open Rail-M licence 'free for research, personal use, and startups under $5M funding/revenue'. Read 30 Aug 2026.
- https://github.com/Mintplex-Labs/anything-llm (README, Telemetry & Privacy section), 'Set DISABLE_TELEMETRY ... to "true" to opt out of telemetry. You can also do this in-app by going to the sidebar > Privacy and disabling telemetry.' Lists installation type, document added/removed events with 'No information about the document', vector DB type, LLM provider and model tag, and 'When a chat is sent'; states 'No IP or other identifying information is collected'. Read 30 Aug 2026.
- https://anythingllm.com/desktop · 'Your models, documents, and chat history stay on your machine. Nothing phones home.' Free; macOS, Windows and Linux builds. Read 30 Aug 2026.
- https://docs.anythingllm.com/installation-desktop/system-requirements · recommended 16 GB RAM, 8-core CPU. Read 30 Aug 2026.
- https://www.recoll.org/pages/recoll-macos.html · 'I don't sign the application...'; 'The ARM-based Mac computers will absolutely refuse to run unsigned code (unlike the Intel Macs which protest but yield to user insistence).' Read 30 Aug 2026.
- https://www.recoll.org/pages/recoll-windows.html · 'I recommend contributing a small fee for downloading and using the Windows binary version'; 'If the payment is a problem for you for any reason, please send me an email, no need for justification'. Version 1.44.1, 2026-08-24. Read 30 Aug 2026.
- https://en.wikipedia.org/wiki/Mata_v._Avianca,_Inc. · Judge P. Kevin Castel imposed a USD 5,000 sanction on 22 June 2023 for a brief containing ChatGPT-fabricated citations. Read 30 Aug 2026.
- https://veracrypt.io/en/Downloads.html · VeraCrypt 1.26.29, released 9 June 2026; free open source disk encryption; Windows and macOS installers listed. Read 30 Aug 2026. (Note: veracrypt.fr now redirects to veracrypt.io.)
- NOT AVAILABLE, the academic AI-hallucination-in-court case database returned HTTP 403 to every attempt on 30 Aug 2026. The claim it supported has been deleted from this recipe rather than softened.
بەرهەمێکی ڕەسەنی CDR، بۆ کۆنفرانسی پۆینتی عێراق ٧ نووسراوە. لە ئابی ٢٠٢٦دا بەرامبەر پەڕەکانی دابینکەران و کۆگای پاکێجەکان پشکنراوە؛ سەرچاوەکانیش لەم پەڕەیەدا هاتوون.
ئەوەی نەمانتوانی پشتڕاستی بکەینەوە
WHAT WAS REMOVED, AND WHY.
Recoll, the desktop search tool, is gone. It worked, but not for this room. Its Windows build asks for a payment contribution before download, the author will send it free to anyone who emails and asks, with no justification required, which is generous, but it means a human rights worker emailing a stranger from a work address before they can search their own archive. And on Apple Silicon Macs, which is every Mac sold since 2020, the download is unsigned and its own macOS page says ARM Macs "will absolutely refuse to run unsigned code" without you code-signing it yourself from the Terminal. A room with no engineer, a majority on Windows and 64% blocked by cost cannot absorb either. Step 22 replaces it with a method that needs no download on any platform: OCRmyPDF's --sidecar option writes a .txt transcript beside each PDF during Step 20, and one built-in command searches all of them. The trade-off is honest: no graphical search box, no ranked relevance, no automatic re-indexing when files change. You get a filename and a matching line, which is what a finding aid needs to produce.
The winget Ghostscript package is gone, because it does not exist. Step 13 downloads it by hand.
The USD 20-22 monthly price figure is gone. It was published in the previous version alongside an admission that the vendors' pricing pages could not actually be read. A price nobody opened is exactly the kind of citation this recipe tells readers to reject.
The claim about an academic database of AI hallucination cases, the institutional attribution and the "over 1,300 instances" figure, is gone. The site returned HTTP 403 on every attempt during this revision. Only the Mata v. Avianca facts, which were verified, remain.
The claim that cloud vision models read Sorani better than Tesseract is gone. It was asserted with the same confidence as claims that had actually been tested, and it had not been tested. What replaces it is narrower and true: Tesseract has no Sorani model at all.
WHAT IS STILL NOT KNOWN.
Nobody has run this on actual Iraqi Official Gazette scans. The claim that old gazette scans OCR badly is an inference from how these tools behave on comparable degraded Arabic print, not a measurement on your archive. Step 18 exists so you spend ten minutes finding out before you spend an afternoon.
The Windows path has not been executed end to end. Every Windows command here was checked against the current package registry, the vendor's own download page, or the tool's official documentation, and the dead ones were replaced. But the machine this recipe was written on is a Mac, so the macOS pipeline was run start to finish on a real folder and the PowerShell block was not. Before this is taught to a room, one person should run the Windows path on a real Windows laptop with OneDrive switched on, including the Step 3 folder check, which is the step most likely to reveal something unexpected.
Kurdish Sorani does not work, and this recipe does not fix it. There is no maintained Sorani Tesseract model; you can confirm that on your own screen in one command in Step 17. Running Sorani through the Arabic model produces text that is searchable for some words and silently wrong everywhere else, which is worse than a visible failure. Budget human transcription time. Do not budget software.
Arabic handwriting does not work either. Handwritten complaints, witness statements and annotated margins should be treated as not OCR-able by any free tool here.
The online lane's quality advantage on damaged scans is real and it is only available for documents that are already public. That asymmetry is not a flaw in the recipe; it is the actual state of the field, and pretending otherwise would push sensitive material into the cloud.
Encryption protects a powered-off, locked laptop. It does not protect a laptop that is on and logged in when it is taken, and it does not protect against someone who can compel you to unlock it. This recipe reduces one risk substantially and does not eliminate the category.
Finally: this recipe makes it faster to find the page. It does not make the page mean what you hope it means. Step 25 is the part that cannot be automated, and it is the part that keeps your organisation out of the news.