لم يراجع هذه الترجمة إنسان بعد
أُنتجت بمساعدة الترجمة الآلية. اقرأ النسخة الإنجليزية قبل أن تتصرف بناءً على أي شيء فيها.
ترجم الكلام المنطوق إلى العربية أو الكردية على حاسوبك، دون إرسال أي شيء إلى الإنترنت
أكثر الوصفات تقنيةً هنا: سطر أوامر، ونحو 4 غيغابايت تُنزَّل مرة واحدة، ثم ترجمة حية والشبكة مفصولة. وهناك حاسوب شائع لا يستطيع تشغيلها إطلاقاً.
- الوقت اللازم
- 100 دقيقة · 23 خطوة
- ما الذي يتطلبه
- تقني: يستخدم سطر الأوامر
- التكلفة
- يوجد مسار مجاني
- آخر تحقق
- 2026-08-31
- صحفي أو محرر
- التواصل والحملات
- التوثيق والرصد
- المنح والشؤون الإدارية
- الترجمة من اللغات المحلية وإليها
- تفريغ الشهادات والمقابلات
- حماية المعلومات الحساسة
هذه تعليمات، وليست دراسات حالة
الوصفة تخبرك بما ينبغي أن تفعله. وهي ليست تقريراً عن شيء حدث في مكان آخر. لا شيء هنا يزعم أن منظمة ما قامت بهذا؛ بل هي مكتوبة لكي تنفّذها أنت الآن، وكل خطوة تشرح لك كيف تتأكد من أنها نجحت.
كلام حي، مترجَم إلى العربية أو الكردية، على حاسوب محمول والإنترنت مفصول. هذه أكثر وصفات هذا القسم تقنيةً: تستخدم سطر الأوامر، وهي الوحيدة التي لا تستطيع بعض الأجهزة تشغيلها إطلاقاً. اقرأ قائمة الأجهزة قبل أن تثبّت أي شيء.
ويندوز 10 أو 11 على معالج إنتل أو AMD بنواة 64 بت هو المسار الأكثر اختباراً، وكل خطوة تعطي أمر ويندوز أولاً. أجهزة ماك بمعالج إنتل لا تستطيع تشغيل هذا، ولا يوجد إعداد يصلحه. وأجهزة ماك بمعالج آبل تحتاج إلى macOS 14 أو أحدث. أما كروم بوك واللوحيات والهواتف فغير مدعومة ولا حلّ بديل لها.
خط الأمان
بعد الإعداد لمرة واحدة، لا شيء يغادر. لا صوت، ولا نص، ولا ترجمة، ولا بيانات وصفية، ولا تتبّع. وأثناء الإعداد وحده، يجلب حاسوبك حزم البرامج وملفات النماذج؛ وتلك المواقع ترى عنوان IP الخاص بك وأي ملفات نزّلتها، لا محتواك، لأنك لا تملك محتوى بعد.
لا تستخدم مخرجات الآلة بوصفها السجل نفسه، لا لشهادة، ولا لمقابلة مع ناجٍ، ولا لمذكرة قضائية، ولا لوثيقة طبية أو قانونية، ولا لطلب لجوء. وفي العامية العراقية، توقّع نسبة خطأ لا تقل سوءاً عن 48% المقيسة على العامية الأردنية، وربما أسوأ بكثير. ولا تحذف تسجيلك الأصلي ولا تكتب فوقه أبداً.
عن الكردية: جوابان مختلفان
ترجمة النص إلى السورانية تعمل: فنموذج الترجمة المستخدم هنا يغطي السورانية المكتوبة بالحرف العربي. أما التعرف على الكردية المنطوقة فلا يعمل إطلاقاً، لأن ويسبرWhisper (running on your own computer)آمن للمواد الحساسةمفتوح المصدرMITله نسخة مجانية: Completely free. You need about 2GB of disk for a usable model.عرض الأداة لا يدعم الكردية بأي صورة. فهذا المسار يستطيع تحويل الكلام العربي إلى نص سوراني. ولا يستطيع أن يسمع السورانية.
ونقطة ترخيص لمديرك: نموذج الترجمة مرخّص للاستخدام غير التجاري. وهذا يتوافق مع التوثيق والصحافة. ولا يتوافق مع بيع خدمة ترجمة، ولا مع تقديم مخرجات آلية بوصفها سجلاً رسمياً.
التفاصيل الكاملة أدناه، الخطوات والأوامر والمصادر، منشورة بالإنجليزية، لأن الأوامر وأسماء الأزرار داخل البرامج نفسها بالإنجليزية، ولأن رقم إصدار يختلف بين ثلاث ترجمات هو رقم لا يمكن الوثوق به. المقدمة والملخص الأمني في أعلى الصفحة مترجمان.
ماذا يحدث لمعلوماتك
لا تستخدم هذا أبداً من أجل
Do not use the machine output as the record itself - of a testimony, an interview with a survivor, a court submission, a medical or legal document, or an asylum claim. On Iraqi colloquial Arabic expect the error rate to be at least as bad as the 48% measured on Jordanian Arabic and quite possibly much worse, and Whisper is documented to sometimes invent whole sentences that nobody said. Never delete or overwrite your original audio: this produces a draft to work from, not a transcript of record. Do not paste the output into a cloud service to clean it up - that undoes the entire recipe. Do not use it on a laptop without full-disk encryption for anything that would endanger someone if the machine were seized.
ما الذي يغادر جهازك
أي قانون ينطبق عليها · الخيار المحلي بالكامل، دون إنترنت
After the one-time setup, nothing. No audio, no transcript, no translation, no metadata, no telemetry. Speech is captured by your microphone, converted to text by a Whisper model file on your disk, and translated by an NLLB model file on your disk. There is no server in the loop and no account tied to it. During setup only, your computer contacts pypi.org (to fetch the Python packages), huggingface.co (to fetch the model files) and, on Linux only, download.pytorch.org. Those hosts see your IP address and which files you downloaded - they do not see any of your content, because you have no content yet. BE PRECISE ABOUT WHY THE RUN-TIME SILENCE IS REAL, because an earlier draft of this recipe got this wrong in the direction that matters. The script sets HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1, but those two settings govern Hugging Face downloads ONLY. They do not govern torch.hub, which is a separate download channel. The earlier draft installed RealtimeSTT with an option that did not include the voice-detection model, and RealtimeSTT then quietly fetched that model from github.com through torch.hub every single time the program started - while the recipe told readers it could not reach the network. That is fixed here by installing RealtimeSTT[recommended] in step 14, which includes the silero-vad package. I verified that package ships its own model file inside the wheel (silero_vad_op18_ifless.onnx, in a 9.1 MB download), so nothing is fetched at run time. Step 15 makes you check this on your own machine rather than trusting this paragraph, and step 23 makes you prove the whole thing with the wifi physically off.
أي قانون ينطبق عليها
None for your data - it never leaves your machine, so no jurisdiction applies to it. For the one-time download: PyPI is operated by the Python Software Foundation (United States), Hugging Face Inc. is a United States company, and download.pytorch.org is operated by the PyTorch Foundation (United States). All three see only download requests, never your recordings.
الخيار المحلي بالكامل، دون إنترنت
This recipe IS the fully local option. That is the entire point of it. It is the strongest answer in this collection to the 86% of organisations that fear sensitive information leaving them: you can physically disconnect the laptop from the internet, or work in a room with no connectivity, and it keeps working. Verify this yourself with step 23 rather than taking anyone's word for it. BUT LOCAL IS NOT THE SAME AS PRIVATE, and a recipe whose whole premise is 'the sensitive material stays on this laptop' owes you the next sentence: and that laptop gets seized, stolen, borrowed, or handed to a repair shop. Three things follow. (1) TURN ON FULL-DISK ENCRYPTION BEFORE you put anything sensitive through this - BitLocker on Windows, FileVault on macOS, LUKS on Linux. This is step 2 and it is not optional for documentation work. (2) SWAP. This recipe deliberately targets an 8 GB laptop that may page memory out to disk. When that happens, fragments of the audio buffer and of the decoded transcript can be written into a swap file and can survive there after the program closes. macOS encrypts swap by default; many Ubuntu installs do not; on Windows the page file is covered once BitLocker is on. Full-disk encryption is what makes this a non-issue. (3) WHERE THE OUTPUT GOES NEXT is where most people actually leak. The moment you copy the translation out of this terminal and paste it into Google Docs, Gmail, WhatsApp, Telegram or any online translator to tidy it up, every guarantee on this page is void - you have just uploaded the sensitive text yourself. Your terminal scrollback is also a retained record, and it may be swept into Time Machine, OneDrive or another backup that syncs to a cloud. Decide where the text is going to live before you start speaking.
التكلفة
الدفع من العراق
Nothing to pay for, so Iraqi card rejection and geo-blocking of payment do not apply at all. This is one of the few recipes with zero payment surface: no sign-up, no phone verification, no sanctions-screened account creation anywhere in the flow. The only access cost is bandwidth and disk. HOW MUCH YOU ACTUALLY DOWNLOAD, corrected from an earlier draft that understated it: about 4 GB on Windows and on Apple Silicon macOS. On Ubuntu it is also about 4 GB IF you do step 13, and about 8-10 GB if you skip step 13, because pip otherwise pulls in NVIDIA CUDA libraries for a graphics card your laptop does not have. Whether huggingface.co and pypi.org download reliably on every Iraqi ISP is unverified in this research pass, and I found no evidence of either host blocking Iraq. IF THE DOWNLOAD WILL NOT COMPLETE ON YOUR CONNECTION, use this and not a mirror: have one colleague on a good connection complete steps 9 to 18 on the SAME operating system, then copy just two things to a USB stick - the folder 'models' and the file 'translate_live.py'. Every recipient still has to run steps 9 to 15 themselves (creating their own .venv and running their own pip install), and then drops the copied 'models' folder and 'translate_live.py' into their own live-translate folder, skipping steps 16 and 17. DO NOT copy the .venv folder: it contains absolute paths and compiled files built for one specific machine and user account, and it will not work on another laptop. An earlier draft told you to copy the whole live-translate folder; that does not work, and the office that most needs the workaround would have discovered it only after losing an afternoon. A previous draft also suggested setting HF_ENDPOINT to the third-party mirror hf-mirror.com. That advice has been REMOVED. It is an unaudited proxy with no published provenance, it would see every model request from your office's IP address, and the NLLB weights file it would serve you is a Python pickle that the conversion tool then loads. For documenters this is not a trade worth making: prefer a colleague's USB copy from a machine you trust.
المسار المجاني
إذا دفعت
The entire recipe is free and stays free. No account, no sign-up, no API key, no credit card, no usage limit, no per-minute charge. You download roughly 4 GB of software and model files once, then it runs unlimited times with the internet switched off. Software licences: RealtimeSTT (MIT), faster-whisper (MIT), CTranslate2 (MIT), silero-vad (MIT), OpenAI Whisper weights (MIT). One licence catch: the NLLB-200 translation model is CC-BY-NC, meaning non-commercial use only. Documentation, journalism, internal reports and advocacy are non-commercial and fine. Selling translation as a paid service using this model is not permitted by its licence.
إذا دفعت
$0. There is no paid tier and nothing to upgrade to.
قبل أن تبدأ
- A supported laptop. Windows 10/11, or an Apple Silicon Mac (M1 or later) running macOS 14 or later, or Ubuntu 22.04+. INTEL MACs CANNOT RUN THIS AT ALL - see the platform notes before you start, because there is no workaround and you will otherwise hit a wall at step 14.
- A laptop with at least 8 GB of RAM. 4 GB will not work. Two of the fourteen organisations surveyed are blocked by hardware and this recipe is one they will be blocked on.
- Free disk space during setup: at least 12 GB on Windows and macOS. On Ubuntu, 12 GB if you do step 13, or 18 GB if you skip it. You get roughly 6 GB back at step 18. Expect to end up using about 4 GB permanently.
- A working microphone. A cheap USB headset microphone will improve accuracy more than any software setting in this recipe.
- Full-disk encryption, turned on before you put anything sensitive through this. Step 2 does it. On Windows Home editions BitLocker may be unavailable, in which case use Device Encryption if your machine offers it - and if it offers neither, do not use this laptop for material that would endanger anyone if the machine were seized.
- Internet once, for the setup only, to download about 4 GB. It does not need to be fast, but it needs to survive a large download - see the USB-stick workaround in the cost note if it will not.
- Permission to install software on the laptop. On a locked-down or managed machine you may not be able to install Python.
- About 100 minutes, most of it waiting for downloads. You can walk away during steps 14, 16 and 17.
على أي أجهزة يعمل هذا
READ THIS BEFORE YOU INSTALL ANYTHING, because one common laptop cannot run this recipe at all.
WINDOWS 10 or 11, 64-bit Intel or AMD: fully supported and the best-tested path. This is the majority platform in the room and every step below gives the Windows command first. Download is about 4 GB. Two Windows-specific things you will meet: you must tick "Add python.exe to PATH" in the Python installer, and you must allow PowerShell to run the activation script (step 11 explains exactly what that means before you type it).
macOS ON APPLE SILICON (M1, M2, M3, M4) RUNNING macOS 14 OR LATER: fully supported and the fastest of the three. Download is about 4 GB. Needs two extra setup steps (Homebrew and portaudio) that Windows does not.
INTEL MACs: THIS RECIPE CANNOT BE MADE TO WORK ON YOUR MACHINE, and you should stop here rather than spend your afternoon on it. PyTorch, which RealtimeSTT requires, stopped publishing Intel-Mac builds after version 2.2.2 in March 2024. I checked the current release (2.13.0) and every macOS build in it is Apple-Silicon-only; there is no Intel Mac build to install. The install will fail at step 14 and no setting fixes it. Use a Windows laptop, an Apple Silicon Mac, or the "someone else runs it" route in the failure modes list.
APPLE SILICON MACs ON macOS 12 OR 13: also blocked. The current PyTorch Apple Silicon build requires macOS 14 or later. Update macOS to 14+ first, or use another machine.
UBUNTU 22.04+ / LINUX x86_64 or ARM64: supported, with one mandatory extra step that Windows and macOS do not need. If you skip step 13, pip will download several gigabytes of NVIDIA CUDA graphics-card libraries onto a laptop that has no NVIDIA graphics card, turning a 4 GB download into an 8-10 GB one. Step 13 is one line and saves roughly 4 GB. Do not skip it on a metered or slow connection.
CHROMEBOOKS, TABLETS, PHONES: not supported, no workaround.
الخطوات
خطوة 1 / 23كل الأنظمة
Check your laptop has enough memory, and check which kind of Mac you have if you are on a Mac. ON WINDOWS: press the Windows key, type 'About your PC', press Enter, and read 'Installed RAM'. ON macOS: click the Apple menu, then About This Mac, and read both the Memory line and the Chip line. ON UBUNTU: open Settings, then About.
تأكد من أنها نجحت
You need 8 GB or more of RAM. On a Mac you additionally need the Chip line to say Apple M1, M2, M3 or M4 (NOT 'Intel'), and the macOS version to be 14 or higher. If all of that is true, continue.
ملاحظة
STOP HERE, AND DO NOT SPEND THE AFTERNOON, IF: it says 4 GB of RAM (no setting will fix this); or your Mac's chip line says Intel (PyTorch has published no Intel-Mac build since March 2024, so the install cannot complete); or your Apple Silicon Mac is on macOS 12 or 13 (update to 14+ first). In any of those cases use the 'someone else runs it and shares the output' route instead, or borrow a Windows laptop.
خطوة 2 / 23كل الأنظمة
Turn on full-disk encryption, before you put anything sensitive through this tool. This is the step that makes 'it stays on my laptop' actually mean something, because laptops get seized, stolen and repaired. ON WINDOWS: press the Windows key, type 'Manage BitLocker', press Enter, and switch BitLocker on for your C: drive; save the recovery key somewhere that is not the same laptop. ON macOS: Apple menu, System Settings, Privacy & Security, scroll to FileVault, and turn it on. ON UBUNTU: whole-disk LUKS encryption can only be chosen while installing Ubuntu; if you did not choose it then, you cannot add it now without reinstalling, so treat this laptop as unencrypted and plan accordingly.
تأكد من أنها نجحت
Windows: the BitLocker panel says 'BitLocker on' for your C: drive. macOS: the FileVault panel says 'FileVault is turned on'. Encryption may then run in the background for an hour or more; you can carry on with the recipe while it does.
ملاحظة
Why this is in a translation recipe: an 8 GB laptop pages memory out to a swap file, and fragments of the audio buffer and of the decoded transcript can land there and survive after you close the program. macOS encrypts swap by default; many Ubuntu installs do not; on Windows the page file is covered once BitLocker is on. This one step closes the most likely place sensitive content actually ends up on disk. If BitLocker is missing on Windows Home, look for 'Device encryption' in Settings, Privacy & security.
خطوة 3 / 23كل الأنظمة
Open a terminal window. This is the black or white text window you will type every command into. ON WINDOWS: press the Windows key, type PowerShell, and press Enter. ON macOS: hold Command and press the space bar, type the word Terminal, and press Enter. ON UBUNTU: hold Ctrl and Alt and press T.
تأكد من أنها نجحت
A window opens with a blinking cursor waiting for you to type. That is all you need.
ملاحظة
Nothing you type here is sent anywhere. It is just a way of giving your own computer instructions. When a command below is on more than one line, copy the whole block at once.
خطوة 4 / 23كل الأنظمة
Find out whether Python is already installed and which version. Copy the line for your system, paste it into the terminal, and press Enter.
انسخ هذا كما هو
WINDOWS: python --version macOS / UBUNTU: python3 --versionتأكد من أنها نجحت
You need it to print a version starting with 3.11 or 3.12 - for example 'Python 3.12.7'. If it does, skip step 5 entirely and go to step 6.
ملاحظة
If it says 3.13, 3.14 or higher, that is TOO NEW: RealtimeSTT 1.1.2 declares its supported Python range as 3.11 up to but not including 3.13, and pip will simply refuse to install it. If it says 3.10 or lower, or 'command not found' or 'not recognized', it is too old or missing. In all of those cases do step 5. Having a newer Python already installed is not a problem - step 5 installs 3.12 alongside it without removing anything.
خطوة 5 / 23كل الأنظمة
ONLY IF step 4 did not print 3.11 or 3.12: install Python 3.12. Open https://www.python.org/downloads/ in your browser, scroll down to 'Looking for a specific release?', find the most recent release whose number starts with 3.12, click it, scroll to the bottom, and download the installer for your system (Windows installer 64-bit, or macOS 64-bit universal2 installer). Run it.
انسخ هذا كما هو
https://www.python.org/downloads/تأكد من أنها نجحت
Close the terminal window completely, open a new one, and run step 4 again. It must now print a version starting with 3.12. If it still does not, the PATH box below is almost always the reason on Windows.
ملاحظة
ON WINDOWS THIS MATTERS MORE THAN ANYTHING ELSE ON THIS PAGE: on the very first installer screen, tick the box that says 'Add python.exe to PATH' before you click Install Now. It is easy to miss and if you miss it, none of the later commands will work and the errors will not tell you why. If you already clicked through without ticking it, run the installer again, choose Modify, and complete it - or simply uninstall and reinstall with the box ticked.
خطوة 6 / 23ماك
macOS ONLY: install Homebrew, the tool macOS needs in order to get microphone support. Skip this step entirely on Windows and Linux. Paste the line below and press Enter, then follow the prompts - it will ask for your Mac login password, and the password will not appear on screen as you type it, which is normal.
انسخ هذا كما هو
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"تأكد من أنها نجحت
When it finishes, close the terminal, open a new one, and run: brew --version It should print something like 'Homebrew 4.x.x'. If it says 'command not found', you missed the two 'Next steps' lines described below.
ملاحظة
UNDERSTAND WHAT YOU ARE AUTHORISING HERE, because this is a pattern criminals also use. This command downloads a script from Homebrew's official address and runs it immediately, with your admin password. It is the genuinely standard way to install Homebrew and it is safe from this exact address - but the shape of it ('paste this line, give it your password') is the shape of many attacks. The rule to carry away: only ever do this for software you went looking for, from the vendor's own documented address that you typed or verified yourself. Never because a message, an email or a stranger told you to. This takes 5-15 minutes. At the end it may print two 'Next steps' lines beginning with 'echo' and 'eval' - copy and run those two lines exactly as printed, then close and reopen the terminal.
خطوة 7 / 23ماك
macOS ONLY: install the microphone library. Skip on Windows and Linux.
انسخ هذا كما هو
brew install portaudioتأكد من أنها نجحت
It ends without a red error, and 'brew list portaudio' prints a list of files. If it says portaudio is already installed, that is fine.
ملاحظة
Without this, the pip install in step 14 fails with an error mentioning 'portaudio.h'. This is the single most common macOS failure in this recipe.
خطوة 8 / 23لينكس
UBUNTU/LINUX ONLY: install the system packages for microphone and audio. Skip on macOS and Windows. It will ask for your login password.
انسخ هذا كما هو
sudo apt-get update && sudo apt-get install -y python3-dev python3-venv portaudio19-dev ffmpeg libsndfile1تأكد من أنها نجحت
The last line of output should be something like 'Processing triggers...' with no 'E:' error lines. If it reports unmet dependencies, run 'sudo apt-get update' on its own first and try again.
ملاحظة
Windows and macOS need no equivalent of this step. Windows users go straight to step 9.
خطوة 9 / 23كل الأنظمة
Make a folder to keep everything in, and move into it. Copy the two lines for your system and paste them together.
انسخ هذا كما هو
WINDOWS PowerShell: mkdir $HOME\live-translate cd $HOME\live-translate macOS / UBUNTU: mkdir -p ~/live-translate cd ~/live-translateتأكد من أنها نجحت
Your terminal prompt should now show that you are inside the live-translate folder. To be sure, run 'pwd' on macOS/Ubuntu or 'pwd' on Windows PowerShell - it must end in live-translate.
ملاحظة
On Windows the folder appears at C:\Users\YourName\live-translate. From here on, every command assumes you are inside this folder. If you close the terminal and come back later, move back into it first with 'cd $HOME\live-translate' on Windows or 'cd ~/live-translate' on macOS and Ubuntu. If mkdir says the folder already exists, that is fine - just run the 'cd' line.
خطوة 10 / 23كل الأنظمة
Create a private workspace for the software so it cannot disturb anything else on your laptop.
انسخ هذا كما هو
WINDOWS: python -m venv .venv macOS / UBUNTU: python3 -m venv .venvتأكد من أنها نجحت
It takes a few seconds and prints nothing. Confirm it worked by listing the folder: run 'dir' on Windows or 'ls -a' on macOS/Ubuntu, and check that a '.venv' folder is now there.
ملاحظة
If you ever want to undo this whole recipe, deleting the live-translate folder removes everything except the Python installation itself.
خطوة 11 / 23كل الأنظمة
Switch the terminal into that workspace. Use the line for your system.
انسخ هذا كما هو
WINDOWS PowerShell: .\.venv\Scripts\Activate.ps1 macOS / UBUNTU: source .venv/bin/activateتأكد من أنها نجحت
'(.venv)' must appear at the start of your terminal line. If it does not, the next steps will install software into the wrong place. Do not continue until you see it.
ملاحظة
IF WINDOWS REFUSES with a message about 'running scripts is disabled on this system', run this first and then try the activation line again: Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass What that does, in plain terms: Windows blocks script files by default as a safety measure, and this lifts the block for THIS TERMINAL WINDOW ONLY - the '-Scope Process' part means it forgets the moment you close the window, and it changes nothing permanently on your computer. It is the normal way to activate a Python workspace. Be wary of any instruction that tells you to run the same command WITHOUT '-Scope Process', or with '-Scope LocalMachine', because that would lower the setting for your whole computer permanently. YOU MUST DO THIS ACTIVATION STEP EVERY SINGLE TIME you open a new terminal to use this tool.
خطوة 12 / 23كل الأنظمة
Update the installer tool itself, so the next steps do not fail on an old version.
انسخ هذا كما هو
python -m pip install -U pip setuptools wheelتأكد من أنها نجحت
It ends with a line beginning 'Successfully installed' (or 'Requirement already satisfied' for all three). Takes under a minute.
ملاحظة
Note it is now 'python', not 'python3' - inside the activated workspace, 'python' is the right one on all systems. If this errors with 'no such file or directory', you are not inside the workspace: go back to step 11.
خطوة 13 / 23لينكس
UBUNTU/LINUX ONLY, AND DO NOT SKIP IT: install the CPU-only version of PyTorch first. Skip this step on Windows and macOS. This single line saves you roughly 4 GB of download.
انسخ هذا كما هو
python -m pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpuتأكد من أنها نجحت
Run: python -c "import torch; print(torch.__version__)" It must print a version ending in '+cpu', for example '2.13.0+cpu'. If it prints a version with no '+cpu' on the end, the plain version was installed instead - uninstall it with 'python -m pip uninstall -y torch torchaudio' and run this step again.
ملاحظة
WHY THIS EXISTS: on Linux, and only on Linux, the ordinary PyTorch package unconditionally drags in several NVIDIA CUDA graphics-card libraries - I checked the current release and nvidia-cudnn-cu13, nvidia-nccl-cu13, nvidia-cusparselt-cu13, nvidia-nvshmem-cu13 and triton are all marked as required whenever the system is Linux. Your laptop almost certainly has no NVIDIA graphics card, and this recipe runs entirely on the CPU anyway, so those gigabytes are pure waste on a metered connection. The Linux CPU-only build is 192 MB against 527 MB plus the CUDA packages. Doing this BEFORE step 14 means step 14 finds PyTorch already present and does not fetch its own copy.
خطوة 14 / 23كل الأنظمة
Install the software: the microphone-and-listening tool with its voice detection, the translation engine, and the helper libraries. This is the longest download of the recipe. Leave it running.
انسخ هذا كما هو
python -m pip install "RealtimeSTT[recommended]" ctranslate2 transformers sentencepiece protobufتأكد من أنها نجحت
It must end with a long line beginning 'Successfully installed' that includes realtimestt, faster-whisper, silero-vad, onnxruntime and ctranslate2. Step 15 checks the critical part properly - do not skip it.
ملاحظة
Expect 10-40 minutes depending on your connection. THE WORD 'recommended' IS LOAD-BEARING AND MUST NOT BE CHANGED. An earlier version of this recipe said RealtimeSTT[faster-whisper] instead. That option installs the speech recogniser but NOT the voice-detection model, and RealtimeSTT then silently downloads that model from github.com every time you start the program - which would have made this recipe's central promise false. 'recommended' installs faster-whisper 1.2.1 plus silero-vad with its CPU ONNX runtime, and the silero-vad package carries its own model file inside it, so nothing is fetched later. If the install stops with a red error mentioning 'portaudio.h', you skipped step 7 (macOS) or step 8 (Ubuntu). If it stops with a network timeout, run the identical command again - it resumes rather than starting over.
خطوة 15 / 23كل الأنظمة
Prove the voice-detection model is really on your disk. This is the check that makes the offline promise true rather than hopeful. Paste the whole line.
انسخ هذا كما هو
python -c "import os, silero_vad, onnxruntime; p=os.path.join(os.path.dirname(silero_vad.__file__),'data','silero_vad_op18_ifless.onnx'); print('VAD model on disk:', os.path.exists(p)); print('onnxruntime version:', onnxruntime.__version__)"تأكد من أنها نجحت
It must print exactly 'VAD model on disk: True' and then an onnxruntime version number. If it prints 'False', or fails with 'No module named silero_vad' or 'No module named onnxruntime', then step 14 did not install the right thing - re-run step 14 and check you typed [recommended] and not [faster-whisper].
ملاحظة
What you just checked: RealtimeSTT looks for this exact file inside the installed silero_vad package before it falls back to downloading a voice-detection model from the internet. Because the file is there, the fallback never runs. If this prints False and you continue anyway, your program will quietly contact github.com every time it starts, and no offline setting in the script will stop it.
خطوة 16 / 23كل الأنظمة
Download the translation model and shrink it so it fits comfortably in memory. This downloads about 2.5 GB and writes out a much smaller compressed model. Paste the whole block as one command.
انسخ هذا كما هو
ct2-transformers-converter --model facebook/nllb-200-distilled-600M --output_dir models/nllb-600m-int8 --quantization int8 --copy_files tokenizer.json tokenizer_config.json special_tokens_map.json sentencepiece.bpe.modelتأكد من أنها نجحت
When it finishes, list the output folder - 'dir models\nllb-600m-int8' on Windows, 'ls -lh models/nllb-600m-int8' on macOS/Ubuntu. You must see a 'model.bin' of roughly 600 MB (about 623 MB), plus tokenizer.json, tokenizer_config.json, special_tokens_map.json and sentencepiece.bpe.model. IF model.bin IS MISSING OR IS ONLY A FEW KILOBYTES, the conversion failed even if the command looked calm - delete the whole models folder and run this step again. A folder that exists but is nearly empty is the failure to watch for here.
ملاحظة
Expect 10-30 minutes. This is NLLB-200 distilled 600M, the model that gives you Kurdish Sorani (ckb_Arab), Kurmanji (kmr_Latn) and Iraqi Arabic (acm_Arab). The '--quantization int8' part is what makes it run on an 8 GB laptop: it turns a 2.46 GB model into roughly 623 MB. Converting it yourself rather than downloading someone's pre-converted copy means you know exactly what you are running. Do NOT delete any download cache yet - step 18 tells you exactly what is safe to delete and when.
خطوة 17 / 23كل الأنظمة
Download the listening model. Paste the whole line.
انسخ هذا كما هو
python -c "from faster_whisper import WhisperModel; WhisperModel('small', device='cpu', compute_type='int8'); print('listening model ready')"تأكد من أنها نجحت
It must print 'listening model ready' as the last line. That fetches about 484 MB. If it ends with an error mentioning a connection or a repository, your download did not finish - run the identical line again.
ملاحظة
'small' is the right starting size for an 8 GB laptop. To try the more accurate but heavier 'medium' (1.53 GB) later, run the same line with 'medium' in place of 'small' - read the model tier table in step 22's note first.
خطوة 18 / 23كل الأنظمة
OPTIONAL, and only worth doing if you are short of disk space: reclaim about 2.5 GB by deleting the NLLB download cache. You already converted that model in step 16, so this copy is no longer needed. Use the line for your system EXACTLY as written.
انسخ هذا كما هو
WINDOWS PowerShell: Remove-Item -Recurse -Force "$HOME\.cache\huggingface\hub\models--facebook--nllb-200-distilled-600M" macOS / UBUNTU: rm -rf ~/.cache/huggingface/hub/models--facebook--nllb-200-distilled-600Mتأكد من أنها نجحت
Immediately afterwards, re-run the step 17 line. It must still print 'listening model ready' WITHOUT downloading anything again. If it starts downloading 484 MB again, you deleted too much - let it finish and do not repeat this step.
ملاحظة
NOTE THE PRECISE PATH: it ends with the named nllb folder. An earlier version of this recipe told you to delete the whole '~/.cache/huggingface/hub' directory, which is also where your Whisper listening model lives - doing that after step 17 silently destroys the listening model and produces an offline failure that is very hard to diagnose. This version deletes only the one folder that is genuinely finished with, so it is safe whenever you run it. If you are not short of space, just skip this step; keeping the cache costs you nothing but disk.
خطوة 19 / 23كل الأنظمة
Create the program file and open it for editing. Use the command for your system.
انسخ هذا كما هو
WINDOWS: notepad translate_live.py macOS / UBUNTU: nano translate_live.pyتأكد من أنها نجحت
A text editor opens. On Windows, Notepad first asks whether to create the file - click Yes. You should be looking at an empty document.
ملاحظة
You now have an empty text editor open, waiting for the program in the next step.
خطوة 20 / 23كل الأنظمة
Copy the entire block below - all of it, from the first line to the last - and paste it into the editor you just opened. Then save: in Notepad, hold Ctrl and press S, then close the window. In nano, hold Ctrl and press O, press Enter, then hold Ctrl and press X.
انسخ هذا كما هو
# Live speech translation that never leaves your laptop. # Adapted from CDR's 'simultane-ceviri' workshop module (November 2024, MIT licence), # rebuilt for 2026 with NLLB-200 in place of M2M100 so that Kurdish Sorani works at all. import os # Refuse to touch the Hugging Face network. The voice-detection model is loaded # from the installed silero-vad package, which ships its own model file, so # nothing reaches the network at run time. os.environ["HF_HUB_OFFLINE"] = "1" os.environ["TRANSFORMERS_OFFLINE"] = "1" import ctranslate2 import transformers from RealtimeSTT import AudioToTextRecorder # ---------------- CHANGE THESE THREE LINES TO PICK YOUR LANGUAGES ---------------- LISTEN_LANGUAGE = "ar" # the language you SPEAK, in Whisper's code NLLB_SOURCE = "arb_Arab" # the SAME language, in NLLB's code NLLB_TARGET = "ckb_Arab" # the language you want to READ on screen # --------------------------------------------------------------------------------- WHISPER_MODEL = "small" # tiny | base | small | medium | large-v3 MODEL_DIR = "models/nllb-600m-int8" CPU_THREADS = 4 # IMPORTANT: the translation model is loaded lazily, inside get_translator(). # RealtimeSTT starts two extra processes with the "spawn" method on macOS and # Windows, and each of those re-imports this file. If the model were loaded at # the top of the file it would load THREE times and waste about 2 GB of memory # on an 8 GB laptop. Do not move this out of the function. _TRANSLATOR = None _TOKENIZER = None def get_translator(): global _TRANSLATOR, _TOKENIZER if _TRANSLATOR is None: print("Loading the translation model. This takes 10 to 60 seconds.") _TRANSLATOR = ctranslate2.Translator( MODEL_DIR, device="cpu", compute_type="int8", inter_threads=1, intra_threads=CPU_THREADS, ) _TOKENIZER = transformers.AutoTokenizer.from_pretrained( MODEL_DIR, src_lang=NLLB_SOURCE ) return _TRANSLATOR, _TOKENIZER def clear_screen(): os.system("cls" if os.name == "nt" else "clear") def translate(text): translator, tokenizer = get_translator() tokens = tokenizer.convert_ids_to_tokens(tokenizer.encode(text)) results = translator.translate_batch( [tokens], target_prefix=[[NLLB_TARGET]], beam_size=2, max_decoding_length=256, ) out = results[0].hypotheses[0][1:] # drop the language tag we forced return tokenizer.decode(tokenizer.convert_tokens_to_ids(out)) def process_text(heard): if not heard or not heard.strip(): return clear_screen() print("HEARD (" + LISTEN_LANGUAGE + "): " + heard) print("") print("READS (" + NLLB_TARGET + "): " + translate(heard)) print("") print("-" * 60) print("Machine output. Check it before you use it. Press Ctrl and C to stop.") if __name__ == "__main__": get_translator() # load once, here in the parent process only recorder = AudioToTextRecorder( model=WHISPER_MODEL, language=LISTEN_LANGUAGE, device="cpu", # the default is "cuda"; on a laptop this MUST be "cpu" compute_type="int8", post_speech_silence_duration=0.7, beam_size=3, spinner=False, ) print("Ready. Speak now.") try: while True: recorder.text(process_text) except KeyboardInterrupt: print("Stopping.") finally: recorder.shutdown()تأكد من أنها نجحت
Back in the terminal, run: python -c "import ast; ast.parse(open('translate_live.py').read()); print('script file is valid')" It must print 'script file is valid'. If it reports a SyntaxError with a line number, the paste was incomplete or the indentation was mangled - open the file again, delete everything, and re-paste the whole block.
ملاحظة
Do not retype this by hand - copy and paste it. Python is sensitive to indentation and a single wrong space stops it running. The three lines under 'CHANGE THESE THREE LINES' are the only ones you ever need to edit. Do not move the model loading out of get_translator(): an earlier version of this recipe loaded it at the top of the file, and because RealtimeSTT starts two helper processes that re-import this file, the translation model was loaded three times over - about 2 GB of wasted memory on exactly the 8 GB laptops this recipe is written for, which then showed up as the machine being unbearably slow for reasons the recipe blamed on the wrong thing.
خطوة 21 / 23كل الأنظمة
Set your language pair. As written, the file listens to Modern Standard Arabic and shows Kurdish Sorani. To change it, reopen the file ('notepad translate_live.py' on Windows, 'nano translate_live.py' on macOS/Ubuntu), edit only the three marked lines, and save the same way as before.
تأكد من أنها نجحت
After saving, run the same validity check as in step 20: python -c "import ast; ast.parse(open('translate_live.py').read()); print('script file is valid')" If you accidentally deleted a quotation mark while editing, this catches it now rather than in the middle of a meeting.
ملاحظة
LANGUAGES YOU CAN SPEAK (LISTEN_LANGUAGE, Whisper codes): ar = Arabic, en = English, tr = Turkish, fa = Persian, fr = French. KURDISH IS NOT ON THIS LIST AND CANNOT BE ADDED - Whisper has 100 languages and no Kurdish in any script. LANGUAGES YOU CAN READ OR WRITE (NLLB_SOURCE and NLLB_TARGET, NLLB codes): arb_Arab = Modern Standard Arabic, acm_Arab = Iraqi/Mesopotamian Arabic, apc_Arab = Levantine Arabic, arz_Arab = Egyptian Arabic, ckb_Arab = Kurdish Sorani, kmr_Latn = Kurdish Kurmanji, eng_Latn = English, tur_Latn = Turkish, pes_Arab = Persian. NLLB_SOURCE must be the same language as LISTEN_LANGUAGE - if you set LISTEN_LANGUAGE to 'en' then NLLB_SOURCE must be 'eng_Latn'. Common working pairs: Arabic in, Sorani out (ar / arb_Arab / ckb_Arab); Arabic in, English out (ar / arb_Arab / eng_Latn); English in, Sorani out (en / eng_Latn / ckb_Arab).
خطوة 22 / 23كل الأنظمة
Run it. Make sure '(.venv)' is still showing at the start of your terminal line, then paste the line below.
انسخ هذا كما هو
python translate_live.pyتأكد من أنها نجحت
The first run takes 1-3 minutes to load. You should see 'Loading the translation model' ONCE - not three times - and then 'Ready. Speak now.' Speak a full sentence and pause for about a second; your words appear on the HEARD line and the translation on the READS line. If 'Loading the translation model' appears three times, the model loading was moved out of get_translator() - re-paste the script from step 20.
ملاحظة
Press Ctrl and C to stop. HOW LONG THE DELAY WILL BE, on a typical 4-core 8 GB laptop with no graphics card: roughly 3-8 seconds between finishing a sentence and seeing the translation. These are estimates, not benchmarks - I did not time this on real hardware. MODEL SIZE TIERS, disk / rough memory while running: tiny 75 MB / ~0.4 GB - fast and too inaccurate for Arabic, use only to check the plumbing works; base 145 MB / ~0.5 GB - still too weak for Arabic; small 484 MB / ~1.5 GB total with the translator - the recommended starting point on 8 GB; medium 1.53 GB / ~3 GB total - noticeably better on Arabic, workable on 8 GB if you close your browser, roughly twice as slow; large-v3 about 3.1 GB / ~4.5 GB total - do not attempt alongside the translator on an 8 GB laptop, it will swap to disk and crawl. On 16 GB, medium is the sweet spot. BEFORE YOU USE THIS ON ANYTHING SENSITIVE, decide where the text is going afterwards: the moment you copy a translation out of this window into Google Docs, Gmail or WhatsApp, you have uploaded it and every promise on this page stops applying.
خطوة 23 / 23كل الأنظمة
Prove to yourself that nothing is leaving the laptop. Stop the program with Ctrl and C. Turn your wifi off completely - click the wifi icon and switch it off - and unplug any network cable. Then run it again and speak.
انسخ هذا كما هو
python translate_live.pyتأكد من أنها نجحت
It must reach 'Ready. Speak now.' and translate your speech exactly as before, with no internet at all. That is the proof. If instead it fails with a message mentioning 'offline', 'Connection error', 'couldn't connect to huggingface.co', or 'Could not initialize any automatic Silero VAD backend', see the note.
ملاحظة
This takes thirty seconds and it is the whole promise of the recipe - do it in front of your colleagues and your director rather than telling them about it. IF IT FAILS: a message about huggingface.co or 'offline' means a model file did not finish downloading, so reconnect and re-run steps 16 and 17. A message containing 'Silero VAD' means the voice-detection model is missing, which means step 14 installed the wrong option - reconnect, re-run step 14 with [recommended] exactly as written, and re-run the check in step 15 until it prints True. Do not present this tool to anyone as offline until this step passes with the wifi genuinely off.
كيف تعرف أن العمل كله قد نجح
Do all five of these before you trust a single word of the output. (1) THE OFFLINE TEST: switch your wifi off completely, unplug any cable, and run the program. If it still translates, nothing is going to a server - this is the proof, not a promise. Do it after the check in step 15 has printed 'VAD model on disk: True', because that check is what makes the offline test meaningful rather than lucky. (2) THE KNOWN-SENTENCE TEST: choose three or four sentences whose correct translation you already know, from a document your organisation has already had translated by a human. Say them into the microphone and compare the output word by word against the human version. Do this in the specific language pair and the specific accent you will actually use, not in English. If it mangles sentences you already know the answer to, it will mangle the ones you do not. (3) THE SILENCE TEST, which catches the most dangerous failure: start the program, then sit in a quiet room and say absolutely nothing for sixty seconds. Nothing should appear on screen. If sentences appear out of silence, you have just watched the model hallucinate - Koenecke and colleagues (FAccT 2024) found about 1% of Whisper transcriptions contained entirely fabricated phrases, that 38% of those included explicit harms such as invented violence or false attributions, and that they cluster exactly around pauses and silence. Repeat this test with a recording that has long gaps in it, because interviews with distressed or hesitant speakers have long gaps. (4) THE NAMES TEST: say five names of people and places you work with. Names, dates and numbers are where machine transcription fails hardest and where an error does the most damage in a documentation context. Check every one against the audio by hand, every time, for as long as you use this tool. (5) THE MEMORY TEST, which takes ten seconds and tells you whether the tool will be usable at all on your machine: start the program and count how many times 'Loading the translation model' appears before 'Ready. Speak now.' It must appear exactly ONCE. If it appears three times, the script has been altered so that the translation model loads in the helper processes too, which wastes about 2 GB on an 8 GB laptop and will make everything crawl - re-paste the script from step 20.
ما الذي قد يفشل، وماذا تفعل حينها
- THE PROGRAM WORKS BUT YOU ARE NOT SURE IT IS REALLY OFFLINE. Do not reason about it - measure it. Run the step 15 check and confirm 'VAD model on disk: True', then run the program with wifi off. If both pass, nothing is reaching the network at run time. This is worth re-checking after any reinstall, because installing RealtimeSTT with the wrong option silently reintroduces a download from github.com at every start-up.
- IT REFUSES TO START WITH 'Could not initialize any automatic Silero VAD backend'. This is the offline-failure message people hit when step 14 was run with [faster-whisper] instead of [recommended], or when the install was interrupted. Reconnect to the internet, re-run step 14 exactly as written, then re-run step 15 until it prints 'VAD model on disk: True'. Note that this message contains neither the word 'offline' nor 'connection', so do not go looking for a download problem.
- YOU ARE ON AN INTEL MAC AND THE INSTALL IN STEP 14 FAILS, often with a message about no matching distribution for torch, or with a long compiler error. There is no fix and nothing you type will help. PyTorch has published no Intel-Mac build since version 2.2.2 in March 2024. Use a Windows laptop or an Apple Silicon Mac. Same answer if you are on an Apple Silicon Mac still running macOS 12 or 13: update to macOS 14 or later first.
- INSTALL FAILS ON macOS WITH AN ERROR MENTIONING 'portaudio.h' OR 'pyaudio'. You skipped step 7. Run 'brew install portaudio', then run step 14 again. On Ubuntu the equivalent is the 'portaudio19-dev' package from step 8.
- 'command not found: python3' OR 'python is not recognized'. Python is not installed, or on Windows you did not tick 'Add python.exe to PATH' during installation. Reinstall Python and tick the box, then close the terminal completely and open a new one.
- PIP REFUSES TO INSTALL RealtimeSTT AT ALL, saying no matching distribution found. Your Python is too new. RealtimeSTT 1.1.2 supports Python 3.11 and 3.12 only. Check with 'python --version' and install 3.12 as in step 5 if it says 3.13 or higher.
- THE INSTALL SUCCEEDS BUT NOTHING HAPPENS WHEN YOU SPEAK, or you get an audio device error. The program is listening to the wrong microphone. Check your operating system's sound settings and set the correct input device as the default, then restart the program. If you have several microphones, RealtimeSTT accepts an input_device_index setting - add input_device_index=1 (then 2, then 3) inside the AudioToTextRecorder brackets until the right one responds.
- 'Loading the translation model' APPEARS THREE TIMES AND THE LAPTOP CRAWLS. The script was altered so the model loads at the top of the file instead of inside get_translator(). RealtimeSTT starts two helper processes that re-import your script, so the model gets built three times and eats roughly 2 GB more memory than it needs. Re-paste the script from step 20 without moving that code.
- IT IS PAINFULLY SLOW - 30 seconds or more per sentence, and the laptop fan roars. First confirm the three-times-loading fault above is not the cause. Otherwise your laptop is running out of memory and swapping to disk: close your browser and every other application. If it is still slow, change WHISPER_MODEL from 'small' to 'base'. If you had set it to 'medium' or 'large-v3' on an 8 GB machine, that is the cause - go back to 'small'.
- IT CRASHES IMMEDIATELY WITH AN ERROR MENTIONING 'cuda' OR 'CUDA driver'. You edited out or mistyped the device="cpu" line. RealtimeSTT defaults to device="cuda", which needs an NVIDIA graphics card that your laptop almost certainly does not have. Put device="cpu" back exactly as written.
- IT REFUSES TO START WITH AN ERROR MENTIONING 'offline' OR 'Connection error' OR 'couldn't connect to huggingface.co'. A model file did not finish downloading. Reconnect to the internet, re-run steps 16 and 17, then try again offline. If you had deleted a cache folder, check you deleted only the nllb folder named in step 18 and not the whole hub directory.
- THE STEP 16 CONVERSION APPEARS TO FINISH BUT models/nllb-600m-int8 IS EMPTY OR model.bin IS TINY. The conversion failed part-way, often after an interrupted download, and it does not always shout about it. Delete the models folder entirely and run step 16 again. Never continue past a models folder that does not contain a roughly 623 MB model.bin.
- THE TRANSLATION IS FLUENT, CONFIDENT, AND ABOUT SOMETHING ELSE ENTIRELY. This is the characteristic failure: the listening step misheard the sentence, and the translation step then translated the mishearing perfectly. The output looks polished because it IS polished - it is just not what was said. Always keep the 'HEARD' line visible next to the 'READS' line, which is why the script prints both, and always check the HEARD line against your memory of the audio before you trust the translation.
- TEXT APPEARS DURING SILENCE, OR THE SAME PHRASE REPEATS OVER AND OVER. This is Whisper hallucination and it is documented, not a fault in your setup. Increase post_speech_silence_duration from 0.7 to 1.2 to reduce it, use a headset microphone, and record in a quieter room. It will not go away completely. Never treat unverified output as a transcript.
- THE OUTPUT IS IN THE WRONG SCRIPT - you asked for Sorani and got Latin letters, or asked for Kurmanji and got Arabic script. You mixed up the codes. Sorani is ckb_Arab (Arabic script) and Kurmanji is kmr_Latn (Latin script). They are different languages in NLLB, not two spellings of one.
- THE OUTPUT IS THE SAME AS THE INPUT, UNTRANSLATED. NLLB_SOURCE and NLLB_TARGET are set to the same language. Check step 21.
- YOU SPOKE KURDISH AND GOT ARABIC-LOOKING NONSENSE. Whisper cannot hear Kurdish at all and is guessing at it as Arabic. There is no setting that fixes this. See the Kurdish section above for the honest state of Sorani speech recognition in 2026.
- THE DOWNLOAD IN STEP 14, 16 OR 17 KEEPS TIMING OUT. Run the identical command again - all three resume where they stopped rather than restarting. If it will not complete on your connection at all, use the USB route in the cost note: a colleague on the same operating system completes steps 9 to 18 and gives you ONLY the 'models' folder and 'translate_live.py'. You still run steps 9 to 15 yourself. Copying their '.venv' folder will not work, because it contains absolute paths and files compiled for their machine.
- YOU COPIED A COLLEAGUE'S WHOLE live-translate FOLDER AND NOTHING RUNS. That is the workaround an earlier version of this recipe described, and it does not work. Delete the copied .venv folder, then do steps 9 to 15 on your own machine, keeping their 'models' folder and 'translate_live.py'.
جودة عمله بلغتك
العربية الفصحى · اللهجة العراقية · الكردية السورانية
العربية الفصحى
Workable as a rough draft, not as a record. Whisper does support Arabic ('ar' is one of its languages) and clear, slow, read Modern Standard Arabic from a good microphone is the best case for this pipeline. I could not retrieve a verified MSA word-error-rate figure in this research pass - treat any specific MSA number you see elsewhere with caution. OpenAI's own large-v3 model card claims only a 10-20% relative error reduction over large-v2 and explicitly warns that 'the predictions may include texts that are not actually spoken in the audio input (i.e. hallucination)' and that the models 'exhibit disparate performance on different accents and dialects'. Expect to correct every paragraph by hand.
اللهجة العراقية
Poor, and you must plan around it. There is a hard fact here: Iraqi/Mesopotamian Arabic is absent from the main public multi-dialect Arabic speech benchmark, so nobody has published a zero-shot Whisper accuracy figure specifically for Iraqi. What we do have is the Casablanca benchmark (EMNLP 2024), which measured Whisper large-v3 zero-shot across eight other Arabic dialects and got word error rates of 48.44% (Jordan), 58.02% (Palestine), 59.11% (Egypt), 62.31% (UAE), 69.94% (Yemen), 83.49% (Algeria), 87.20% (Morocco), 87.44% (Mauritania). The paper's conclusion: 'all models exhibited high WER and CER across each dialect, indicating their inability to effectively generalize.' A separate 2026 study reports zero-shot multilingual Whisper at 78.8% WER on Sudanese. Iraqi sits inside that range, closest in kind to the Gulf and Levantine numbers. SAY IT THIS WAY RATHER THAN QUOTING A SINGLE NUMBER: expect Iraqi colloquial speech to come out at least as badly as Jordanian at 48% of words wrong, and possibly much worse, with a phone recording or a noisy room pushing it further. No one has measured Iraqi, so anyone who gives you a precise Iraqi figure is guessing. Practical mitigation that genuinely helps: ask the speaker to move toward MSA, use an external microphone or headset rather than the laptop's built-in one, record in a quiet room, and use the 'medium' Whisper model rather than 'small' if your laptop can carry it. Separately, NLLB has a distinct code for written Iraqi Arabic, acm_Arab, so on the TEXT side you can at least ask for Iraqi Arabic rather than being forced into MSA - though acm_Arab is a very low-resource direction in a model whose own card says research-only, and I retrieved no quality score for it. The useful distinction stands: you can ask for Iraqi Arabic TEXT even where you cannot recognise Iraqi Arabic SPEECH.
الكردية السورانية
Read this part carefully, because it contains the single most important correction to the original module. Two separate answers. (1) TRANSLATION of Kurdish Sorani: yes, this works. The original 2024 module used M2M100, and M2M100 cannot do Kurdish at all - I checked the facebook/m2m100_418M and facebook/m2m100_1.2B model cards directly and the 100-language list is 'af, am, ar, ast, az, ba, be, bg, bn, br, bs, ca, ceb, cs, cy, da, de, el, en, es, et, fa, ff, fi, fr, fy, ga, gd, gl, gu, ha, he, hi, hr, ht, hu, hy, id, ig, ilo, is, it, ja, jv, ka, kk, km, kn, ko, lb, lg, ln, lo, lt, lv, mg, mk, ml, mn, mr, ms, my, ne, nl, no, ns, oc, or, pa, pl, ps, pt, ro, ru, sd, si, sk, sl, so, sq, sr, ss, su, sv, sw, ta, th, tl, tn, tr, uk, ur, uz, vi, wo, xh, yi, yo, zh, zu' - there is no 'ku' entry, no Kurmanji, no Sorani. The 'ku' code people remember belongs to other model families, not to M2M100. This recipe therefore replaces M2M100 with NLLB-200, which has Central Kurdish as ckb_Arab (Sorani, Arabic script - your locale) and Northern Kurdish as kmr_Latn (Kurmanji, Latin script) as two genuinely distinct codes. Quality is a research model's quality: NLLB's own card says it is 'primarily intended for research', 'not intended to be used with domain specific texts, such as medical domain or legal domain', not for document translation, and that its translations 'can not be used as certified translations'. A July 2026 cross-lingual transfer study found Sorani results 'insufficient for high-stakes clinical deployment'. Use it for gist and for a first draft that a Kurdish speaker then fixes. (2) SPEECH RECOGNITION of Kurdish Sorani: no. This is the loud part. Whisper's language table contains 100 entries and Kurdish is not one of them, in any script or variant - I parsed whisper/tokenizer.py directly this session to confirm: no 'ku', no 'ckb', no 'kmr', no entry whose name contains 'kurd'. (An earlier draft said 99 entries; the correct count is 100, because 'yue' for Cantonese was added for large-v3. The substantive point is untouched: no Kurdish either way.) So the real-time script in this recipe cannot listen to Sorani. It can listen to Arabic and write Sorani; it cannot listen to Sorani. There is no drop-in fix. What does exist in 2026, honestly assessed: Meta's Omnilingual ASR (Apache 2.0) lists 1,424 language-script codes and ckb_Arab, kur_Arab, kmr_Latn and acm_Arab are all present in its lang_ids.py - genuinely the best hope - but RealtimeSTT's own docs mark that engine 'Linux/WSL2 Python 3.11.x only', its published VRAM figures run from about 2 GiB (300M CTC) to about 17-20 GiB (7B LLM), and it only accepts audio clips under 40 seconds on the standard models. It is a batch tool for a technical colleague with a Linux machine, not a live-translation tool for a five-person office. Meta's MMS (facebook/mms-1b-all, CC-BY-NC) has a ckb adapter among its 1,162 languages and will run on CPU, slowly, on recorded files rather than live. Beyond that, the Hugging Face Sorani ASR models are small community fine-tunes with double- and low-triple-digit download counts and no independent evaluation. The absence of a dependable, laptop-runnable, offline Sorani speech recogniser in 2026 is itself the finding, and it is worth saying out loud in Erbil.
على ماذا يستند هذا
M2M100 has no Kurdish: language lists read directly from https://huggingface.co/facebook/m2m100_418M/raw/main/README.md and https://huggingface.co/facebook/m2m100_1.2B (100 codes, no 'ku'). NLLB Kurdish and Arabic codes: FLORES-200 table at https://github.com/facebookresearch/flores/blob/main/flores200/README.md giving 'Central Kurdish: ckb_Arab', 'Northern Kurdish: kmr_Latn', 'Mesopotamian Arabic: acm_Arab', 'Modern Standard Arabic: arb_Arab', 'North Levantine Arabic: apc_Arab', 'Egyptian Arabic: arz_Arab'; re-confirmed this session by reading special_tokens_map.json from facebook/nllb-200-distilled-600M, which contains acm_Arab, apc_Arab, arb_Arab, arz_Arab, ckb_Arab, eng_Latn and kmr_Latn as literal token strings. Whisper has no Kurdish: LANGUAGES dict parsed programmatically this session from https://raw.githubusercontent.com/openai/whisper/main/whisper/tokenizer.py - 100 entries, 'ar' present, 'yue' present, zero entries matching ku/ckb/kmr/kur or the word 'kurdish'. Arabic dialect WER: Casablanca, arXiv 2410.04527, Table 3, Whisper-large-v3 zero-shot WER/CER per dialect. Sudanese zero-shot Whisper 78.8% WER: arXiv 2601.06802. Sorani MT quality: arXiv 2607.22300. Hallucination: Koenecke, Choi, Mei, Schellmann and Sloane, 'Careless Whisper: Speech-to-Text Hallucination Harms', FAccT 2024, arXiv 2402.08021 - about 1% of transcriptions contained entirely hallucinated phrases, 38% of those hallucinations included explicit harms, and they clustered around speakers with longer silent pauses. Omnilingual ASR Kurdish coverage: ckb_Arab, kur_Arab, kmr_Latn and acm_Arab all present in the 1,424-entry supported_langs list in facebookresearch/omnilingual-asr src/omnilingual_asr/models/wav2vec2_llama/lang_ids.py. UNTESTED: I did not run this pipeline end to end on Arabic or Kurdish audio; the code is syntax-checked and every API is verified against current documentation, but the accuracy claims come from published benchmarks, not from my own recording.
البرامج المذكورة في هذه الوصفة
- RealtimeSTT 1.1.2 (MIT, KoljaB) - microphone capture, voice activity detection and the loop. Python 3.11-3.12 only. Install with the [recommended] extra, not [faster-whisper]
- silero-vad 6.2.1 (MIT) with its onnx-cpu extra - the voice-detection model, installed via RealtimeSTT[recommended]; the 9.1 MB wheel bundles its own ONNX model files, which is what removes the run-time download
- faster-whisper 1.2.1 (MIT, SYSTRAN, released 31 October 2025) - the speech recogniser
- CTranslate2 4.8.1 (MIT, OpenNMT) - the CPU inference engine, with cp311/cp312 wheels for macOS arm64 and x86_64, manylinux_2_28 x86_64 and aarch64, and win_amd64
- onnxruntime (MIT, Microsoft) - runs the voice-detection model on the CPU; pulled in automatically by silero-vad[onnx-cpu]
- OpenAI Whisper small/medium (MIT weights) via Systran/faster-whisper-small (model.bin is 483,546,902 bytes) and Systran/faster-whisper-medium
- facebook/nllb-200-distilled-600M (CC-BY-NC-4.0) - the translation model, quantised to int8 locally
- PyTorch 2.13.0 - required by RealtimeSTT. On Linux install the CPU-only build from download.pytorch.org/whl/cpu (192 MB) instead of the PyPI default (527 MB plus several GB of NVIDIA CUDA packages)
- transformers and sentencepiece - only for the NLLB tokenizer
- REPLACED AND NO LONGER USED: facebook/m2m100_1.2B and michaelfeil/ct2fast-m2m100_1.2B (M2M100 has no Kurdish at all), and hf_hub_ctranslate2, whose last PyPI release was 2.13.1 on 26 June 2024 and which is treated here as unmaintained
- ROUTE DELETED IN THIS REVISION: Intel Macs. PyTorch published its last Intel-Mac wheels in version 2.2.2 (March 2024); the current 2.13.0 release ships macOS builds for Apple Silicon only, so there is no working install path
- ADVICE DELETED IN THIS REVISION: the hf-mirror.com download mirror, which was an unaudited third party for this audience
- MENTIONED BUT NOT USED IN THIS RECIPE: facebook/omnilingual-asr (Apache 2.0, 1,424 languages including ckb_Arab, kur_Arab and acm_Arab, but Linux/WSL2 and GPU-shaped) and facebook/mms-1b-all (CC-BY-NC, 1,162 languages including ckb, CPU-runnable but not live)
المصادر
- https://github.com/cdrorgtr/cdr-ai/tree/main/simultane-ceviri - the original CDR workshop module, November 2024, MIT licensed, adapted here with credit
- https://pypi.org/pypi/RealtimeSTT/json - version 1.1.2, requires_python '<3.13,>=3.11', and the extras table showing that 'recommended' = faster-whisper==1.2.1 plus silero-vad[onnx-cpu]>=6.2.1 while 'faster-whisper' = faster-whisper==1.2.1 alone; base requirements include torch and torchaudio unpinned
- RealtimeSTT/core/silero_vad.py, read from the realtimestt-1.1.2 wheel this session - the _create_auto_vad backend chain, which tries silero_vad_op18_ifless.onnx and silero_vad.onnx from the installed silero_vad package before falling back to torch.hub.load('snakers4/silero-vad')
- RealtimeSTT/core/initialization.py and core/runtime.py, read from the same wheel - mp.set_start_method('spawn') and mp.Process, the reason the script must not build the translation model at module level
- https://pypi.org/pypi/silero-vad/json - version 6.2.1, released 24 February 2026, wheel 9,146,242 bytes, extras onnx-cpu / onnx-gpu / test; onnx-cpu pulls onnxruntime>=1.16.1
- silero_vad-6.2.1-py3-none-any.whl contents, listed this session - silero_vad/data/ contains silero_vad_op18_ifless.onnx (2,845,718 bytes), silero_vad.onnx, silero_vad.jit and others, confirming the model ships inside the package
- https://pypi.org/pypi/ctranslate2/json and .../4.8.1/json - version 4.8.1, requires_python >=3.9, cp311 and cp312 wheels for macosx_11_0_arm64, macosx_11_0_x86_64, manylinux_2_28 x86_64 and aarch64, and win_amd64
- https://pypi.org/pypi/faster-whisper/json - version 1.2.1
- https://pypi.org/pypi/torch/json and .../2.13.0/json - torch 2.13.0 requires nvidia-cudnn-cu13, nvidia-nccl-cu13, nvidia-cusparselt-cu13, nvidia-nvshmem-cu13 and triton whenever platform_system == 'Linux'; Linux x86_64 wheel 526.6 MB, Windows 122.1 MB, macOS wheels are macosx_14_0_arm64 only at 111.2 MB, and there is no Intel-Mac wheel
- https://pypi.org/pypi/torch/2.2.2/json - the last torch release containing macosx_10_9_x86_64 wheels, i.e. the last with Intel-Mac support; checked 2.3.0 and 2.4.1 and both have none
- https://download.pytorch.org/whl/cpu/torch/ - the CPU-only index, containing torch-2.13.0+cpu for cp311 and cp312 on manylinux_2_28 x86_64 and aarch64; the x86_64 CPU wheel is 191,817,609 bytes against 526.6 MB for the default
- https://raw.githubusercontent.com/KoljaB/RealtimeSTT/master/README.md - current engine list and quickstart API
- https://raw.githubusercontent.com/KoljaB/RealtimeSTT/master/docs/installation.md - exact per-OS install commands
- https://raw.githubusercontent.com/KoljaB/RealtimeSTT/master/docs/configuration.md - AudioToTextRecorder parameter names and defaults, including device defaulting to cuda
- https://raw.githubusercontent.com/KoljaB/RealtimeSTT/master/docs/transcription-engines.md - engine strings, and the Linux/WSL2 Python 3.11 restriction on the Omnilingual ASR engine
- https://pypi.org/pypi/hf-hub-ctranslate2/json - last release 2.13.1, 26 June 2024, treated as unmaintained
- https://opennmt.net/CTranslate2/guides/transformers.html - official NLLB conversion command and the exact src_lang / target_prefix inference pattern used in the script
- https://huggingface.co/api/models/facebook/nllb-200-distilled-600M - licence cc-by-nc-4.0 and the repository file list, confirming that all four files passed to --copy_files (tokenizer.json, tokenizer_config.json, special_tokens_map.json, sentencepiece.bpe.model) actually exist
- https://huggingface.co/facebook/nllb-200-distilled-600M - CC-BY-NC licence, research-only intended use, not for medical or legal domains, not certified translation, 512-token limit
- https://huggingface.co/facebook/nllb-200-distilled-600M/raw/main/special_tokens_map.json - contains acm_Arab, apc_Arab, arb_Arab, arz_Arab, ckb_Arab, eng_Latn and kmr_Latn as literal language tokens
- https://huggingface.co/facebook/m2m100_418M/raw/main/README.md - the 100-language list, showing no Kurdish
- https://huggingface.co/facebook/m2m100_1.2B - same, for the 1.2B model the original module used
- https://github.com/facebookresearch/flores/blob/main/flores200/README.md - ckb_Arab, kmr_Latn, acm_Arab, arb_Arab, apc_Arab, arz_Arab, ars_Arab
- https://raw.githubusercontent.com/openai/whisper/main/whisper/tokenizer.py - the LANGUAGES dict, parsed programmatically this session: 100 entries, 'ar' and 'yue' present, no Kurdish under any of ku, ckb, kmr or kur
- https://huggingface.co/openai/whisper-large-v3 - 1.55B parameters, explicit hallucination warning, disparate performance across dialects
- https://arxiv.org/html/2410.04527v1 - Casablanca, EMNLP 2024, Table 3 zero-shot Whisper large-v3 WER by Arabic dialect
- https://arxiv.org/abs/2402.08021 - Koenecke et al., 'Careless Whisper: Speech-to-Text Hallucination Harms', FAccT 2024
- https://arxiv.org/abs/2506.02627 - 'Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning', includes Iraqi among five dialects studied
- https://arxiv.org/abs/2607.22300 - cross-lingual transfer study finding Sorani Kurdish translation insufficient for high-stakes deployment
- https://raw.githubusercontent.com/facebookresearch/omnilingual-asr/main/src/omnilingual_asr/models/wav2vec2_llama/lang_ids.py - 1,424 codes including ckb_Arab, kur_Arab, kmr_Latn, acm_Arab
- https://github.com/facebookresearch/omnilingual-asr - model sizes, VRAM table, 40-second audio limit
- https://huggingface.co/facebook/mms-1b-all - 1,162 languages including ckb, CC-BY-NC
- https://huggingface.co/Systran/faster-whisper-small - model.bin content length 483,546,902 bytes, i.e. 484 MB
من إنتاج CDR، وكُتبت لمؤتمر بوينت العراق 7. جرى التحقق منها مقابل صفحات المزوّدين وسجلات الحزم في آب/أغسطس 2026؛ والمصادر مدرجة في هذه الصفحة.
ما لم نتمكن من التحقق منه
WHAT I VERIFIED AND WHAT I DID NOT, updated after this repair pass. Verified by opening the actual source, this session: that RealtimeSTT's 'recommended' extra really is faster-whisper==1.2.1 plus silero-vad[onnx-cpu]>=6.2.1, and that the plain 'faster-whisper' extra really does omit the voice detector; that the silero-vad wheel really does contain silero_vad_op18_ifless.onnx, which is exactly the first file RealtimeSTT's automatic backend looks for, so the run-time download to github.com genuinely does not happen once it is installed - I read both the wheel contents and RealtimeSTT's own _create_auto_vad chain to confirm this rather than trusting the package description; that Python multiprocessing under 'spawn' re-imports the main script into every child process, which I demonstrated with a six-line test on this machine, confirming the triple model load the earlier draft would have caused; that torch 2.13.0 makes five NVIDIA CUDA packages unconditional on Linux and that the CPU-only index carries torch 2.13.0+cpu for Python 3.11 and 3.12 at 192 MB against 527 MB; every package version and release date; that all four files passed to --copy_files exist in the NLLB repository; that NLLB's own token list contains ckb_Arab, kmr_Latn and acm_Arab; that Whisper's language table has no Kurdish in any script; the model file sizes, including 483,546,902 bytes for faster-whisper-small.
CORRECTED SINCE THE LAST VERSION, AND WORTH KNOWING BECAUSE THE OLD CLAIM WAS WRONG: Whisper's language table has 100 entries, not 99 - the earlier count predated 'yue' being added for large-v3. No Kurdish either way, so nothing that depends on it changes.
ROUTES DELETED RATHER THAN PATCHED. Intel Macs are now excluded outright. PyTorch's last Intel-Mac wheels were in version 2.2.2 in March 2024 and the current release has none, so there is no honest way to keep that path; pinning an eighteen-month-old PyTorch against current RealtimeSTT, scipy and onnxruntime is an untested combination I will not put in front of someone with two hours a week. Apple Silicon Macs on macOS 12 or 13 are excluded for the same kind of reason: the current Apple Silicon wheel requires macOS 14. The hf-mirror.com fallback has also been removed rather than caveated - recommending an unaudited proxy to human rights documenters, for files that a converter then loads, is not a trade worth making for bandwidth.
STILL NOT VERIFIED, and you should treat these as informed estimates. I have still not run this pipeline end to end on real Arabic or Kurdish audio with a real microphone. The script is syntax-checked, the multiprocessing fix is proven, and every function call comes from current official documentation, but nobody has yet spoken into it and read the result. Someone should run it once on a Windows laptop and once on an Apple Silicon Mac before Erbil; that single test is what would have caught both defects in the previous version, and it is the last thing standing between this recipe and full confidence. The 3-8 second latency figure and the per-model memory figures are calculated from model sizes and general CPU behaviour, not measured. I could not obtain a verified word-error-rate for Whisper on Modern Standard Arabic, so the MSA assessment is qualitative only. There is no published zero-shot Whisper figure for Iraqi/Mesopotamian Arabic at all, because Iraq is absent from the Casablanca benchmark and from NADI 2025 - which is why this version says 'at least as bad as Jordanian at 48%, possibly much worse' instead of a single invented number. I retrieved no FLORES chrF++ score for NLLB on ckb_Arab or acm_Arab, so the Sorani and Iraqi-Arabic translation quality claims rest on the model card's own research-only framing plus one 2026 study finding it insufficient for clinical use; the true quality on ordinary NGO prose is unmeasured. Whether huggingface.co, pypi.org and download.pytorch.org download reliably from Iraqi ISPs is unverified, though I found no evidence of any of them blocking Iraq. Meta's Omnilingual ASR is confirmed to LIST ckb_Arab and kur_Arab; its actual accuracy on Sorani is not established by that fact and I found no evaluation of it. The community Sorani ASR models on Hugging Face have download counts in the tens and low hundreds and no independent evaluation - I am not recommending them, only recording that they exist.
ONE LICENCE POINT FOR YOUR DIRECTOR RATHER THAN A FOOTNOTE: NLLB-200 is CC-BY-NC and its own card says it is not for production deployment, not for legal or medical text, and not usable as certified translation. That is compatible with documentation and journalism. It is not compatible with selling translation, and it is not compatible with submitting machine output as an official record.