لم يراجع هذه الترجمة إنسان بعد

أُنتجت بمساعدة الترجمة الآلية. اقرأ النسخة الإنجليزية قبل أن تتصرف بناءً على أي شيء فيها.

الوقت اللازم
90 دقيقة · 20 خطوة
ما الذي يتطلبه
يتطلب تثبيت برامج
التكلفة
يوجد مسار مجاني
آخر تحقق
2026-08-31
آمن للمواد الحساسة
  • التوثيق والرصد
  • صحفي أو محرر
  • باحث أو محلل
  • تفريغ الشهادات والمقابلات
  • حماية المعلومات الحساسة
  • التحقق من الصور والفيديو والادعاءات
هذه تعليمات، وليست دراسات حالة

الوصفة تخبرك بما ينبغي أن تفعله. وهي ليست تقريراً عن شيء حدث في مكان آخر. لا شيء هنا يزعم أن منظمة ما قامت بهذا؛ بل هي مكتوبة لكي تنفّذها أنت الآن، وكل خطوة تشرح لك كيف تتأكد من أنها نجحت.

تسجيل المقابلة هو غالباً أخطر ملف تحتفظ به منظمة. هذه الوصفة تفرّغه داخل حاسوبك المحمول نفسه والشبكة مفصولة، والخطوة العاشرة تجعلك تثبت أن الشبكة مفصولة فعلاً بدل أن تصدّقنا.

خط الأمان

الصوت والنص لا يغادران الجهاز أبداً. أربعة أشياء تعبر الإنترنت مرة واحدة فقط: ملف تثبيت التطبيق، وملف النموذج بحجم 1.62 غيغابايت، وإحصاءات الاستخدام المجهولة التي تُطفئها الخطوة السابعة، وفحص التحديثات.

لا تعامل النص الخام الصادر عن الآلة على أنه سجل لما قاله الشخص فعلاً. إنه مسودة تسرّع استماعك أنت. وقد وُثّق عن أنه يخترع جملاً كاملة لم يقلها أحد.

ولا تستخدم المحرك السحابي أو زر التلخيص داخل هذه التطبيقات مع مقابلة حساسة. ولا تعمل داخل مجلد يتزامن مع OneDrive أو iCloud أو Google Drive أو Dropbox. ولا تعتمد على حذف الملف لاحقاً كإجراء أمني: فالحماية الحقيقية الوحيدة هي تشفير القرص بالكامل، مفعّلاً قبل كتابة التسجيل، وفقط بينما يكون الجهاز مطفأً.

عن الكردية، بصراحة

ويسبر لا يدعم الكردية إطلاقاً. لا يوجد في قائمة لغاته أي مدخل كردي بأي صورة: لا ku ولا ckb ولا kmr. وإذا ضُبط على الكشف التلقائي فسيخمّن الفارسية أو العربية أو التركية ويعيد هراءً واثقاً. لا شيء في هذه الوصفة يساعدك على الصوت الكردي.

التفاصيل الكاملة أدناه، الخطوات والأوامر والمصادر، منشورة بالإنجليزية، لأن الأوامر وأسماء الأزرار داخل البرامج نفسها بالإنجليزية، ولأن رقم إصدار يختلف بين ثلاث ترجمات هو رقم لا يمكن الوثوق به. المقدمة والملخص الأمني في أعلى الصفحة مترجمان.

ماذا يحدث لمعلوماتك

لا تستخدم هذا أبداً من أجل

Never treat the raw machine transcript as a record of what a person actually said. It is a draft to speed up your own listening, nothing more. Do not paste it into a case file, a submission, a court annex, a report quotation or a published article until a human has listened to that specific passage and confirmed the words. Whisper is documented to invent whole sentences that were never spoken. Never use the cloud engine or 'Summarize' options inside these apps for a sensitive interview; that sends your content to a foreign company. Never work in a folder that syncs to OneDrive, iCloud, Google Drive or Dropbox. And never rely on deleting files afterwards as a security measure: deletion does not erase data from a disk, and only full-disk encryption switched on BEFORE the recording was written offers real protection, and only while the machine is powered off.

ما الذي يغادر جهازك

أي قانون ينطبق عليها · الخيار المحلي بالكامل، دون إنترنت

YOUR AUDIO AND YOUR TRANSCRIPT: nothing, once steps 2 and 7 are done. The recording is never uploaded. Transcription happens inside your laptop's own processor using a model file on your own disk, and step 10 makes you prove it with the network switched off. FOUR THINGS DO CROSS THE INTERNET, AND NONE OF THEM IS YOUR AUDIO. (1) The app installer, downloaded once from GitHub. (2) The Whisper model file, 1.62 GB, downloaded once from Hugging Face. Both are ordinary file downloads; a network observer sees that you downloaded a transcription tool, not what you transcribed. (3) ANONYMOUS USAGE ANALYTICS, ON BY DEFAULT, WHICH STEP 7 TURNS OFF. Vibe's privacy policy states it 'collects anonymous, privacy-friendly analytics through Aptabase', specifically 'Event names (e.g. transcription started, failed)', 'Error messages (no file names or content)', 'OS name, app version, and country (coarse, no IP stored)', and that 'Analytics can be disabled at any time in Settings.' That is not your audio and not your transcript, but it is a signal leaving the machine and you should switch it off. (4) Automatic update checks, the policy says updates 'come directly from GitHub releases. No other cloud storage or external servers are involved.' TWO THINGS WILL SEND YOUR CONTENT ABROAD IF YOU LET THEM, AND YOU MUST NOT. First, Vibe's 'Summarize' option in the More Options menu: the app's own description reads 'Requires API key. Once the transcription is complete, it will be sent to Claude API along with the chosen prompt', and the privacy policy confirms 'By default, this option is disabled.' Never enable it for a sensitive interview. Buzz, Subtitle Edit and MacWhisper have equivalent optional cloud engines, anything whose name contains 'API', 'Cloud', or a company name. Second, and far more dangerous because it is silent: CLOUD FOLDER SYNC. Microsoft's own documentation says Known Folder Move covers 'Windows known folders (Desktop, Documents, Pictures, Screenshots, and Camera Roll)' and describes its silent policy as being used 'to redirect and move known folders to OneDrive without any user interaction'; macOS iCloud Drive has a matching 'Desktop & Documents Folders' switch under System Settings > [your name] > iCloud > Drive. On a default-configured laptop, a witness interview saved to the Desktop is uploaded automatically, with no prompt, and the offline test in step 10 CANNOT detect it because the sync client simply uploads later when the network returns. That is why step 2 is a prerequisite and not a suggestion. One more thing worth knowing, in Vibe's own words: 'There is no encryption involved in the app.' The files it writes are ordinary files. Your disk encryption is the thing protecting them.

أي قانون ينطبق عليها

For your recording: none, provided step 2 was done. It stays on your laptop and is subject only to the physical security of that laptop and to whoever can compel access to it. For the one-time downloads: the app installer and the model file are hosted in the United States (GitHub/Microsoft, Hugging Face) and your download of them is visible to those companies and to your internet provider. The default-on usage analytics go to Aptabase; step 7 switches them off. If your organisation considers even the fact of downloading transcription software sensitive, do that download on a different network or a different device and carry the files across on a USB stick. If step 2 was NOT done and the file sat in a synced Desktop or Documents folder, the audio and the transcript are in the United States under Microsoft's or Apple's control and are subject to US legal process, treat that recording as disclosed.

الخيار المحلي بالكامل، دون إنترنت

This recipe IS the fully local option. That is the entire point of it. Vibe's README describes it as 'Ultimate privacy: fully offline transcription, no data ever leaves your device', and its privacy policy says 'After the initial setup, where the model files are downloaded, all transcription happens entirely on your local device.' After those one-time downloads, the whole workflow runs with the network cable out and Wi-Fi off, and steps 10 and 14 make you do exactly that. If you need to be fully offline from the very start, ask a colleague to download the Vibe installer and the ggml-large-v3-turbo.bin model file elsewhere and bring them on a USB stick; place the .bin in the path shown under 'Models Folder' in Vibe's Settings and the app will use it without ever reaching the internet. For the strongest version of this: keep one laptop that has BitLocker or FileVault on, has never had OneDrive or iCloud Desktop sync enabled, has Vibe's analytics switched off, and does interview transcription only.

التكلفة

الدفع من العراق

The recommended route needs no payment at all, which is deliberate: for an Iraqi organisation this removes the payment problem entirely rather than working around it. Vibe, Buzz, Subtitle Edit, whisper.cpp, faster-whisper and Audacity are downloaded from GitHub, SourceForge, Hugging Face and audacityteam.org, none of which asks for a card, an email address or an account, and none of which is known to geo-block Iraq. If you did want a paid tier, MacWhisper and Whisper Notes are sold through their vendors' own channels; we could NOT verify first-hand whether an Iraqi-issued card is accepted, Iraqi cards are frequently declined by international SaaS billing, and organisations should assume a foreign card or a colleague abroad would be needed. Treat the paid tiers as optional extras that may be unbuyable from Iraq, and plan on the free route. THE REAL CONSTRAINT HERE IS BANDWIDTH, NOT PAYMENT, and this repair improves it substantially. The previously recommended app, Buzz, ships its Windows build on GitHub as three files, Buzz-1.4.5-windows.exe (4,392,262 bytes), Buzz-1.4.5-windows-1.bin (2,095,607,552 bytes) and Buzz-1.4.5-windows-2.bin (594,799,916 bytes), about 2.7 GB in total, all three of which you need. Vibe's Windows installer is 43,592,944 bytes, roughly sixty times smaller for the same underlying whisper.cpp engine. Added to the 1.62 GB model file, the whole setup is about 1.66 GB. The model file is also a one-time office-wide asset: download it once, copy it to a USB stick, and every other laptop in the office can be configured fully offline from it. If even 1.62 GB is out of reach, ggml-large-v3-turbo-q5_0.bin is 574 MB at a small accuracy cost.

المسار المجاني

إذا دفعت

Fully free, with no account, no card, no trial and no subscription. Vibe (MIT licence, github.com/thewh1teagle/vibe) is the recommended app and is free on Windows, macOS, both Intel and Apple silicon, and Linux, with no paid tier of any kind. Buzz (MIT licence, github.com/chidiwilliams/buzz) is free on the same platforms, subject to the Intel Mac limit described in platformNotes. Subtitle Edit (github.com/SubtitleEdit/subtitleedit) is a third free option on Windows and includes a built-in Whisper engine. whisper.cpp and faster-whisper are free open-source software. Audacity is free under the GPL. The Whisper model files themselves are free, anonymous downloads from Hugging Face. The only things you spend are: about 44 MB for the Vibe installer, 1.62 GB once for the Large v3 Turbo model, and CPU time. NOTE ON WHAT CHANGED: an earlier version of this recipe recommended MacWhisper's free tier as the primary route for Apple-silicon Macs. We have removed it. MacWhisper's published pricing page compares free and Pro on export formats, speaker recognition, batch processing and dictation, but says nothing whatever about which Whisper MODELS the free tier can run, and this recipe's own step 8 declares Tiny and Base unusable for Arabic. We could not verify from the vendor's own materials that a free MacWhisper user can reach Large v3 Turbo or Medium, so we are not going to route Mac users down a path whose free tier we cannot confirm reaches a usable model. Vibe runs on Apple-silicon Macs perfectly well and its free-ness is not in question.

إذا دفعت

USD 0/month on the recommended route, and USD 0/month on every route in this recipe. There is no subscription anywhere here and nothing to buy. Two paid products are named elsewhere in this recipe only so you can recognise them and ignore them: MacWhisper (https://www.macwhisper.com/) and Whisper Notes (https://whispernotes.app/). WE ARE DELIBERATELY NOT PRINTING THEIR PRICES OR TERMS. Prices change within weeks and this page is read for a year, click the vendor pages if you are curious. Neither is needed to complete this recipe. Vibe, Buzz, Subtitle Edit, whisper.cpp, faster-whisper and Audacity have no paid tier at all; check any of those claims yourself at the repository links in sources.

قبل أن تبدأ

  • A laptop with at least 8 GB of RAM. 4 GB works but only with the 'Small' model and noticeably worse Arabic. Step 1 shows you how to check.
  • A CPU that supports AVX2 (roughly, any laptop made after about 2013, but NOT low-end Celeron/Pentium/Atom chips). If Vibe says 'Your CPU is not supported by this version of Vibe. Please use a more modern PC.', the point-and-click route is closed on that machine and you need step 19 or a different laptop.
  • At least 4 GB of free disk space for the app, one model file and your working copies.
  • An internet connection for the first downloads only: about 44 MB for the app plus 1.62 GB for the Large v3 Turbo model. On a slow or metered connection, do this once and keep the model file. It can be copied to every other laptop in the office on a USB stick.
  • THE FOLDER CHECK, WHICH IS A PREREQUISITE AND NOT A SUGGESTION. Before any recording goes on this machine, confirm that the folder you will work in is NOT synced to OneDrive, iCloud Drive, Google Drive, Dropbox or Nextcloud. Microsoft's own documentation confirms OneDrive Known Folder Move covers 'Windows known folders (Desktop, Documents, Pictures, Screenshots, and Camera Roll)' and that its silent policy is used 'to redirect and move known folders to OneDrive without any user interaction'. Apple's iCloud Drive has an equivalent 'Desktop & Documents Folders' setting. Either one uploads your witness audio and your finished transcript to a US company automatically, with no prompt, and the offline test in step 10 cannot detect it because the sync client simply uploads later when the network comes back. Step 2 walks you through this.
  • Full-disk encryption turned on BEFORE the recording ever touches the laptop: BitLocker on Windows, FileVault on macOS. Turning it on afterwards does not retroactively protect what was already written to the disk. Vibe's own privacy policy is blunt about this: 'There is no encryption involved in the app.'
  • The audio or video file itself. Phone voice memos (.m4a), WhatsApp voice notes (.opus, .ogg), .mp3, .wav and .mp4 all work, Vibe's architecture document confirms that 'macOS and Windows builds also bundle ffmpeg', so it converts them for you.
  • A separate, non-sensitive test recording you are willing to throw away. Step 3 has you make one. Every piece of setup and proof in this recipe happens on that throwaway file, and your real interview does not come onto the machine until step 12.
  • Headphones, for the listening pass in step 16. This step is not optional and it is the only reason the output can be trusted.
  • Permission to have this recording on this laptop at all. A local transcript is only as private as the machine it sits on. If the laptop is shared, unencrypted, or likely to be searched, solve that before you start.
  • No account, no email address, no credit card, no engineer.

على أي أجهزة يعمل هذا

WINDOWS FIRST, because most of this room is on Windows. The recommended route, Vibe, has a build for Windows x64, for Apple-silicon Macs, for Intel Macs and for Linux (.deb and .rpm), and the steps below are identical on all four. Sizes matter here more than usual: the Vibe Windows installer is about 44 MB, versus roughly 2.7 GB for Buzz on Windows (Buzz ships its Windows build as one small .exe plus two large .bin files). On a throttled Iraqi connection that difference is hours, which is why this repair moved the recommended app from Buzz to Vibe.

PLATFORM LIMITS, STATED BEFORE YOU INSTALL ANYTHING:

1. INTEL MACS. Buzz no longer supports them. Its README says verbatim: "Intel Macs: Buzz now requires Apple silicon. The last version to support Intel Macs is 1.4.5." If you are on an Intel Mac, use Vibe (there is a dedicated Intel build, the .dmg whose name ends _x64.dmg), do not use Buzz, and specifically do not follow any advice to "take the newest Buzz version", which will hand you a build that cannot run on your machine and will not tell you why.

2. OLD OR LOW-END CPUs. Vibe requires a CPU with AVX2. If your processor is older than roughly 2013, or is a low-end Celeron/Pentium/Atom, Vibe shows exactly this message: "Your CPU is not supported by this version of Vibe. Please use a more modern PC." If you see that string, the GUI route is closed on that machine. Your options are a different laptop, or the whisper.cpp terminal route (step 19), which compiles for whatever CPU you actually have and will run, slowly, without AVX2.

3. THE TERMINAL ROUTE IS macOS AND LINUX ONLY in this recipe. Building whisper.cpp on Windows needs Visual Studio Build Tools and is a different, longer guide; we have deliberately not included a half-working Windows version of it. Windows readers should use Vibe (steps 5–18), which covers everything the terminal route covers except the last few percent of speed.

4. LINUX. Vibe ships .deb and .rpm only, no AppImage. Debian/Ubuntu: the file ending _amd64.deb. Fedora/RHEL: the file ending -1.x86_64.rpm. (There are also arm64/aarch64 builds of each if you are on ARM Linux.) If you are on another distribution, use the whisper.cpp terminal route instead of hunting for a package.

5. THE CLOUD-SYNC CHECK IN STEP 2 IS DIFFERENT ON EVERY PLATFORM and is not optional on any of them. Windows OneDrive and macOS iCloud both back up the Desktop by default configuration in many organisations. Do not skip step 2 because it looks like housekeeping. It is the step that makes the rest of the recipe true.

الخطوات

  1. خطوة 1 / 20كل الأنظمة

    Find out how much memory (RAM) your laptop has and how much free disk space. WINDOWS: press the Windows key, type the words About your PC, press Enter, and read the line 'Installed RAM'. Then open File Explorer, click 'This PC' in the left column, and read the free space under 'Local Disk (C:)'. macOS: click the Apple logo at the top-left of the screen, click 'About This Mac', and read the 'Memory' line; for disk space click the Apple logo, then System Settings, then General, then Storage. LINUX: open a terminal and run the command below.

    انسخ هذا كما هو

    free -h && df -h /home

    تأكد من أنها نجحت

    You have written down two numbers. The RAM number must be at least 4, and 8 or more is what this recipe assumes. The free-disk number must be at least 4 GB. If free disk is under 4 GB, stop and clear space now, a half-downloaded model file fails in a way the app reports as a generic error.

    ملاحظة

    Write both numbers down on paper. The RAM number decides which model you pick in step 8, and choosing a model too big for your RAM is the single most common reason this fails. The command shown is for Linux only; Windows and macOS use the menus described in the instruction.

  2. خطوة 2 / 20كل الأنظمة

    MAKE A WORKING FOLDER THAT IS NOT SYNCED TO ANY CLOUD. Do not use the Desktop and do not use Documents. WINDOWS: open File Explorer, click 'This PC', open 'Local Disk (C:)', open 'Users', open the folder with your own name, and there create a new folder named exactly: transcribe. Your path is now C:\Users\YourName\transcribe. Then press the Windows key, type OneDrive, open it, click the gear icon, choose Settings, and find the backup section, recent clients call it 'Sync and backup' and older ones call it 'Backup'; in either case click 'Manage backup' and look at whether Desktop, Documents and Pictures are switched on. macOS: click the Apple logo, System Settings, click your own name at the top, click iCloud, click 'Drive', and look at whether 'Desktop & Documents Folders' is switched on. Then in Finder use the 'Go' menu, choose 'Home', and create a new folder there named exactly: transcribe. Your path is now ~/transcribe. LINUX: create ~/transcribe and confirm no Dropbox, Nextcloud, Insync or Google Drive client is syncing your home folder.

    انسخ هذا كما هو

    mkdir -p ~/transcribe && echo ~/transcribe

    تأكد من أنها نجحت

    WINDOWS: open your new folder in File Explorer and look at the address bar. The path must read C:\Users\YourName\transcribe and must NOT contain the word OneDrive anywhere. Switch File Explorer to Details view; the folder and everything in it must show no blue-cloud icon and no green-circle tick in the Status column. Those icons mean the file is being synced. macOS: with the folder open in Finder, press Command+Option+P to show the path bar at the bottom. It must read Macintosh HD > Users > YourName > transcribe. If it reads iCloud Drive anywhere, you are in the wrong place. If 'Desktop & Documents Folders' was switched on and you had already put a recording on the Desktop, assume that recording has already been uploaded to Apple.

    ملاحظة

    This is the most important step in the recipe. Microsoft's documentation confirms OneDrive Known Folder Move covers 'Windows known folders (Desktop, Documents, Pictures, Screenshots, and Camera Roll)', and describes its silent policy as being used 'to redirect and move known folders to OneDrive without any user interaction'. Apple's support page confirms the matching 'Desktop & Documents Folders' switch under System Settings > [your name] > iCloud > Drive. On a default-configured laptop, a witness interview saved to the Desktop is uploaded to a US company within minutes, and no test later in this recipe can detect it, a sync client just waits for the network and uploads then. Your home folder itself (C:\Users\YourName\ on Windows, ~/ on macOS) is not covered by either feature, which is why we work there. The Linux command shown also works on macOS. If you would rather not touch the OneDrive settings at all: you do not have to. Simply working inside C:\Users\YourName\transcribe is enough, because Known Folder Move only redirects the named folders.

  3. خطوة 3 / 20كل الأنظمة

    MAKE A THROWAWAY TEST RECORDING. Take a printed page or a screen of Modern Standard Arabic you can read from, a news article, a page of one of your own reports, and record yourself on your phone reading it aloud for sixty seconds. Keep the page. Do not use anything sensitive: this file is going to be used to install, configure, prove and test everything, while your real interview stays off this laptop entirely until step 12.

    تأكد من أنها نجحت

    You have a sixty-second audio file on your phone, and you still have the exact text you read from, on paper or on screen. If you improvised instead of reading from a page, record it again, the whole value of this file is that you know the correct answer.

    ملاحظة

    This is the change that makes the recipe safe. In earlier versions, the real interview was transcribed online first and the privacy proof came afterwards, which meant that if you had accidentally chosen a cloud engine, the disclosure had already happened and could not be undone. Everything risky now happens on a recording you would not mind losing. It also gives you a known-answer test, which is the only way to find out how good this actually is on your voice, your accent and your microphone, in five minutes, rather than after a three-hour interview.

  4. خطوة 4 / 20كل الأنظمة

    Copy the test recording from your phone into the transcribe folder you made in step 2, and rename the copy so its name has no spaces and no Arabic letters: name it exactly test-01.m4a (or test-01.mp3, or whatever extension it already has, change only the part before the dot).

    تأكد من أنها نجحت

    The file is inside C:\Users\YourName\transcribe (Windows) or ~/transcribe (macOS and Linux), it is named test-01 followed by its extension, and double-clicking it plays your own voice. If the file is 0 KB, the copy from the phone failed, copy it again.

    ملاحظة

    Arabic characters and spaces in filenames are the most common cause of 'file not found' and silently empty output in every tool in this recipe, and none of them tells you that is the problem. Always rename the copy and never the original.

  5. خطوة 5 / 20كل الأنظمة

    DOWNLOAD VIBE. Open your browser and go to the releases page below. Under 'Assets', download exactly one file, the one whose name ENDS with the suffix for your machine, whatever the version number in the middle happens to be. WINDOWS: the file ending _x64-setup.exe. MAC WITH APPLE SILICON (M1, M2, M3, M4): the file ending _aarch64.dmg. MAC WITH AN INTEL PROCESSOR: the file ending _x64.dmg. LINUX, DEBIAN OR UBUNTU: the file ending _amd64.deb. LINUX, FEDORA OR RHEL: the file ending -1.x86_64.rpm. Ignore every file ending in .sig, every file ending in .app.tar.gz, and the file named latest.json.

    انسخ هذا كما هو

    https://github.com/thewh1teagle/vibe/releases/latest

    تأكد من أنها نجحت

    The file you downloaded is tens of megabytes, not a few kilobytes. If it is a few hundred bytes you took the .sig signature file by mistake, delete it and download the one without .sig at the end. For scale, in release v3.1.6, which was the current release when we checked on 30 August 2026, the assets were: Windows installer 43,592,944 bytes; Apple-silicon .dmg 38,364,114; Intel .dmg 44,111,146; .deb 36,727,394; .rpm 36,721,050. Newer releases will differ a little. TAKE WHATEVER VERSION THE PAGE SHOWS, unlike Buzz, Vibe builds every platform including Intel Macs from the same release, so 'take the newest version' is safe advice here.

    ملاحظة

    Vibe is MIT-licensed, free, has no paid tier, no account and no trial. Its README describes it as 'Ultimate privacy: fully offline transcription, no data ever leaves your device'; its architecture document confirms the engine is 'Rust + whisper.cpp bindings' and that 'macOS and Windows builds also bundle ffmpeg', so it will accept your phone recording directly. IF YOU PREFER A DIFFERENT APP: Buzz (MIT, also free, github.com/chidiwilliams/buzz) does the same job, but its Windows download is one small .exe plus two large .bin files totalling about 2.7 GB, and on Intel Macs you must take exactly the 1.4.5 mac-X64 build and never a higher version, because 1.4.5 is the last release that supports Intel at all. Subtitle Edit (github.com/SubtitleEdit/subtitleedit) is a good free Windows alternative if you already work with subtitles. We have removed MacWhisper from this recipe, see honestLimitations for why.

  6. خطوة 6 / 20كل الأنظمة

    INSTALL AND OPEN IT. WINDOWS: double-click the _x64-setup.exe file you downloaded and follow the installer. If a blue box appears saying 'Windows protected your PC', click the small words 'More info' and then the button 'Run anyway'. macOS: double-click the .dmg and drag the Vibe icon onto the Applications folder shown beside it. If macOS then refuses to open it and says the developer cannot be verified, click the Apple logo, System Settings, Privacy & Security, scroll down to the message about Vibe, and click 'Open Anyway'. LINUX (Debian/Ubuntu): open a terminal, change into the folder where the .deb downloaded, and run the command below.

    انسخ هذا كما هو

    sudo apt install ./vibe_*_amd64.deb

    تأكد من أنها نجحت

    Vibe opens and shows a window containing the words 'Drop audio, video or a folder here'. If instead you see 'Your CPU is not supported by this version of Vibe. Please use a more modern PC.', this machine lacks AVX2 and the point-and-click route is closed on it, go to step 19, or use a different laptop.

    ملاحظة

    The security warning is normal for small independent open-source software that has not paid for a commercial code-signing certificate. It is not a virus warning. You should only click through it because you downloaded from the official releases page in step 5. Never do this for software that arrived by email, WhatsApp, or a link someone sent you. The apt command uses a wildcard so it works whatever version you downloaded; it assumes you are in the folder where the .deb landed. On Fedora use: sudo dnf install ./vibe-*.x86_64.rpm

  7. خطوة 7 / 20كل الأنظمة

    TURN OFF THE THINGS THAT PHONE HOME, AND DECIDE WHAT THE APP IS ALLOWED TO KEEP. Open Vibe's Settings (the gear icon). Find the toggle labelled 'Anonymous analytics' and switch it OFF. Ignore the message beside it that says 'We strongly recommend keeping this enabled so we can find and fix issues faster.' Still in Settings, find 'Save projects', whose description reads 'Automatically save transcripts and media in the projects folder.' Switch that OFF. Also in Settings, find the lines 'Projects folder', 'Models Folder' and 'Temporary files', and write those three paths down on paper. You need them in steps 8, 9 and 18. Finally: never touch the 'Summarize' option in the 'More Options' menu.

    تأكد من أنها نجحت

    The 'Anonymous analytics' switch is visibly off, and the 'Save projects' switch is visibly off. Both must read off before you go any further, and you must have the three folder paths written down. If you cannot find these settings, update Vibe to the current release.

    ملاحظة

    Vibe's privacy policy states that by default it 'collects anonymous, privacy-friendly analytics through Aptabase', specifically 'Event names (e.g. transcription started, failed)', 'Error messages (no file names or content)' and 'OS name, app version, and country (coarse, no IP stored)', and that 'Analytics can be disabled at any time in Settings.' That is not your audio and not your transcript, but it is a signal leaving the machine, and for an organisation that considers even the fact of using transcription software sensitive it matters. 'Save projects' matters for a different reason: Vibe's own interface text says 'Vibe stores each transcript and its media together in this folder' and describes 'Saved projects' as 'Every transcription Vibe saved, with its copy of the audio.' With it on, deleting your own copy of the interview leaves Vibe holding a full duplicate, audio and text, in a folder you did not choose. 'Summarize' is worse: Vibe's own description says 'Requires API key. Once the transcription is complete, it will be sent to Claude API along with the chosen prompt'. It is off by default and does nothing without a key, and you must never turn it on for a sensitive interview.

  8. خطوة 8 / 20كل الأنظمة

    GET A MODEL, THEN SET THE TWO SETTINGS THAT MATTER. Drag test-01.m4a onto the Vibe window (or click 'Select File' and choose it). Then, before pressing anything: (1) GET A MODEL. Vibe does not ship with a menu of named models. It has a models folder and a link box. In Settings look for 'Download model' and 'Paste Model Link', paste the direct link for the model you want, and wait for the download to finish; then use 'Select Model' to pick it. IF YOUR LAPTOP HAS 8 GB OF RAM OR MORE, use Large v3 Turbo, the link is in the command field below. IF YOUR LAPTOP HAS ONLY 4 GB, use Small instead: https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.bin?download=true, Do not use Tiny or Base. (2) Find 'Select Language' and choose Arabic explicitly, do not leave it on auto-detect. (3) Find 'Translate To English' and make sure it is switched OFF.

    انسخ هذا كما هو

    https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3-turbo.bin

    تأكد من أنها نجحت

    Before you press Transcribe, read the screen back to yourself and confirm all three: a model is selected and its name contains large-v3-turbo (or small on a 4 GB machine), the language shows Arabic, and 'Translate To English' is off. If Vibe says 'No models installed' or 'Select a model in Settings to get started', the model download did not finish, do that first. If any one of these three is wrong you will get a page of confident, fluent, useless text and it will not look like an error.

    ملاحظة

    Vibe's own model page (https://github.com/thewh1teagle/vibe/blob/main/docs/models.md) is where those links come from; it labels the turbo model '(Recommended)' and also offers one-click 'Magic Setup' links that open a model directly in Vibe. Vibe's Settings text explains the manual alternative: 'Download any supported model in ggml or gguf format (with the file extension ending in .bin). Transfer it to the models directory, then select from the dropdown. No need to restart', so you can equally drop a .bin file into the 'Models Folder' path you wrote down in step 7. FILE SIZES, read from Hugging Face on 30 August 2026: tiny 77.7 MB, base 148 MB, small 488 MB, medium 1.53 GB, large-v3 3.1 GB, large-v3-turbo 1.62 GB, large-v3-turbo-q5_0 574 MB. MEMORY, from whisper.cpp's own table: tiny ~273 MB, base ~388 MB, small ~852 MB, medium ~2.1 GB, large ~3.9 GB. That table has no row for turbo, so we cannot quote you a figure for it; in practice assume it needs at least as much as medium. Tiny and Base produce gibberish in Arabic, do not use them. Small is the realistic minimum, and expect heavy correction. Auto-detect on Arabic dialect audio regularly guesses Persian, Urdu or Turkish and then returns fluent nonsense; setting the language by hand is free accuracy. 'Translate To English' does exactly what Vibe's own help text says, 'Translate transcription into English from any language by enabling this option', so if your transcript comes back in English, this is why. If Vibe crashes or your GPU driver misbehaves, switch on 'Transcribe on the CPU', described in the app as 'Runs transcription without the GPU. Slower, but it avoids driver crashes and frees the GPU for other work.'

  9. خطوة 9 / 20كل الأنظمة

    PRESS TRANSCRIBE ON THE THROWAWAY FILE. If the model is still downloading, wait for it to finish. That is the 1.62 GB download for Large v3 Turbo, and it needs internet. Then let the transcription run. Do not close the lid or let the laptop sleep while it works.

    تأكد من أنها نجحت

    Arabic script appears on screen and it recognisably resembles the page you read from. THREE WAYS THIS FAILS QUIETLY, all of which look like success: (a) the transcript is empty or a single blank line, almost always the filename, go back to step 4; (b) the text is in Latin letters, or in Persian or Turkish script, the language was not set, go back to step 8; (c) the text is in English, 'Translate To English' was left on, go back to step 8. Do not continue until you have Arabic script in front of you.

    ملاحظة

    If the model download stalls or crawls on a throttled connection, you can fetch the file by hand instead: go to https://huggingface.co/ggerganov/whisper.cpp/tree/main in a browser that can resume downloads, take ggml-large-v3-turbo.bin (1.62 GB), and put it in the 'Models Folder' path you wrote down in step 7. Or have a colleague abroad download it and bring it on a USB stick, the app will use a model file you place there without ever reaching the internet. The model file is a permanent office asset: copy it once and every other laptop can be set up fully offline from it. If you are very short of disk space, ggml-large-v3-turbo-q5_0.bin is 574 MB instead of 1.62 GB at a small accuracy cost.

  10. خطوة 10 / 20كل الأنظمة

    NOW PROVE NOTHING IS LEAVING THE ROOM, STILL ON THE THROWAWAY FILE, BEFORE YOUR INTERVIEW IS ANYWHERE NEAR THIS MACHINE. Disconnect completely. WINDOWS: press the Windows key, type Airplane mode, press Enter, and switch it on. macOS: click the Wi-Fi icon in the menu bar and turn Wi-Fi off. LINUX: turn off networking from the system tray. Unplug any network cable on every platform. Then transcribe test-01.m4a a second time, with the network down.

    تأكد من أنها نجحت

    The second transcription runs to completion and produces Arabic text with the network off. That proves the engine is doing the work on your own processor. If it fails, hangs, or reports a network error, you have selected a cloud engine somewhere, go back to step 8, fix it, and repeat this test. Because you did this on a throwaway recording, a failure here costs you nothing.

    ملاحظة

    Be precise about what this test does and does not prove. It proves the engine can run entirely locally. It does not prove that the earlier online run in step 9 transmitted nothing, and it cannot detect folder sync at all, a OneDrive or iCloud client simply waits and uploads when the network returns, which is why step 2 exists and is separate. Run this once per machine, then write the result into your organisation's AI policy: 'transcription verified to run with networking disabled on [machine], [date]'.

  11. خطوة 11 / 20كل الأنظمة

    SCORE THE KNOWN-ANSWER TEST. Put the page you read from in step 3 beside the transcript on screen. Count the first hundred words of the transcript and count how many of them are wrong, wrong word, missing word, or invented word. Write that number down. It is your honest error rate on clear read-aloud Modern Standard Arabic on this machine, with this microphone, in this voice.

    تأكد من أنها نجحت

    You have a number out of 100. Interpretation: this is the BEST this setup will ever do for you, because read-aloud MSA from a page is the easiest possible input. Real interview audio will be worse, and dialect will be very much worse. If the number is above about 45 out of 100 on clear read-aloud MSA, something is misconfigured rather than merely difficult, check that the model really is Large v3 Turbo and not Small or Tiny, and that the language is set to Arabic.

    ملاحظة

    We give 45 as a tripwire rather than a promise, because the only published anchor we can verify is 45.93% word error rate for the much smaller whisper-small model on MSA before fine-tuning (Ozyilmaz, Coler and Valdenegro-Toro, Interspeech 2025). We could find no published figure for large-v3-turbo on MSA, so we are not going to invent one. Large v3 Turbo should be clearly better than that anchor on clean read-aloud speech; if it is not, your settings are wrong. Whatever number you measure here IS your number, trust it over anything anyone tells you, including us.

  12. خطوة 12 / 20كل الأنظمة

    ONLY NOW BRING THE REAL INTERVIEW ONTO THE LAPTOP. Copy the recording into the same transcribe folder from step 2, C:\Users\YourName\transcribe on Windows, ~/transcribe on macOS and Linux. Rename the copy to exactly interview-01 plus its existing extension, for example interview-01.m4a. Leave your master recording untouched wherever it already lives; you will need it for the listening pass and you must never be editing your only copy.

    تأكد من أنها نجحت

    Double-click interview-01.m4a and confirm it plays and is the right recording and the right length. Confirm again that the folder path contains no cloud service name and shows no sync icons, repeat the step 2 check now that a sensitive file is in it, because a sync client that was switched on since then would begin uploading immediately.

    ملاحظة

    Everything before this point happened on a file you did not care about. That was deliberate. From here on, treat this folder the way you would treat a paper file of the same interview.

  13. خطوة 13 / 20كل الأنظمة

    OPTIONAL BUT WORTH IT: CUT THE LONG SILENCES BEFORE YOU TRANSCRIBE. Install the free open-source editor Audacity from the address below. Open your working copy interview-01.m4a in it, find any stretch longer than about fifteen seconds where nobody is speaking, select it, and delete it, or simply split the recording into the sections where someone is talking. Save the result as a new file, interview-01-trimmed.wav, in the same folder. Never edit your master copy, only this working copy.

    انسخ هذا كما هو

    https://www.audacityteam.org/download/

    تأكد من أنها نجحت

    The trimmed file plays, is shorter than the original, and still contains everything anyone actually said, play the joins to be sure you have not cut off the start or end of a sentence. Write down, on paper, which stretches you removed and where, so that your timestamps and your chain of custody stay honest.

    ملاحظة

    Silence is where Whisper invents sentences. Removing it removes most of the opportunity, and this is the only way to do that without a terminal. Skip this step if your recording has little dead air, or if you are short of time, steps 16 and 17 will still catch the problem, this step just makes there be less of it to catch. Audacity is free, GPL-licensed, and available for Windows, macOS and Linux; check the download page for which builds it currently ships. Two warnings: Audacity keeps unsaved-project temporary files in a directory of its own, which you will clean up in step 18; and if you do trim, use interview-01-trimmed.wav as the file you feed to Vibe in the next step instead of interview-01.m4a.

  14. خطوة 14 / 20كل الأنظمة

    TRANSCRIBE THE REAL INTERVIEW, WITH THE NETWORK OFF FROM THE START. Turn Airplane mode on (Windows) or Wi-Fi off (macOS) before you load the file. Then drag interview-01.m4a, or interview-01-trimmed.wav if you did step 13, onto the Vibe window, confirm the three settings from step 8 are still model = large-v3-turbo, language = Arabic, Translate To English = off, and press Transcribe.

    تأكد من أنها نجحت

    The transcript is in Arabic script, and its length is roughly proportional to the length of the recording, a one-hour interview should not produce three lines. If the output is a single short paragraph for an hour of audio, or the same sentence repeated down the page, stop: that is the failure described in step 17 and in the failure modes, not a finished transcript.

    ملاحظة

    There is no reason to be online for this, so do not be. TIME, HONESTLY: we have no benchmark for this model on this class of laptop that we can stand behind, so we are not giving you a figure dressed up as one. Plan for somewhere between one and five times the length of the recording, a one-hour interview may take one to five hours on a CPU-only machine. Start it and go to lunch. Do not start a three-hour recording at 5pm expecting it before you go home. TIME A TEN-MINUTE FILE FIRST and multiply: that measurement on your own laptop is worth more than any published table. Split anything over an hour into 20–30 minute pieces so you get partial results you can stop between.

  15. خطوة 15 / 20كل الأنظمة

    EXPORT THE RESULT TWICE: once as a subtitle file and once as plain text. In Vibe click 'Save Transcript', choose Format, and pick 'SRT'. Save it into your transcribe folder. Then do it again and pick the plain-text format. Save both into C:\Users\YourName\transcribe (Windows) or ~/transcribe (macOS and Linux), not the Desktop, not Documents.

    تأكد من أنها نجحت

    OPEN THE .SRT FILE IN NOTEPAD (Windows) OR TEXTEDIT (macOS) AND LOOK AT IT. Good output looks like this: a line containing just the number 1, then a line reading something like 00:00:04,120 --> 00:00:09,880, then a line of Arabic text, then a blank line, then 2, and so on. If the file is 0 bytes, if it contains numbers and timestamps but no Arabic at all, or if it is one enormous block with no timestamps, the export failed and you must redo it, do not carry on and discover this in step 17.

    ملاحظة

    The .srt is the important one and most people skip it. It carries a timestamp against every line, so during the listening pass you can jump straight to the second of audio a sentence claims to come from. Without timestamps you cannot check a transcript efficiently, and a transcript you cannot check is not evidence. Vibe offers several other export formats as well; you do not need them for this.

  16. خطوة 16 / 20كل الأنظمة

    DO THE LISTENING PASS. Open the .srt beside the original master recording, put your headphones on, and listen to the whole recording while reading along. Correct as you go. Mark clearly, in your own file, every passage you have verified by ear and every passage you have not.

    تأكد من أنها نجحت

    Every passage in the document is now marked either verified-by-ear or not-verified. There is no partial credit here: an unmarked document is an unverified document, and whoever reads it next will assume it is a record.

    ملاحظة

    This is the step that turns a machine draft into a record, and there is no shortcut around it. The machine saves you the typing, not the listening. For a documenter the rule is absolute: no sentence goes into a case file, a submission, a report or a published quotation unless a human has heard it. If time is short, do a full listening pass on the passages you intend to quote or rely on, and stamp the rest of the document 'machine draft, unverified' so that nobody downstream mistakes it for a record. Budget the full length of the recording for this step, on top of the transcription time.

  17. خطوة 17 / 20كل الأنظمة

    DO THE HALLUCINATION SWEEP, IN ADDITION TO LISTENING. Search your transcript for any sentence that appears three or more times in a row. Then open the .srt and look for any single subtitle line whose two timestamps are more than about ten seconds apart, and play that exact stretch of the audio.

    تأكد من أنها نجحت

    You have played the audio behind every repeated block and every over-long subtitle line, and either confirmed the words or deleted them. If you found none of either, say so explicitly in your notes, 'swept for repeated lines and over-long segments, none found', so the next person knows the check was done rather than skipped.

    ملاحظة

    Whisper does not only mis-hear words: it invents whole sentences that nobody said, most often over silence, background noise and long pauses. Koenecke, Choi, Mei, Schellmann and Sloane ('Careless Whisper: Speech-to-Text Hallucination Harms', ACM FAccT 2024, arXiv:2402.08021) ran 13,140 audio segments and confirmed that 187 of them reliably produced hallucinations, about 1.4%, and found that '38% of hallucinations include' explicit harms: invented violence, invented racial references, invented medications, and statements phrased with false authority. They found hallucination concentrated among speakers 'who speak with longer shares of non-vocal durations', which is exactly what an interview with a traumatised or hesitant witness sounds like. Long silences and repeated lines are therefore the two highest-yield places to look.

  18. خطوة 18 / 20كل الأنظمة

    CLEAN UP THE COPIES YOU DID NOT KNOW YOU HAD MADE. In Vibe's Settings, open 'Saved projects', Vibe's own description is 'Every transcription Vibe saved, with its copy of the audio', and if anything is listed there, click 'Delete all projects' and confirm. Then open the 'Temporary files' folder whose path you wrote down in step 7 and empty it. If you used Audacity in step 13, open Audacity, go to Preferences and then Directories, find the temporary-files path, close Audacity, and empty that folder too. Finally delete your working copy interview-01.m4a, any interview-01-trimmed.wav, and any exports you do not need to keep.

    تأكد من أنها نجحت

    Search the whole disk for the filename and confirm only the copies you meant to keep come back. WINDOWS: open File Explorer, click 'This PC', type interview-01 in the search box at the top right, and wait for the search to finish. macOS: press Command+Space, or in Finder press Command+F, click 'This Mac', and search for interview-01. If Vibe's projects folder, Vibe's temp folder or Audacity's temp folder appear in those results, they were not emptied, go back and empty them.

    ملاحظة

    Turning off 'Save projects' in step 7 prevents most of this; this step removes what was created before you did, and whatever Audacity left behind. Vibe's own confirmation dialogue for 'Delete all projects' reads 'Their transcripts and copies of the audio will be removed. This cannot be undone.' BE HONEST WITH YOURSELF ABOUT WHAT DELETING ACHIEVES: sending a file to the recycle bin, or even emptying the bin, does not erase it from the disk, and a forensic tool can often recover it. The only thing that genuinely protects a seized laptop is full-disk encryption that was already switched on before the file was ever written, BitLocker on Windows, FileVault on macOS, and even that protects only a powered-off machine, not one that is running and unlocked. Deleting afterwards is good hygiene. It is not a security control.

  19. خطوة 19 / 20ماك

    OPTIONAL, TERMINAL ROUTE, macOS AND LINUX ONLY, SKIP THIS UNLESS YOU HAVE A TECHNICAL VOLUNTEER. This is not available on Windows in this recipe; Windows readers should stop at step 18. whisper.cpp is the fastest and lightest way to run Whisper on a CPU-only laptop, it will run on CPUs without AVX2 that Vibe refuses, and it has a built-in silence filter. Open the Terminal application and type these lines one at a time, pressing Enter after each and waiting for it to finish before typing the next. Lines starting with # are comments, do not type them.

    انسخ هذا كما هو

    # macOS only: install the build tools first (skip the first line if you already have Homebrew)
    /bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/Homebrew/install/HEAD/install.sh)"
    brew install cmake ffmpeg git
    
    # Linux (Debian/Ubuntu) instead of the two lines above:
    # sudo apt update && sudo apt install -y git cmake build-essential ffmpeg
    
    # From here the commands are the same on macOS and Linux:
    git clone https://github.com/ggml-org/whisper.cpp.git
    cd whisper.cpp
    cmake -B build
    cmake --build build -j --config Release
    sh ./models/download-ggml-model.sh large-v3-turbo
    sh ./models/download-vad-model.sh silero-v6.2.0
    ffmpeg -i ~/transcribe/interview-01.m4a -ar 16000 -ac 1 -c:a pcm_s16le ~/transcribe/interview-01.wav
    ./build/bin/whisper-cli -m models/ggml-large-v3-turbo.bin --vad --vad-model models/ggml-silero-v6.2.0.bin -f ~/transcribe/interview-01.wav -l ar -otxt -osrt

    تأكد من أنها نجحت

    Check in three places, because each one can fail while printing something that looks like progress. (1) After the ffmpeg line, run: ls -l ~/transcribe/interview-01.wav, the size must be roughly 1.9 MB per minute of audio and must not be 0. A 44-byte .wav is an empty file with only a header, and whisper-cli will happily transcribe it into nothing. (2) After the build, run: ls -l ./build/bin/whisper-cli, the file must exist. (3) After the last line, run: wc -c ~/transcribe/interview-01.wav.txt && head -20 ~/transcribe/interview-01.wav.srt, the byte count must not be 0, and the head must show numbered blocks with timestamps and Arabic text. Note the output filenames: whisper-cli appends to the input name, so you get interview-01.wav.txt and interview-01.wav.srt, not interview-01.txt.

    ملاحظة

    The -l ar flag sets Arabic explicitly; do not omit it. The ffmpeg line is required because, in the whisper.cpp README's own words, 'the whisper-cli example currently runs only with 16-bit WAV files, so make sure to convert your input before running the tool'. That command is the README's own conversion line with your paths substituted. ffmpeg is not present on a clean macOS or Linux machine, which is why the install line is here rather than assumed. The two --vad lines matter: Silero VAD detects and skips non-speech, which is precisely where Whisper invents sentences, so this is the terminal equivalent of the Audacity trimming in step 13. The download script writes the VAD file as models/ggml-silero-v6.2.0.bin, which is exactly the path passed to --vad-model. If you are very short of disk space, download ggml-large-v3-turbo-q5_0.bin (574 MB against 1.62 GB) instead, at a small accuracy cost, and change the -m path to match. Everything here runs offline after the model downloads. This route does NOT replace steps 16 and 17.

  20. خطوة 20 / 20كل الأنظمة

    OPTIONAL, TERMINAL ROUTE, SILENCE FILTERING WITH A SCRIPT, again, skip unless you have a technical volunteer. faster-whisper is the other main local engine and it has a built-in voice-activity filter that skips silence automatically. It works on Windows too, but only if Python is already installed from python.org, which is a longer detour than the Vibe route is worth for most readers. Install it, then save the Python code below as transcribe.py and run it with: python3 transcribe.py

    انسخ هذا كما هو

    python3 -m pip install faster-whisper
    
    # --- save everything below as transcribe.py, then run: python3 transcribe.py ---
    # Replace YOURNAME with your own username in both paths before running.
    from faster_whisper import WhisperModel
    
    AUDIO = "/Users/YOURNAME/transcribe/interview-01.m4a"
    OUT   = "/Users/YOURNAME/transcribe/interview-01-fw.txt"
    
    model = WhisperModel("large-v3-turbo", device="cpu", compute_type="int8")
    segments, info = model.transcribe(AUDIO, language="ar", vad_filter=True, beam_size=5)
    
    with open(OUT, "w", encoding="utf-8") as f:
        for s in segments:
            line = "[%.2fs -> %.2fs] %s" % (s.start, s.end, s.text)
            print(line)
            f.write(line + "\n")
    
    print("Wrote:", OUT)

    تأكد من أنها نجحت

    You must see Arabic lines scrolling past on screen while it runs, and the last line must read 'Wrote: /Users/YOURNAME/transcribe/interview-01-fw.txt'. Then check the file is not empty: wc -c /Users/YOURNAME/transcribe/interview-01-fw.txt, if that reports 0, the transcription produced nothing and the script still exited cleanly, which is exactly the silent-empty-success failure to watch for. Open the file and confirm it contains Arabic script and timestamps in the form [12.40s -> 18.90s].

    ملاحظة

    vad_filter=True is the setting that matters: it detects and skips non-speech, which is where Whisper invents sentences. compute_type="int8" is what makes this fit on a CPU-only 8 GB laptop. It uses substantially less memory than the default; the project publishes its own CPU benchmark table in its README at https://github.com/SYSTRAN/faster-whisper if you want figures, and we would rather you read them there than trust a number copied into this page. The model name "large-v3-turbo" is a valid key in faster-whisper's own model map, so the string above works as written. faster-whisper was version 1.2.1, MIT-licensed, and required Python 3.9 or newer when we checked on 30 August 2026, confirm the current state at https://pypi.org/project/faster-whisper/. On first run it downloads the model from Hugging Face, so that run needs internet; after that it is offline. Note that vad_filter reduces hallucination, it does not eliminate it, and it does not replace the listening pass in step 16.

كيف تعرف أن العمل كله قد نجح

Four checks, in this order, and do all four the first time. The order is the point: the first three all happen on a throwaway recording, before your real interview is ever on the laptop. (1) THE FOLDER CHECK, from step 2, done BEFORE anything else. On Windows, open your working folder and confirm the address bar reads C:\Users\YourName\transcribe with no mention of OneDrive, and that File Explorer's Status column shows no blue-cloud icon and no green tick beside any file. On macOS, press Command+Option+P in Finder and confirm the path bar reads Macintosh HD > Users > YourName > transcribe and not iCloud Drive. This is the check that most often makes the difference between a private transcript and a copy of a witness interview sitting on a US server. Nothing later in this recipe can detect a sync client, because a sync client just waits for the network. (2) THE KNOWN-ANSWER TEST, from steps 3 and 11. Record yourself for sixty seconds reading a paragraph of Arabic you have in front of you. Transcribe it with the exact settings you plan to use. Compare word by word against the page. Count the errors in the first hundred words. That number is your honest best case on your voice, your accent and your microphone, and you get it in five minutes instead of after a three-hour interview. As a tripwire rather than a promise: if more than about 45 out of 100 words are wrong on clear read-aloud MSA, your settings are wrong rather than the task being hard, check that the model really is Large v3 Turbo and that the language is set to Arabic. (3) THE OFFLINE TEST, from step 10, run on that same throwaway file. Switch on Airplane mode or turn Wi-Fi off completely and transcribe again. If it completes, the engine is running on your own processor. If it fails or hangs, you have a cloud engine selected, and because you did this on a file you do not care about, fixing it costs you nothing. Be precise about what this proves: it proves the engine CAN run locally; it does not prove the earlier online run transmitted nothing, and it cannot see folder sync at all. Do it once per machine and write the result into your organisation's AI policy. (4) THE SPOT-CHECK ON THE REAL INTERVIEW, after step 15. Open the .srt, pick a two-minute stretch from the MIDDLE of the recording, never the beginning, which is always the cleanest part, and listen to it while reading the matching lines. Count the errors. That count is your honest error rate for this recording, and it tells you whether this transcript can be quoted from at all or only used to find your way around the audio. Then search the whole transcript for any line repeated three or more times, and play the audio behind every subtitle line whose timestamps span more than ten seconds; those are the two places fabricated sentences appear. FINALLY, THE CLEAN-UP CHECK from step 18: search the whole disk for the filename interview-01 and confirm only the copies you meant to keep come back. If Vibe's projects folder, Vibe's temporary folder or Audacity's temporary folder appear in the results, you have duplicate copies of a witness interview you did not know about.

ما الذي قد يفشل، وماذا تفعل حينها

  • The transcript contains the same sentence repeated over and over, sometimes filling a whole page. This is Whisper looping, almost always over silence, music, or background noise. FIX: find that timestamp in the .srt, listen to the audio there, delete the invented block, and cut the silence out of the audio with Audacity (step 13) or use the --vad flags in the whisper.cpp route (step 19) before re-running. Do not simply delete the repetition and assume the lines either side of it are fine, check those too.
  • A short, fluent, entirely plausible sentence appears that nobody ever said, often about violence, medication, or attributed to some authority. This is the documented hallucination failure and it is the dangerous one precisely because it does not look like an error. FIX: there is no automatic fix. Only the human listening pass in step 16 catches it. This is why no sentence may be quoted before it has been heard.
  • THE SILENT ONE: your witness interview has been uploaded to Microsoft or Apple and nothing on screen ever told you. You put the file on the Desktop or in Documents on a laptop where OneDrive Known Folder Move or iCloud 'Desktop & Documents Folders' is switched on. FIX: this cannot be undone, treat that recording as disclosed and tell whoever needs to know. PREVENT: work only in C:\Users\YourName\transcribe or ~/transcribe, and run the step 2 check before any sensitive file exists on the machine. The offline test in step 10 will NOT catch this, because a sync client simply uploads later when the network comes back.
  • You deleted your copy of the interview and the transcript, and the app still has both. Vibe's 'Save projects' setting stores, in its own words, 'Every transcription Vibe saved, with its copy of the audio', audio and text together, in a folder you did not choose. Buzz and MacWhisper keep equivalent transcription histories. FIX: Vibe Settings, 'Saved projects', 'Delete all projects'. PREVENT: switch 'Save projects' off in step 7 before you transcribe anything real.
  • Vibe refuses to start and says 'Your CPU is not supported by this version of Vibe. Please use a more modern PC.' Your processor lacks AVX2, typical of pre-2013 machines and of low-end Celeron, Pentium and Atom chips. FIX: there is no setting for this. Use a different laptop, or use the whisper.cpp terminal route in step 19, which compiles for the CPU you actually have and will run, slowly, without AVX2.
  • ON AN INTEL MAC: you downloaded the newest Buzz and it will not run, with no useful error. Buzz's README states 'Intel Macs: Buzz now requires Apple silicon. The last version to support Intel Macs is 1.4.5.' FIX: use Vibe instead, the Vibe .dmg ending _x64.dmg is built for Intel Macs and is the recommended route in this recipe. If you specifically want Buzz on an Intel Mac, you must take exactly the 1.4.5 mac-X64 build and never a higher version number.
  • Vibe says 'No models installed' or 'Select a model in Settings to get started' and nothing happens when you press Transcribe. Vibe ships with no model. You have to fetch one. FIX: Settings, 'Download model' / 'Paste Model Link', paste https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3-turbo.bin , wait for it to complete, then 'Select Model'. Or drop a .bin file into the 'Models Folder' path by hand, Vibe's own text says 'No need to restart'.
  • The transcript comes back in English instead of Arabic. Vibe's 'Translate To English' switch was left on, or in another app the Task setting was left on 'Translate'. FIX: switch it off and re-run. Do not use a machine English translation of a witness statement for anything. It is an unreviewed translation of an unreviewed transcription, two layers of error stacked.
  • The transcript comes back in Latin letters, or in Persian or Turkish script, or is fluent nonsense. The language was left on auto-detect and Whisper guessed wrong, very common with Iraqi dialect. FIX: set 'Select Language' explicitly to Arabic and re-run. If the speaker is actually speaking Kurdish, no setting will fix it: Whisper has no Kurdish language at all.
  • The app freezes, crashes, or the laptop becomes unusable partway through. The model is too large for your RAM and the machine is swapping to disk. FIX: close everything else, then drop from Large v3 to Large v3 Turbo, or from Turbo to Medium, or from Medium to Small. In Vibe, 'Transcribe on the CPU' also helps if the crash is a GPU driver problem, the app describes it as 'Runs transcription without the GPU. Slower, but it avoids driver crashes'. On the terminal route, use a quantised model file (the ones ending -q5_0.bin), which uses substantially less memory.
  • Nothing happens when you drop the file in, or the output is empty while the app reports success. Almost always the filename. FIX: rename the working copy to plain Latin letters with no spaces (interview-01.m4a) and try again. If it still fails, the audio format is unusual, convert it with the ffmpeg command in step 19, or re-export it as a .wav from any audio player. Always open the exported file and confirm it is not empty; step 15 tells you exactly what good .srt output looks like.
  • The model download never completes, or crawls. Common on constrained or throttled connections. FIX: download ggml-large-v3-turbo.bin by hand from https://huggingface.co/ggerganov/whisper.cpp/tree/main using a browser that can resume, and place it in the path shown under 'Models Folder' in Vibe's Settings; or have a colleague download it and bring it on a USB stick. It is a one-time cost and the file can be copied to every other laptop in the office. If bandwidth is the hard limit, take ggml-large-v3-turbo-q5_0.bin at 574 MB instead of 1.62 GB.
  • On Windows, Buzz starts and immediately dies with no clear message. Buzz's GitHub release ships the Windows build as one small .exe plus two large .bin files (Buzz-1.4.5-windows.exe, Buzz-1.4.5-windows-1.bin and Buzz-1.4.5-windows-2.bin) and you took only the .exe. FIX: download all three and put them in the same folder before running the .exe. Note that other download mirrors package Buzz differently, if you got a single self-contained archive somewhere else and are hunting for .bin files, they do not exist there and nothing is wrong. This confusion is one reason the recipe now recommends Vibe.
  • The speaker labels are wrong. It merges two people or invents a third. Speaker identification in these tools is approximate and degrades badly on overlapping speech and phone-quality audio. FIX: do not rely on it. Assign speakers yourself during the listening pass. For a two-person interview, record on two devices if you can, or have the interviewer state the speaker's name at each turn.
  • The first minutes are good and then quality collapses. Usually the speaker relaxed out of MSA into dialect. FIX: nothing technical helps much. Accept that the dialect portions are a navigation aid only, and budget more listening time for them.
  • It is taking far longer than you expected and you need the laptop. FIX: this is normal, expect one to five times the length of the recording on a CPU-only machine. Time a ten-minute file first so you know your own multiplier. Split long recordings into 20–30 minute pieces so you get partial results and can stop between them, and run anything over an hour overnight.
  • On the whisper.cpp terminal route, the transcript file is created but empty, and no error was printed. Almost always the ffmpeg conversion produced a header-only WAV. FIX: run ls -l ~/transcribe/interview-01.wav, a valid file is roughly 1.9 MB per minute of audio; 44 bytes means it is empty. Re-run the ffmpeg line and read its output for the real error, which is usually a wrong path or an unreadable source file.

جودة عمله بلغتك

العربية الفصحى · اللهجة العراقية · الكردية السورانية

العربية الفصحى

Works, and is genuinely useful, but is not clean. Expect to correct the text rather than read it. Modern Standard Arabic is one of the 100 languages listed in Whisper's tokenizer and is well represented in its training data. A concrete published anchor: Ozyilmaz, Coler and Valdenegro-Toro (Interspeech 2025) report 45.93% word error rate for the un-fine-tuned whisper-small model on MSA, the paper's exact sentence is 'Figure 1 showed a 45.93% WER on the pre-trained model before fine-tuning', roughly one word in two wrong at that model size. The large-v3 and large-v3-turbo models are substantially better than small, which is why this recipe tells you to use turbo rather than small whenever your laptop can carry it; we have no published turbo-on-MSA figure and are not going to invent one, which is precisely why step 11 makes you measure your own. Numerals, proper names, place names and organisation names are the weakest points; check every one by ear. Whisper also outputs unvocalised Arabic script and its punctuation is guesswork, so sentence boundaries, which matter a great deal when you are quoting someone, are not reliable.

اللهجة العراقية

MUCH WORSE THAN MSA, AND WORSE THAN EARLIER VERSIONS OF THIS RECIPE ADMITTED. An earlier draft told you to expect roughly one word in three to one word in two wrong on heavy Iraqi dialect. The honest range is closer to HALF TO FOUR-FIFTHS OF WORDS WRONG, and here is the evidence, all of which we opened and read directly in this session. There is still no published disaggregated zero-shot word error rate for Iraqi/Mesopotamian Arabic specifically. That gap is itself a finding and we are stating it rather than filling it with a guess. What the literature does establish is the shape and the scale of the problem. First, Iraqi is desperately data-poor: the Ozyilmaz paper states verbatim that in the MASC corpus 'the Egyptian (385h) and Levantine (148h) datasets were clearly overrepresented, while Iraqi (13h) and Maghrebi (17h) were more limited'. Iraqi is the second most data-starved dialect in the set. Second, and this is the number that should govern your planning: in that paper's cross-dialect confusion matrix, a whisper-small model FINE-TUNED ON IRAQI and TESTED ON IRAQI scored 79.96% WER (± 3.37). That is a model deliberately trained on the dialect, and four words in five still came out wrong. An out-of-the-box model that has never seen Iraqi has no reason to do better. Third, across the whole experiment the pre-trained models averaged 83.87% WER (σ 13.24) on dialectal test sets. Fourth, Talafha, Waheed and Abdul-Mageed (Interspeech 2023) found that Whisper's 'performance deteriorates significantly in the zero-shot setting for five unseen dialects (i.e., Algeria, Jordan, Palestine, UAE, and Yemen)'. Fifth, a 2026 Sudanese-dialect paper reports 'zero-shot multilingual Whisper (78.8% WER)' before fine-tuning. WE HAVE REMOVED two figures earlier drafts leaned on: the Algerian 53.8% WER, because on re-reading the abstract of arXiv:2411.13424 that number follows 'well-designed data processing pipelines and advanced decoding techniques' and is a best case rather than a baseline; and a 55.1% figure attributed to arXiv:2406.04512 for five under-represented dialects, because that attribution could not be confirmed and we would rather cut a number than misattribute one. THE OPERATIONAL CONSEQUENCE, which is the only part that matters to you: at 33% error a transcript is a draft you correct; at 80% it is barely a navigation aid. It tells you roughly where in the audio a topic sits and nothing more. Budget for the second, and be pleasantly surprised. Interviews where the speaker leans toward MSA (officials, prepared statements, formal testimony) come out far more usable than kitchen-table Baghdadi or southern dialect. One caveat in the other direction, stated for fairness: the Ozyilmaz figures are for the small model on at most 20 hours of fine-tuning data, and the paper's own limitations section discusses the absence of dialect-aware text normalisation, which can inflate error rates even where the phonetic content was recognised correctly. Large v3 Turbo out of the box may do somewhat better than 80%. It will not do 33%.

الكردية السورانية

WHISPER DOES NOT DO KURDISH AT ALL, and you should not spend an afternoon finding that out. Kurdish is not a supported Whisper language. We checked the source directly in this session: openai/whisper's tokenizer.py LANGUAGES dictionary has 100 entries, Arabic ('ar') is among them, and there is no Kurdish entry in any form, no 'ku', no 'ckb' (Sorani), no 'kmr' (Kurmanji). Because there is no language token, the models shipped inside Vibe, Buzz, Subtitle Edit, MacWhisper and whisper.cpp cannot be told to transcribe Sorani; set to auto-detect they will guess Persian, Arabic or Turkish and return confident nonsense. Nothing in this recipe helps you with Kurdish audio, and we would rather say so on the first page than let you discover it in the fourth hour. WHAT HAS CHANGED, AND IT IS REAL BUT OUT OF REACH. NVIDIA Parakeet TDT 0.6B v3, offered in Vibe, is confirmed at 25 European languages with no Kurdish and no Arabic. We read its model card. But Meta's Omnilingual ASR, released under Apache 2.0, does cover Central Kurdish, and we verified this first-hand rather than taking the marketing at its word: we downloaded the supported-language list from the source file src/omnilingual_asr/models/wav2vec2_llama/lang_ids.py and confirmed that 'ckb_Arab' is present among 1,668 language-and-script entries, alongside kmr_Arab, kmr_Cyrl and kmr_Latn for Kurmanji. Meta's own per-language results table for its 7B LLM-ASR model reports a character error rate of 6.0 for ckb_Arab, from 59.6 hours of training data. THAT IS GENUINELY GOOD NEWS AND IT IS ALSO NOT AVAILABLE TO YOU ON A LAPTOP. That CER 6.0 figure is for omniASR_LLM_7B, which the repository's own model table lists at a 30.0 GiB download needing about 17 GiB of inference VRAM. The smallest model in the family (omniASR_CTC_300M) still needs about 2 GiB of VRAM, and the smallest LLM-ASR variant about 5 GiB, and Meta publishes no per-language accuracy for the small models, so we cannot tell you what they would actually give you in Sorani. There is no graphical interface: it is 'pip install omnilingual-asr', a Python script, and a GPU. Note also that it is character error rate, not word error rate, and the two are not comparable. Community Whisper fine-tunes for Sorani also exist on Hugging Face, but they are small unbenchmarked research uploads with no published accuracy figures, they cannot be loaded by the point-and-click route in this recipe, and we have not tested them, search Hugging Face yourself if you want to try one. Loading an arbitrary community model is also a supply-chain decision, not just an accuracy one: prefer files in .safetensors format over .bin pickles, and treat weights from an unknown uploader the way you would treat an executable from an unknown sender. SO: FOR SORANI DOCUMENTATION WORK ON A LAPTOP IN ERBIL IN 2026, PLAN ON HUMAN TRANSCRIPTION. What has changed is that Sorani ASR is now a fundable ask rather than an impossibility, an organisation with a GPU server, or a technical partner with one, can now get real Central Kurdish speech recognition from an Apache-2.0 model. That is worth putting in a grant application. It is not worth putting on your laptop this week.

على ماذا يستند هذا

MSA figure: Ozyilmaz, Coler & Valdenegro-Toro, 'Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning', Interspeech 2025, arXiv:2506.02627. We downloaded the PDF and extracted its text in this session and found verbatim: 'Figure 1 showed a 45.93% WER on the pre-trained model before fine-tuning'. Iraqi data scarcity, same paper, verbatim: 'the Egyptian (385h) and Levantine (148h) datasets were clearly overrepresented, while Iraqi (13h) and Maghrebi (17h) were more limited'. Iraqi fine-tuned-on-Iraqi WER of 79.96 (± 3.37) read from the cross-dialect confusion matrix in the same extracted text; we confirmed the diagonal orientation by checking that the Levantine and Maghrebi rows also have their own dialect in the corresponding column position. Overall pre-trained mean, verbatim: 'Pre-trained models (µ = 83.87, σ = 13.24)'. Dialect deterioration: Talafha, Waheed & Abdul-Mageed, Interspeech 2023, arXiv:2306.02902, abstract read in this session, verbatim: 'its performance deteriorates significantly in the zero-shot setting for five unseen dialects (i.e., Algeria, Jordan, Palestine, UAE, and Yemen)'. Sudanese: arXiv:2601.06802, abstract read in this session, verbatim 'substantially outperforming zero-shot multilingual Whisper (78.8% WER)'. REMOVED IN THIS PASS: (a) the Algerian 53.8% figure from arXiv:2411.13424, because the abstract we read in this session says the ASR figures follow 'well-designed data processing pipelines and advanced decoding techniques', making 'Word Error Rate (WER) of 0.538' a post-optimisation result rather than a zero-shot baseline; (b) a 55.1% WER attributed to arXiv:2406.04512 for five under-represented Arabic dialects, deleted because we could not confirm that attribution and a prior check indicated the number is an overall average rather than the five-dialect score. Kurdish absence from Whisper: verified in this session by downloading https://raw.githubusercontent.com/openai/whisper/main/whisper/tokenizer.py and parsing the LANGUAGES dictionary, 100 entries, 'ar' present, no 'ku'/'ckb'/'kmr'. Kurdish presence in Omnilingual ASR: verified in this session by downloading lang_ids.py from facebookresearch/omnilingual-asr and grepping it, 'ckb_Arab' present, 1,668 language-script entries; the line 'ckb_Arab,59.6,6.0' read from the repository's own per_language_results_table_7B_llm_asr.csv; the 30.0 GiB / ~17 GiB and 1.3 GiB / ~2 GiB and 6.1 GiB / ~5 GiB rows read from the model table in the repository README. Parakeet: model card at huggingface.co/nvidia/parakeet-tdt-0.6b-v3 read in this session, 25 European languages, no Kurdish, no Arabic. UNTESTED BY US: no Iraqi-dialect WER is published in disaggregated zero-shot form in any source we located, and we did not run our own Iraqi-dialect test, the 50-to-80% expectation above is inference from the fine-tuned Iraqi figure and neighbouring dialects and is explicitly labelled as such. We also have not run Omnilingual ASR on Kurdish audio ourselves.

البرامج المذكورة في هذه الوصفة

  • Vibe
  • Buzz
  • Subtitle Edit
  • whisper.cpp
  • faster-whisper
  • OpenAI Whisper
  • Whisper large-v3-turbo
  • Silero VAD
  • NVIDIA Parakeet TDT v3
  • Meta Omnilingual ASR
  • Hugging Face
  • ffmpeg
  • Audacity
  • BitLocker
  • FileVault
  • MacWhisper
  • Whisper Notes

المصادر

  • https://github.com/thewh1teagle/vibe/releases/latest · release metadata read via the GitHub API in this session: current tag v3.1.6, published 2026-08-28. Asset names and byte sizes read from that same response: vibe_3.1.6_x64-setup.exe 43,592,944; vibe_3.1.6_aarch64.dmg 38,364,114; vibe_3.1.6_x64.dmg 44,111,146; vibe_3.1.6_amd64.deb 36,727,394; vibe-3.1.6-1.x86_64.rpm 36,721,050 (plus arm64/aarch64 .deb and .rpm, .app.tar.gz archives, and .sig files).
  • https://raw.githubusercontent.com/thewh1teagle/vibe/main/website/public/privacy_policy.md · read in this session, last updated 2026-02-08. Verbatim: 'Vibe operates fully offline. After the initial setup, where the model files are downloaded, all transcription happens entirely on your local device'; 'Vibe collects anonymous, privacy-friendly analytics through Aptabase'; 'Event names (e.g. transcription started, failed)'; 'Error messages (no file names or content)'; 'OS name, app version, and country (coarse, no IP stored)'; 'Analytics can be disabled at any time in Settings'; 'When using the Summarize option in the More Options menu (which is off by default), the transcription is sent to the Claude API'; 'There is no encryption involved in the app.'; updates 'come directly from GitHub releases. No other cloud storage or external servers are involved.'
  • https://raw.githubusercontent.com/thewh1teagle/vibe/main/docs/architecture.md · read in this session. Verbatim: 'macOS and Windows builds also bundle `ffmpeg` from the Sona release archives'; engine described as 'Rust + whisper.cpp bindings'.
  • https://raw.githubusercontent.com/thewh1teagle/vibe/main/docs/models.md · read in this session. Source of the model download links used in step 8; heading reads '### 🚀 Large v3 Turbo (Recommended)'; page states 'To install a model, use the "Magic Setup" link to open it in Vibe, or copy and paste the direct download link in Vibe settings.'
  • https://raw.githubusercontent.com/thewh1teagle/vibe/main/i18n/translations/en-US/desktop.json · read in this session. Every Vibe interface string quoted in this recipe was found verbatim in this file: 'Anonymous analytics'; 'We strongly recommend keeping this enabled so we can find and fix issues faster.'; 'Save projects' / 'Automatically save transcripts and media in the projects folder.'; 'Saved projects' / 'Every transcription Vibe saved, with its copy of the audio.'; 'Delete all projects' / 'Their transcripts and copies of the audio will be removed. This cannot be undone.'; 'Projects folder' / 'Vibe stores each transcript and its media together in this folder.'; 'Temporary files'; 'Models Folder'; 'Download model'; 'Paste Model Link'; 'Select Model'; 'No models installed'; 'Select a model in Settings to get started'; 'Download any supported model in ggml or gguf format (with the file extension ending in '.bin'). Transfer it to the models directory, then select from the dropdown. No need to restart'; 'Select Language'; 'Translate To English' / 'Translate transcription into English from any language by enabling this option'; 'Transcribe on the CPU' / 'Runs transcription without the GPU. Slower, but it avoids driver crashes and frees the GPU for other work.'; 'Select File'; 'Drop audio, video or a folder here'; 'Summarize' / 'Requires API key. Once the transcription is complete, it will be sent to Claude API along with the chosen prompt'; and the AVX2 error 'Your CPU is not supported by this version of Vibe. Please use a more modern PC.'
  • https://github.com/thewh1teagle/vibe/tree/main/i18n/translations · we probed this folder in this session for the plausible Arabic and Kurdish locale codes (ar, ar-SA, ar-EG, ku, ckb, ku-IQ) and every one returned 404, while he-IL and tr-TR returned 200. Vibe's interface is therefore not available in Arabic or Kurdish; check the folder yourself for the current list.
  • https://raw.githubusercontent.com/chidiwilliams/buzz/main/README.md · read in this session. Verbatim: 'Intel Macs: Buzz now requires Apple silicon. The last version to support Intel Macs is 1.4.5.'
  • https://github.com/chidiwilliams/buzz/releases/latest · read via the GitHub API in this session: tag v1.4.5, published 2026-08-23. Assets and byte sizes: Buzz-1.4.5-windows.exe 4,392,262; Buzz-1.4.5-windows-1.bin 2,095,607,552; Buzz-1.4.5-windows-2.bin 594,799,916 (about 2.7 GB in total); Buzz-1.4.5-mac-ARM64.dmg 469,325,527; Buzz-1.4.5-mac-X64.dmg 540,127,994; a manylinux Python wheel 60,590,705.
  • https://github.com/SubtitleEdit/subtitleedit · free, offline, with a built-in Whisper engine. We are not quoting a version or a file size; check the releases page for the current ones.
  • https://github.com/ggml-org/whisper.cpp · README read in this session. Source of every command in step 19: 'cmake -B build', 'cmake --build build -j --config Release', './build/bin/whisper-cli', 'sh ./models/download-ggml-model.sh', the '--vad' and '--vad-model' flags, and verbatim 'the whisper-cli example currently runs only with 16-bit WAV files, so make sure to convert your input before running the tool' with the line 'ffmpeg -i input.mp3 -ar 16000 -ac 1 -c:a pcm_s16le output.wav'. Its Memory usage table is the source of the per-model memory figures: tiny ~273 MB, base ~388 MB, small ~852 MB, medium ~2.1 GB, large ~3.9 GB (no row for turbo).
  • https://raw.githubusercontent.com/ggml-org/whisper.cpp/master/models/download-vad-model.sh and .../download-ggml-model.sh, both read in this session. 'silero-v6.2.0' is a valid target (models="silero-v5.1.2 silero-v6.2.0") and the script writes ggml-"$model".bin, i.e. models/ggml-silero-v6.2.0.bin; 'large-v3-turbo' and 'large-v3-turbo-q5_0' are valid targets in download-ggml-model.sh.
  • https://huggingface.co/ggerganov/whisper.cpp/tree/main · file sizes read from the Hugging Face API in this session, in bytes: ggml-tiny.bin 77,691,713; ggml-base.bin 147,951,465; ggml-small.bin 487,601,967; ggml-medium.bin 1,533,763,059; ggml-large-v3.bin 3,095,033,483; ggml-large-v3-turbo.bin 1,624,555,275; ggml-large-v3-turbo-q5_0.bin 574,041,195.
  • https://pypi.org/project/faster-whisper/ · read via the PyPI JSON API in this session: version 1.2.1, licence MIT, requires_python >=3.9. https://raw.githubusercontent.com/SYSTRAN/faster-whisper/master/faster_whisper/utils.py confirms '"large-v3-turbo": "mobiuslabsgmbh/faster-whisper-large-v3-turbo"', so the model string used in step 20 resolves.
  • Koenecke, A., Choi, A. S. G., Mei, K. X., Schellmann, H. & Sloane, M. (2024). 'Careless Whisper: Speech-to-Text Hallucination Harms.' ACM FAccT '24. https://arxiv.org/abs/2402.08021, PDF downloaded and text extracted in this session. Verbatim: 'running 13,140 audio segments'; 'Manual review confirmed that 187 audio segments reliably result in Whisper hallucinations'; '38% of hallucinations include' explicit harms; hallucination concentrated among speakers 'who speak with longer shares of non-vocal durations'.
  • Ozyilmaz, O. T., Coler, M. & Valdenegro-Toro, M. (2025). 'Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning.' Interspeech 2025. https://arxiv.org/abs/2506.02627, abstract and full PDF text read in this session. Verbatim: 'Figure 1 showed a 45.93% WER on the pre-trained model before fine-tuning'; 'the Egyptian (385h) and Levantine (148h) datasets were clearly overrepresented, while Iraqi (13h) and Maghrebi (17h) were more limited'; 'Pre-trained models (µ = 83.87, σ = 13.24)'. The Iraqi-on-Iraqi cell 79.96 (± 3.37) was read from the cross-dialect confusion matrix in the extracted text.
  • Talafha, B., Waheed, A. & Abdul-Mageed, M. (2023). 'N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition.' Interspeech 2023. https://arxiv.org/abs/2306.02902, abstract read in this session. Verbatim: 'its performance deteriorates significantly in the zero-shot setting for five unseen dialects (i.e., Algeria, Jordan, Palestine, UAE, and Yemen)'.
  • 'Doing More with Less: Data Augmentation for Sudanese Dialect ASR' (2026). https://arxiv.org/abs/2601.06802, abstract read in this session. Verbatim: 'substantially outperforming zero-shot multilingual Whisper (78.8% WER)'.
  • CITED IN ORDER TO EXPLAIN A REMOVAL: 'CAFE: Code-switching Dataset for Algerian Dialect' (2024). https://arxiv.org/abs/2411.13424, abstract read in this session, verbatim: 'we benchmark CAFE data with the aforementioned Whisper models and show how well-designed data processing pipelines and advanced decoding techniques can improve the ASR performance in terms of Mixed Error Rate (MER) of 0.310, C[ER]...'. The 53.8% WER is therefore a post-optimisation result, not a zero-shot baseline, and stays out of this recipe.
  • https://raw.githubusercontent.com/openai/whisper/main/whisper/tokenizer.py · downloaded and parsed in this session. The LANGUAGES dictionary has 100 entries (the last is 'yue'), includes 'ar', and contains no 'ku', 'ckb' or 'kmr'.
  • https://github.com/facebookresearch/omnilingual-asr · Apache 2.0. In this session we downloaded src/omnilingual_asr/models/wav2vec2_llama/lang_ids.py and parsed it: 1,668 language-script entries including 'ckb_Arab', 'kmr_Arab', 'kmr_Cyrl' and 'kmr_Latn'. per_language_results_table_7B_llm_asr.csv contains the line 'ckb_Arab,59.6,6.0'. The README model table gives omniASR_LLM_7B at 30.0 GiB download and ~17 GiB inference VRAM, omniASR_LLM_300M at 6.1 GiB and ~5 GiB, omniASR_CTC_300M at 1.3 GiB and ~2 GiB; installation is 'pip install omnilingual-asr'.
  • https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 · model card read in this session: 25 European languages listed, no Kurdish and no Arabic.
  • Microsoft, 'Redirect and move Windows known folders to OneDrive.' https://learn.microsoft.com/en-us/sharepoint/redirect-known-folders, read in this session. Verbatim: 'Windows known folders (Desktop, Documents, Pictures, Screenshots, and Camera Roll)'; and of the silent policy, 'Use this setting to redirect and move known folders to OneDrive without any user interaction.'
  • Apple, 'Add your Desktop and Documents files to iCloud Drive.' https://support.apple.com/en-us/109344, read in this session. Confirms the 'Desktop & Documents Folders' setting under System Settings > [your name] > iCloud > Drive, and: 'When you store your Desktop and Documents folders in iCloud Drive, you can access the files from your Mac on all your devices.'
  • https://www.audacityteam.org/download/ · free, GPL-licensed audio editor. Linked rather than quoted; check the page for the current builds.
  • https://www.macwhisper.com/ and https://whispernotes.app/, named so you can recognise them and ignore them. This recipe deliberately quotes no price or tier detail from either; open the pages if you want current terms.
  • https://brew.sh/ · the Homebrew install one-liner in step 19 is copied from this page. No Homebrew formula version numbers are quoted in this recipe.

من إنتاج CDR، وكُتبت لمؤتمر بوينت العراق 7. جرى التحقق منها مقابل صفحات المزوّدين وسجلات الحزم في آب/أغسطس 2026؛ والمصادر مدرجة في هذه الصفحة.

ما لم نتمكن من التحقق منه

Things we could not verify, things we removed, and things that will disappoint you.

FIRST, A NOTE ON HOW THIS VERSION WAS EDITED. Two earlier repairs of this recipe each fixed the problems they were handed and each introduced new invented specifics, because the format asks for exact prices, versions and benchmark figures, and a writer asked for a number will produce one. This pass reversed the default. Every number, price, version, model identifier, benchmark figure, date and citation that survives here was opened at its source during the editing session; everything else was deleted and replaced with a link you can click. That is why some sentences now say "check the vendor page" where they used to state a figure. It is not laziness. A shorter page where every claim is true is worth more to you than a longer one where some are invented.

SECOND, THE IRAQI-DIALECT NUMBER STILL DOES NOT EXIST. There is no published study reporting a disaggregated zero-shot Whisper word error rate for Iraqi/Mesopotamian Arabic. An earlier version filled that gap with "roughly one word in three to one word in two", which was better than every dialect figure it itself cited, an error in the reader's disfavour, because it is the difference between a draft you correct and a file you cannot use. We have raised the expectation to roughly half to four-fifths of words wrong, anchored on a number we read directly out of the Interspeech 2025 paper: a whisper-small model fine-tuned on Iraqi and tested on Iraqi still scored 79.96% WER. It remains an inference, not a measurement, and we did not run our own Iraqi-dialect test. Run the known-answer test in step 11 on your own speakers before planning any project around this.

THIRD, WE DELETED A MISATTRIBUTED DIALECT FIGURE. Earlier drafts stated that Waheed et al. (arXiv:2406.04512) measured Whisper-large-v2 at 55.1% WER "across five under-represented Arabic dialects". The paper is real and 55.1% appears in it, but a check indicated that number is an overall average rather than the five-dialect score. Rather than restate it with a corrected gloss we could not confirm ourselves, we cut it and the source entry with it. Nothing in the argument depended on it: the Iraqi, Sudanese and cross-dialect figures that remain are stronger evidence anyway.

FOURTH, WE FIXED A COUNT WE HAD ASSERTED AS A FIRST-HAND CHECK. Earlier drafts said Whisper's tokenizer lists "exactly 99 languages". We re-downloaded and parsed the file: it has 100 entries, the hundredth being 'yue' (Cantonese). The substantive claim was and remains correct: Arabic is present, Kurdish in every form is absent. But a wrong count inside a sentence claiming to be a source check is exactly the kind of thing that erodes trust in everything around it.

FIFTH, WE REWROTE STEP 8 BECAUSE IT DESCRIBED A SCREEN THAT DOES NOT EXIST. Earlier drafts told you to "find the model selector, click 'Download model', and choose 'Large v3 Turbo', which Vibe marks as Recommended". Vibe has no in-app menu of named models. The "(Recommended)" label lives on Vibe's models documentation page on the web, and the app instead offers a models folder, a "Paste Model Link" box and a "Select Model" dropdown of whatever you have downloaded. Step 8 now gives you the actual mechanism and the two direct model URLs, including the Small one, which earlier drafts named but never linked, leaving 4 GB readers with an instruction they could not follow.

SIXTH, WE REMOVED THE SOURCEFORGE FILE SIZES FOR BUZZ. Earlier drafts quoted a 2.69 GB Windows archive and a 4.83 GB Linux archive to the byte, sourced from a SourceForge feed. That feed returned nothing when we tried to re-read it, so we cut every SourceForge filename and figure. The bandwidth argument survives on GitHub numbers we did read: Buzz's Windows download there is three files totalling about 2.7 GB against Vibe's single 43,592,944-byte installer. We also corrected Buzz's macOS asset names, which are now .dmg files rather than the .zip names we had listed.

SEVENTH, WE DELETED ALL PRICES. MacWhisper and Whisper Notes are still named so you can recognise them, but no price, currency, trial length or licensing term for either appears in this recipe any more. Those go stale within weeks and this page will be read for a year. Click the vendor links. Nothing in this recipe costs money, so nothing depends on those figures.

EIGHTH, WE DELETED THE SPEED BENCHMARK. Earlier drafts extrapolated from the faster-whisper project's own CPU table, specific memory figures and run times for a specific model on server-class hardware, to a consumer laptop. We did not re-open that table, so it is gone. In its place: time a ten-minute file on your own machine and multiply. That measurement is worth more than the extrapolation was.

NINTH, WE REMOVED MACWHISPER RATHER THAN PATCHING IT. An earlier version made MacWhisper's free tier the primary path for Apple-silicon Macs, on the assumption that it can run a model large enough for Arabic. MacWhisper's own pricing page compares free and Pro on export formats, speaker recognition, batch processing and dictation, and says nothing about which Whisper models the free tier can download, and this recipe's own step 8 declares Tiny and Base unusable for Arabic. Rather than ship a free path whose free-ness we cannot confirm, we removed the route. Vibe serves Apple-silicon Macs with no such ambiguity. If someone in the room has a Mac and confirms MacWhisper's free model list first-hand, that is worth reporting back.

TENTH, WE REPLACED BUZZ WITH VIBE AS THE RECOMMENDED APP, and Buzz's Intel Mac situation is why. Buzz's README states that 1.4.5 is the last version supporting Intel Macs. An earlier recipe both routed Intel Macs to Buzz and told everyone to "take the higher version number", which together produce a download that cannot run and does not say why. Rather than add a caveat, we changed the recommendation: Vibe builds for Intel Macs, Apple silicon, Windows and Linux from the same release, so the advice is uniform. The bandwidth difference turned out to matter as much. Buzz remains a good app and is still listed as an alternative with the Intel pin spelled out. We have also written the Vibe download instructions by filename suffix rather than by full versioned filename, so they stay correct as new releases ship.

ELEVENTH, WE CUT THE whisper.cpp TERMINAL ROUTE ON WINDOWS ENTIRELY. Building it on Windows needs Visual Studio Build Tools and is a materially different, longer procedure. Rather than write half of it and let the majority platform dead-end at a missing compiler, we labelled the route macOS and Linux only, before the reader installs anything. Windows readers lose nothing they need.

TWELFTH, KURDISH SORANI STILL DOES NOT WORK ON YOUR LAPTOP. Whisper has no Kurdish token; that is verified against source and is not a settings problem. Meta's Omnilingual ASR (Apache 2.0) does cover Central Kurdish. We verified ckb_Arab is in its supported-language list and that its published table reports CER 6.0 for the 7B model. But that model is a 30 GiB download needing about 17 GiB of GPU memory, it is pip-and-Python only with no interface, Meta publishes no per-language accuracy for the small models, and CER is not WER. The practical conclusion is unchanged, plan on human transcription for Sorani, while the strategic conclusion has changed: this is now something to ask a funder or a technical partner for. We have not run it on Kurdish audio ourselves.

THIRTEENTH, VIBE'S INTERFACE IS NOT AVAILABLE IN ARABIC OR KURDISH. We probed the translations folder for every plausible Arabic and Kurdish locale code and found none, while Hebrew and Turkish are present. We are no longer printing the full list of locales that do exist, because we could not enumerate the folder cleanly, open it yourself if you want it. Your colleagues will be operating an English menu to transcribe Arabic. Vibe accepts translation contributions, which is a small, concrete thing an organisation in this room could do that would help everyone else in it.

FOURTEENTH AND MOST IMPORTANT: THIS RECIPE PROTECTS YOUR AUDIO FROM FOREIGN SERVERS. IT DOES NOT MAKE THE TRANSCRIPT TRUE. A local Whisper transcript of an Arabic interview is a fast, private, free way to find your way around a recording, and it is not a record of what a witness said until a human has listened. If your organisation adopts this, adopt the listening pass and the folder rule with it, and write all three into your AI policy at the same time.

كل الوصفات