Choose a Transcription Model
Compare local on-device and cloud speech models by privacy, setup difficulty, language coverage, and resource usage.
Open Settings → Speech to select and test the model used for new recordings. The best choice depends on your macOS version, meeting languages, privacy constraints, and whether your Mac will be offline.
Quick comparison
| Model | Runs Where | Setup | Best Fit |
|---|---|---|---|
| Apple Speech | macOS native on your Mac | Grant Speech Recognition permission; depends on downloaded Apple language assets | Lightweight, zero-setup local transcription on supported Macs |
| Qwen3-ASR 0.6B | On your Mac | One-click model download (~1.2 GB) in Kapinote | 100% offline use, complex multilingual discussions, and 22 Chinese dialects |
| Soniox Realtime | Soniox cloud | API key | Live multilingual and mid-sentence code-switching conversations |
| Deepgram Nova | Deepgram cloud | API key | Low latency, speaker diarization, endpointing, and custom keyterm prompting |
| AssemblyAI Streaming | AssemblyAI cloud | API key | Automatic language detection across supported streaming languages |
| ElevenLabs Scribe | ElevenLabs cloud | API key | High accuracy and broad cloud language coverage with real-time streaming |
Cloud provider capabilities and quota policies can change. Kapinote displays options currently integrated with the app; consult linked provider documentation for billing, regional availability, and model lifecycles.
Local on-device models (privacy first)
Local transcription keeps your audio entirely on your Mac and works reliably without an active internet connection once assets are installed. It utilizes your Mac's CPU, GPU, unified memory, and disk storage. Run the built-in test to verify real-time processing speed on your hardware.
- Apple Speech: Minimal battery and RAM footprint; ideal for straightforward meetings in system-supported languages.
- Qwen3-ASR: State-of-the-art open speech model supporting 30 global languages and 22 Chinese dialects/accents. See Qwen3-ASR languages and dialects.
Cloud transcription providers
Cloud models stream captured audio directly to your configured provider. They offload local resource consumption and offer advanced diarization and endpoint controls, but require an active internet connection, provider account, and valid API key. See Cloud transcription providers.
Practical decision checklist
- Confidential & NDA Discussions: Verify organization policies; select on-device processing (Apple Speech / Qwen3-ASR) to keep audio local.
- Multilingual or Code-Switching: Ensure your selected model and mode cover all expected conversation languages (e.g., Qwen3-ASR or Soniox).
- Specialized Jargon & Names: Add Custom vocabulary.
- Travel or In-flight Recording: Download and test Qwen3-ASR locally before traveling.
Note: Switching models only affects future recordings. Past meetings are not automatically re-transcribed.