Kapinote logoDocumentation
Transcription

Choose a Transcription Model

Compare local on-device and cloud speech models by privacy, setup difficulty, language coverage, and resource usage.

Open Settings → Speech to select and test the model used for new recordings. The best choice depends on your macOS version, meeting languages, privacy constraints, and whether your Mac will be offline.

Quick comparison

ModelRuns WhereSetupBest Fit
Apple SpeechmacOS native on your MacGrant Speech Recognition permission; depends on downloaded Apple language assetsLightweight, zero-setup local transcription on supported Macs
Qwen3-ASR 0.6BOn your MacOne-click model download (~1.2 GB) in Kapinote100% offline use, complex multilingual discussions, and 22 Chinese dialects
Soniox RealtimeSoniox cloudAPI keyLive multilingual and mid-sentence code-switching conversations
Deepgram NovaDeepgram cloudAPI keyLow latency, speaker diarization, endpointing, and custom keyterm prompting
AssemblyAI StreamingAssemblyAI cloudAPI keyAutomatic language detection across supported streaming languages
ElevenLabs ScribeElevenLabs cloudAPI keyHigh accuracy and broad cloud language coverage with real-time streaming

Cloud provider capabilities and quota policies can change. Kapinote displays options currently integrated with the app; consult linked provider documentation for billing, regional availability, and model lifecycles.

Local on-device models (privacy first)

Local transcription keeps your audio entirely on your Mac and works reliably without an active internet connection once assets are installed. It utilizes your Mac's CPU, GPU, unified memory, and disk storage. Run the built-in test to verify real-time processing speed on your hardware.

  • Apple Speech: Minimal battery and RAM footprint; ideal for straightforward meetings in system-supported languages.
  • Qwen3-ASR: State-of-the-art open speech model supporting 30 global languages and 22 Chinese dialects/accents. See Qwen3-ASR languages and dialects.

Cloud transcription providers

Cloud models stream captured audio directly to your configured provider. They offload local resource consumption and offer advanced diarization and endpoint controls, but require an active internet connection, provider account, and valid API key. See Cloud transcription providers.

Practical decision checklist

  • Confidential & NDA Discussions: Verify organization policies; select on-device processing (Apple Speech / Qwen3-ASR) to keep audio local.
  • Multilingual or Code-Switching: Ensure your selected model and mode cover all expected conversation languages (e.g., Qwen3-ASR or Soniox).
  • Specialized Jargon & Names: Add Custom vocabulary.
  • Travel or In-flight Recording: Download and test Qwen3-ASR locally before traveling.

Note: Switching models only affects future recordings. Past meetings are not automatically re-transcribed.