Kapinote logoDocumentation
Transcription

Qwen3-ASR Languages and Dialects

Install the local on-device Qwen3-ASR model and review its full list of 30 supported global languages and 22 Chinese dialects.

Kapinote provides out-of-the-box support for the open-source Qwen3-ASR 0.6B local on-device transcription model. Once downloaded to your Mac, speech recognition runs completely offline without consuming cloud quotas or transmitting audio data externally.

Setup steps

  1. Open Settings → Speech.
  2. Select Qwen3-ASR 0.6B from the model dropdown.
  3. Start the one-click model download (~1.2 GB) and keep Kapinote open until the download and checksum verify.
  4. Select your preferred language (or keep automatic detection) and run the built-in test.
  5. Record a short test sample containing accents and technical vocabulary common in your meetings.

Note: Initial setup requires an internet connection and at least 2 GB of free disk space. Real-time transcription speed depends on your Mac's Apple silicon configuration (M-series unified memory and GPU cores).

30 officially supported languages

According to the official Qwen3-ASR repository specifications, the model is trained and optimized for the following 30 languages:

Language & CodeLanguage & CodeLanguage & CodeLanguage & Code
Chinese (zh)English (en)Cantonese (yue)Arabic (ar)
German (de)French (fr)Spanish (es)Portuguese (pt)
Indonesian (id)Italian (it)Korean (ko)Russian (ru)
Thai (th)Vietnamese (vi)Japanese (ja)Turkish (tr)
Hindi (hi)Malay (ms)Dutch (nl)Swedish (sv)
Danish (da)Finnish (fi)Polish (pl)Czech (cs)
Filipino (fil)Persian (fa)Greek (el)Hungarian (hu)
Macedonian (mk)Romanian (ro)

22 officially supported Chinese dialects & accents

Qwen3-ASR offers exceptional adaptability for Chinese regional dialects and accents:

  • Northern & Central Plain: Dongbei, Hebei, Tianjin, Shandong, Henan, Shaanxi, Shanxi
  • Southwestern & Central: Sichuan, Hubei, Hunan, Guizhou, Yunnan, Jiangxi, Anhui
  • Northwestern: Gansu, Ningxia
  • Southern & Southeastern: Cantonese (Hong Kong accent), Cantonese (Guangdong accent), Wu, Minnan, Fujian, Zhejiang

Language support indicates the acoustic model is specifically trained on these accents. Real-world accuracy still depends on microphone quality, background noise, multi-speaker crosstalk, and vocabulary complexity. See the official Qwen3-ASR GitHub repository for ongoing updates.

Tips for better local transcription

  • Microphone Placement: Use a dedicated headset or keep your Mac's microphone close to speakers to minimize room reverberation.
  • Language Hints: If a meeting is primarily conducted in one language, explicitly select it rather than relying on automatic language detection.
  • Custom Vocabulary: Add specialized project codenames, team member names, and acronyms in Custom vocabulary. While local Qwen transcription runs acoustically on-device, this vocabulary is supplied to downstream AI summary generation to guarantee perfect spelling in notes.