At least English, German, French, Russian, Italian, and Spanish. You can either select them right in the app or let it auto-detect (which it does pretty well). I think more languages are supported, though, just did not try them.
Handy is honestly more mature than mine — cross platform, MIT, and you can pick between Whisper and Parakeet. The real difference is just the model family. I run Qwen3-ASR through MLX instead of Whisper/Parakeet. In my own use it holds up way better with background noise and it punctuate noticeably more naturally, Take that as impression and not a measurement. Qwen3-ASR also accept vocabulary hints, which helps a lot with names and jargon. And mine does files and video with SRT export, not only dictation.
Downside: Apple Silicon only.
Wispr Flow is a different category — it's paid and your audio goes to their cloud. The text formatting is nicer than mine. Mine is local only and Apache-2.0, which for me is the entire point.
1. I see it uses python, but I was wondering if it would use mlx-swift instead. Just curious if you considered it.
2. So far, qwen ASR seem usually less well supported than parakeet or whisper so far in most dictation applications (voxtral is similar in that regard). Why do you think that is?
way better from my point of view. I frequently have recordings made in noisy environments. Whisper does not do the job well under such conditions. Also, Qwen respects gramma a lot and splits sentences properly.
What languages does it support?
At least English, German, French, Russian, Italian, and Spanish. You can either select them right in the app or let it auto-detect (which it does pretty well). I think more languages are supported, though, just did not try them.
How does it compare to models from handy.computer and whisprflow?
Handy is honestly more mature than mine — cross platform, MIT, and you can pick between Whisper and Parakeet. The real difference is just the model family. I run Qwen3-ASR through MLX instead of Whisper/Parakeet. In my own use it holds up way better with background noise and it punctuate noticeably more naturally, Take that as impression and not a measurement. Qwen3-ASR also accept vocabulary hints, which helps a lot with names and jargon. And mine does files and video with SRT export, not only dictation. Downside: Apple Silicon only. Wispr Flow is a different category — it's paid and your audio goes to their cloud. The text formatting is nicer than mine. Mine is local only and Apache-2.0, which for me is the entire point.
Congrats on the launch!
1. I see it uses python, but I was wondering if it would use mlx-swift instead. Just curious if you considered it.
2. So far, qwen ASR seem usually less well supported than parakeet or whisper so far in most dictation applications (voxtral is similar in that regard). Why do you think that is?
Thanks and good luck with your roadmap!
is it better than whisper from oai?
way better from my point of view. I frequently have recordings made in noisy environments. Whisper does not do the job well under such conditions. Also, Qwen respects gramma a lot and splits sentences properly.
I kinda love that you misspelled 'gramma' when saying this model respects it a lot.
Well, at least one of us respects it :)