Why faster-whisper large-v3-turbo Is the Optimal Model for Interviews
In short: large-v3-turbo in faster-whisper delivers nearly the maximum accuracy of large-v3 but noticeably faster — the optimal balance for real time. Smaller models lose on technical vocabulary.
What faster-whisper is
An optimized Whisper implementation on CTranslate2. Faster than the original at the same quality, supports GPU via CUDA. In Mockingbird, recognition runs locally.
Why turbo
- Speed. Turbo reduces the number of decoding steps.
- Accuracy. Quality is close to large-v3.
- Real time. It keeps up with speech as it arrives.
Model comparison
| Model | Accuracy | Speed | When to choose |
|---|---|---|---|
tiny / base | low | very high | CPU only |
small / medium | medium | high | weak GPU |
large-v3-turbo | high | high | CUDA GPU (recommended) |
large-v3 | maximum | lower | strong GPU |
Languages and terms
Whisper is multilingual and works confidently across languages. Larger models recognize technical terms better, and phonetic correction (e.g. “Zabix” → “Zabbix”) runs on top.
Conclusion
If you have a GPU — start with large-v3-turbo. Not enough resources — move down the table. Instructions — in the documentation.
FAQ
Which model is best for my language?
large-v3-turbo gives a good balance. For maximum accuracy with no speed constraints — large-v3.
How much video memory do I need?
It depends on the model. For large-v3-turbo, 6 GB or more is desirable; if it's short, the app switches to CPU.
Can it run without a graphics card?
Yes, but recognition will be slower. An NVIDIA GPU with CUDA is recommended for fast operation.
Mockingbird — an offline interview assistant: local speech recognition, a knowledge base built from your résumé, answers in 2–5 seconds. Installation takes minutes.
