Choosing an STT model
In short: the default is large-v3-turbo — the optimal balance of accuracy and speed on GPU. If resources are tight, pick a smaller model from the table below. Mockingbird recognizes speech locally viafaster-whisper, so the model choice affects accuracy, speed and video memory usage.
Model size
The larger the model, the more accurate the recognition and the more resources it needs. The default is large-v3-turbo — the optimum of quality and speed on GPU.
| Model | Accuracy | Speed | When to choose |
|---|---|---|---|
tiny / base | low | very high | CPU only, weak hardware |
small / medium | medium | high | spare CPU capacity or a weak GPU |
large-v3-turbo | high | high | CUDA GPU (recommended) |
large-v3 | maximum | lower | strong GPU, accuracy matters |
What to pay attention to
- GPU or CPU. Large models run slowly on CPU — the toolbar indicator shows
CPUif CUDA is unavailable. - Language. Pick models that support your language; Whisper is multilingual, but quality depends on the training data for each language.
- Terms. Technical jargon and abbreviations are recognized better by larger models; phonetic term correction works on top of any model.
- Latency. A smaller model means lower latency from the end of the question to the answer.
- Memory. Make sure you have enough free video memory for the chosen size.
Changing the model
The model is selected in the app settings. After switching, the first load may take a while — model files are downloaded once and cached locally.
Start with large-v3-turbo on GPU. If resources are not enough — move down the table until the speed is comfortable.