Mockingbirddocs

Choosing an STT model

In short: the default is large-v3-turbo — the optimal balance of accuracy and speed on GPU. If resources are tight, pick a smaller model from the table below. Mockingbird recognizes speech locally viafaster-whisper, so the model choice affects accuracy, speed and video memory usage.

Model size

The larger the model, the more accurate the recognition and the more resources it needs. The default is large-v3-turbo — the optimum of quality and speed on GPU.

ModelAccuracySpeedWhen to choose
tiny / baselowvery highCPU only, weak hardware
small / mediummediumhighspare CPU capacity or a weak GPU
large-v3-turbohighhighCUDA GPU (recommended)
large-v3maximumlowerstrong GPU, accuracy matters

What to pay attention to

Changing the model

The model is selected in the app settings. After switching, the first load may take a while — model files are downloaded once and cached locally.

Start with large-v3-turbo on GPU. If resources are not enough — move down the table until the speed is comfortable.