Offline or Cloud: Where to Process Speech in an Interview
In short: for interviews, local speech recognition is preferable — it doesn't send your voice to a server, works without internet and gives predictable latency. The cloud is justified only if your computer is very weak.
Privacy
In an interview you say questions and reasoning out loud. Sending that to someone else's server means depending on their storage policy. Offline eliminates the risk by definition.
Latency
- Cloud: network transfer + queue + processing — latency fluctuates.
- Offline on GPU: a stable 2–5 seconds.
Internet dependency
Wi-Fi drops mid-interview — a cloud assistant stops working. Offline keeps recognizing speech.
Cost
Cloud STT APIs charge per minute of audio. An offline model is downloaded once and works without limits.
Comparison
| Criterion | Offline | Cloud |
|---|---|---|
| Audio leak | impossible | depends on the provider |
| Latency | 2–5 s (GPU) | fluctuating |
| Internet | not needed | required |
| Cost | zero | per minute |
| PC requirements | some (GPU recommended) | minimal |
Conclusion
For a technical interview the priorities are privacy and stability, so offline wins. That's exactly how Mockingbird is built: speech, résumé and the knowledge base are processed locally; only the LLM request goes out.
FAQ
Can speech be recognized without internet?
Yes, local faster-whisper models run fully offline. Internet is only needed for the language model request.
Is cloud recognition more accurate?
Not necessarily. Local large-v3-turbo models match cloud accuracy, especially on technical vocabulary.
What goes to the cloud while the assistant runs?
In Mockingbird — only the LLM request text. Audio, résumé and knowledge base stay on your device.
Mockingbird — an offline interview assistant: local speech recognition, a knowledge base built from your résumé, answers in 2–5 seconds. Installation takes minutes.
