Local AI
Vibe
A desktop transcription tool that uses Whisper models to turn speech in recordings into text.
Turn recorded speech into text locally
Vibe is an open source transcription application that can convert spoken language in audio and video recordings into written text.
It uses Whisper speech-recognition models and can perform transcription locally on the user's computer.
An existing audio or video file can be loaded and processed by a selected speech-recognition model. The result can then be used as normal text or as the basis for subtitles.
This is useful for interviews, lectures, meetings, voice notes, podcasts and video recordings. Local processing is particularly interesting when users prefer not to upload recordings to an online transcription service. Processing time depends on the computer, the selected model and the length of the recording.
Why we selected it
Vibe shows how modern AI technology can also be available through open source desktop software rather than only through cloud services. The application gives ordinary Windows users a relatively accessible way to experiment with automatic speech recognition without having to build their own Whisper environment.
Learn more and download
For current versions, system requirements and downloads, use the official project website. Our German catalogue also contains the OpenSource-DVD entry for this program.