1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
|
High-performance inference of OpenAI's Whisper automatic speech
recognition (ASR) model, plus NVIDIA Parakeet models.
- Plain C/C++ implementation without dependencies
- AVX intrinsics support for x86 architectures
- Mixed F16 / F32 precision
- Integer quantization support
- Vulkan and OpenVINO GPU backends
- OpenVINO encoder support
Installs whisper-cli, whisper-server, whisper-bench, whisper-quantize,
whisper-vad-speech-segments, parakeet-cli and parakeet-quantize, plus
the SDL2 microphone tools whisper-stream (live transcription),
whisper-command (voice commands), whisper-talk-llama and whisper-lsp.
The tools decode input audio through ffmpeg, so any format ffmpeg reads
is accepted, not only WAV, MP3 and FLAC.
OpenVINO support is built in. The encoder must be converted to OpenVINO
IR format first, see "OpenVINO Support" in the bundled README.md.
Both GPU backends are built in; whisper uses the first GPU device it
finds. Pick one with whisper-cli -dev N (the startup log lists the
devices) and compare the timings it prints at exit.
OpenBLAS is autodetected: if installed, it is used for the BLAS backend.
|