Xenobd/whisper.cpp
0
1# whisper.cpp/examples/cli
2
3This is the main example demonstrating most of the functionality of the Whisper model.
4It can be used as a reference for using the `whisper.cpp` library in other projects.
5
6```
7./build/bin/whisper-cli -h
8
9usage: ./build/bin/whisper-cli [options] file0 file1 ...
10supported audio formats: flac, mp3, ogg, wav
11
12options:
13 -h, --help [default] show this help message and exit
14 -t N, --threads N [4 ] number of threads to use during computation
15 -p N, --processors N [1 ] number of processors to use during computation
16 -ot N, --offset-t N [0 ] time offset in milliseconds
17 -on N, --offset-n N [0 ] segment index offset
18 -d N, --duration N [0 ] duration of audio to process in milliseconds
19 -mc N, --max-context N [-1 ] maximum number of text context tokens to store
20 -ml N, --max-len N [0 ] maximum segment length in characters
21 -sow, --split-on-word [false ] split on word rather than on token
22 -bo N, --best-of N [5 ] number of best candidates to keep
23 -bs N, --beam-size N [5 ] beam size for beam search
24 -ac N, --audio-ctx N [0 ] audio context size (0 - all)
25 -wt N, --word-thold N [0.01 ] word timestamp probability threshold
26 -et N, --entropy-thold N [2.40 ] entropy threshold for decoder fail
27 -lpt N, --logprob-thold N [-1.00 ] log probability threshold for decoder fail
28 -nth N, --no-speech-thold N [0.60 ] no speech threshold
29 -tp, --temperature N [0.00 ] The sampling temperature, between 0 and 1
30 -tpi, --temperature-inc N [0.20 ] The increment of temperature, between 0 and 1
31 -debug, --debug-mode [false ] enable debug mode (eg. dump log_mel)
32 -tr, --translate [false ] translate from source language to english
33 -di, --diarize [false ] stereo audio diarization
34 -tdrz, --tinydiarize [false ] enable tinydiarize (requires a tdrz model)
35 -nf, --no-fallback [false ] do not use temperature fallback while decoding
36 -otxt, --output-txt [false ] output result in a text file
37 -ovtt, --output-vtt [false ] output result in a vtt file
38 -osrt, --output-srt [false ] output result in a srt file
39 -olrc, --output-lrc [false ] output result in a lrc file
40 -owts, --output-words [false ] output script for generating karaoke video
41 -fp, --font-path [/System/Library/Fonts/Supplemental/Courier New Bold.ttf] path to a monospace font for karaoke video
42 -ocsv, --output-csv [false ] output result in a CSV file
43 -oj, --output-json [false ] output result in a JSON file
44 -ojf, --output-json-full [false ] include more information in the JSON file
45 -of FNAME, --output-file FNAME [ ] output file path (without file extension)
46 -np, --no-prints [false ] do not print anything other than the results
47 -ps, --print-special [false ] print special tokens
48 -pc, --print-colors [false ] print colors
49 -pp, --print-progress [false ] print progress
50 -nt, --no-timestamps [false ] do not print timestamps
51 -l LANG, --language LANG [en ] spoken language ('auto' for auto-detect)
52 -dl, --detect-language [false ] exit after automatically detecting language
53 --prompt PROMPT [ ] initial prompt (max n_text_ctx/2 tokens)
54 -m FNAME, --model FNAME [models/ggml-base.en.bin] model path
55 -f FNAME, --file FNAME [ ] input audio file path
56 -oved D, --ov-e-device DNAME [CPU ] the OpenVINO device used for encode inference
57 -dtw MODEL --dtw MODEL [ ] compute token-level timestamps
58 -ls, --log-score [false ] log best decoder scores of tokens
59 -ng, --no-gpu [false ] disable GPU
60 -fa, --flash-attn [false ] flash attention
61 -sns, --suppress-nst [false ] suppress non-speech tokens
62 --suppress-regex REGEX [ ] regular expression matching tokens to suppress
63 --grammar GRAMMAR [ ] GBNF grammar to guide decoding
64 --grammar-rule RULE [ ] top-level GBNF grammar rule name
65 --grammar-penalty N [100.0 ] scales down logits of nongrammar tokens
66```
67 