whisper.cpp

mirror of https://github.com/ggerganov/whisper.cpp.git synced 2025-06-01 15:36:27 +02:00

Author	SHA1	Message	Date
Georgi Gerganov	7fd6fa8097	talk-llama : sync llama.cpp ggml-ci	2025-06-01 15:14:44 +03:00
Daniel Bevenius	73a8c5fb94	whisper : remove whisper_load_backends function (#3196 ) * whisper : remove whisper_load_backends function This commit removes the `whisper_load_backends` function, which was used to load all GGML backends. The motivation for this change push the responsibility of loading backends to user applications to give them more control over which backends to load and when. See the references below for more context. Resolves: https://github.com/ggml-org/whisper.cpp/issues/3182 Refs: https://github.com/ggml-org/whisper.cpp/pull/3042#issuecomment-2801778733 Refs: https://github.com/ggml-org/whisper.cpp/pull/3042#issuecomment-2801928990 * ruby : add check for rwc is NULL This commit adds a check to ensure that the `rwc` pointer is not NULL before attempting to mark its members in the garbage collector. The motivation for this is an attempt to see if this fixed the CI build as I'm not able to reproduce the issue locally. Refs: https://github.com/ggml-org/whisper.cpp/actions/runs/15299612277/job/43036694928?pr=3196	2025-05-29 08:03:17 +02:00
Georgi Gerganov	26eb48cb08	talk-llama : sync llama.cpp ggml-ci	2025-05-27 18:03:00 +03:00
Daniel Bevenius	450de0787e	node : enable no_prints to suppress all output (#3189 ) This commit enable the node addon to suppress all output, even the result of the transcription if the no_prints parameter is set to true. The motivation for this is that for the node addon there is a fullfilment handler/success callback to process the transcription result. And it might be useful to be able to disable the printing of the transcription result to the console, so that the user can handle the result in their own way. Refs: https://github.com/ggml-org/whisper.cpp/issues/3176	2025-05-27 05:51:47 +02:00
matteng1	ea9f206f18	talk-llama : fix for swedish umlauts + expose model inference settings in talk-llama.cpp (#3187 ) Quick fix for not removing swedish umlauts. * Update talk-llama.cpp Expose model inference settings to user instead of hard coding them. Same defaults as previous defaults. * Update examples/talk-llama/talk-llama.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-05-26 07:57:39 +02:00
Sacha Arbonel	78b31ca782	server : Add k6 Load Testing Script (#3175 ) * add load testing script and update README for k6 integration	2025-05-22 10:03:04 +02:00
Georgi Gerganov	6b6cf19c65	talk-llama : sync llama.cpp ggml-ci	2025-05-19 14:58:39 +03:00
Daniel Bevenius	f389d7e3e5	examples : add --print-confidence option to cli (#3150 ) * examples : add --print-confidence option to cli This commit adds a new command-line option `--print-confidence` to the whisper-cli. When enabled, this option prints the confidence level of each token in the transcribed text using ANSI formatting codes. The confidence levels are represented using different styles: ```console main: confidence: highlighted (low confidence), underlined (medium), dim (high confidence) ``` Refs: https://github.com/ggml-org/whisper.cpp/issues/3135	2025-05-14 19:21:48 +02:00
Daniel Bevenius	3882a099e1	server : add --flash-attn usage output (#3152 ) This commit adds the `--flash-attn` option to the usage output of the server example. The motivation for this change is that while it is possible to set this option it is not printed in the usage output.	2025-05-14 15:22:05 +02:00
Georgi Gerganov	f890560575	talk-llama : sync llama.cpp ggml-ci	2025-05-13 13:59:21 +03:00
Daniel Bevenius	fbad8058c4	examples : add VAD speech segments example (#3147 ) This commit adds an example that demonstrates how to use a VAD (Voice Activity Detection) model to segment an audio file into speech segments. Resolves: https://github.com/ggml-org/whisper.cpp/issues/3144	2025-05-13 12:31:00 +02:00
Daniel Bevenius	b2513a6208	vad : remove shortform for --vad option in cli.cpp (#3145 ) This commit removes the shortform for the --vad option in cli.cpp. The motivation for this is that `-v` is often used for verbose or version is many tools and this might cause confusion. Refs: https://github.com/ggml-org/whisper.cpp/pull/3065#issuecomment-2873243334	2025-05-13 06:04:05 +02:00
Tomer Schlesinger	587ea01f55	docs : update README.md for whisper.objc app (#2569 )	2025-05-13 06:03:50 +02:00
Daniel Bevenius	e41bc5c61a	vad : add initial Voice Activity Detection (VAD) support (#3065 ) * vad : add initial Voice Activity Detection (VAD) support This commit add support for Voice Activity Detection (VAD). When enabled this feature will process the audio input and detect speech segments. This information is then used to reduce the number of samples that need to be processed by whisper_full. Resolves: https://github.com/ggml-org/whisper.cpp/issues/3003 --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-05-12 16:10:11 +02:00
Daniel Bevenius	186855e38b	cli : print color scheme info for --print-colors (#3141 ) This commit adds a description of the color scheme used in the CLI when the --print-colors option is enabled. The motivation for this is that it is not immediately clear what the color scheme is when using the CLI with the --print-colors option. Example output: ```console $ ./build/bin/whisper-cli -f samples/jfk.wav --print-colors ... main: color scheme: red (low confidence), yellow (medium), green (high confidence) [00:00:00.000 --> 00:00:11.000] And so my fellow Americans, ask not what your country can do for you, ask what you can do for your country. ``` The description will not be dispayed if the `--no-prints` options is set. Refs: https://github.com/ggml-org/whisper.cpp/issues/3135	2025-05-12 10:43:04 +02:00
Daniel Bevenius	4730950492	examples : update link to Paul Tol's color scheme [no ci] (#3140 ) This commit updates the link to Paul Tol's color scheme in the `examples/common.h` file. The previous link was outdated and pointed to a non-existent page.	2025-05-12 09:02:06 +02:00
Enes Grahovac	5d4390d281	examples : add HEAPU8 to all of the exported runtime methods (#3134 ) This commit adds HEAPU8 to the list of exported methods. The motivation for this commit is that currently this is causing an error on Window systems where HEAPU8 in undefined, which results in the following error message in the web console: main.js:1 Uncaught TypeError: Cannot read properties of undefined (reading 'buffer') at __emval_get_property (main.js:1:1363125) at 003a453a:0xc4a47 at 003a453a:0xc51cd at Object.full_default (eval at craftInvokerFunction (main.js:1:1347011), <anonymous>:9:10) at whisper.cpp/:647:42 danbev originally fixed this for whisper.wasm, stream.wasm, and command.stream, but the issue still exists on the other examples which I patch in this code. Resolves: #3059	2025-05-10 06:44:13 +02:00
Daniel Bevenius	9791647653	wasm : add note about worker.js file generation [no ci] (#3133 ) This commit updates the documentation for the WASM examples to include a note about the generation of the `worker.js` file. As of Emscripten 3.1.58 (April 2024), separate worker.js files are no longer generated and the worker is embedded in the main JS file. The motivation for this change is to inform users about the new behavior of Emscripten and why the `worker.js` file may not be present. Refs: https://github.com/ggml-org/whisper.cpp/issues/3123	2025-05-09 15:42:45 +02:00
Daniel Bevenius	b6f3fa4059	stream.wasm : add HEAPU8 to exported runtime methods (#3130 ) * stream.wasm : add HEAPU8 to exported runtime methods This commit adds HEAPU8 to the list of exported methods for stream.wasm. The motivation for this is that without it HEAPUD8 will be undefined and when its 'buffer' attribute is accessed this will cause error as reported in the referenced issue. Note that to test this make sure that the web browsers caches is cleared first. Resolves: https://github.com/ggml-org/whisper.cpp/issues/3123 * command.wasm : add HEAPU8 to exported runtime methods	2025-05-08 16:58:34 +02:00
Georgi Gerganov	4a512cb153	cli : avoid std::exchange ggml-ci	2025-05-07 15:39:32 +03:00
Daniel Bevenius	09846f4e12	whisper: remove MSVC warnings pragmas (#3090 ) * ggml : remove MSVC warnings pragmas This commit removes the MSVC-specific pragmas as these are now handled in CMakeLists.txt. * whisper : remove MSVC warning pragmas This commit removes the MSVC-specific pragmas. These are now handled in the CMakeLists.txt file.	2025-05-05 13:09:35 +02:00
Sacha Arbonel	bcf1ed0163	server: update abort mechanism to handle HTTP connection closure (#3112 )	2025-05-05 07:16:54 +02:00
Daniel Tang	934d4b3083	cli : support "-" for stdout like stdin (#3050 ) This changes examples/cli/cli.cpp to be like examples/common-whisper.cpp. "-of -" can be specified (or this can be inferred from "-" as the input file) to output to stdout. This is useful for piping to other applications. Log fname_out consistently when not stdout - Terminals have stdout=stderr, so remove the message before successful output to ease copying - Don't affect actual error messages - Move opening the ofstream into the factory, fixing missing open and/or error messages in output_score/output_wts - Fix struct naming convention Closes #3048	2025-05-05 07:15:39 +02:00
Arpit Jain	988dcd4b5b	docs : Update cli documentation (#3102 ) * docs : Update cli documentation This updates the documentation of cli based on the actual output In the longterm this should ideally be auto generated to prevent mismatch * docs : Update cli documentation This updates the documentation of cli based on the actual output In the longterm this should ideally be auto generated to prevent mismatch	2025-05-02 14:18:33 +02:00
Sacha Arbonel	1fa17bc752	server : update httplib.h to version 0.20.0 (#3101 )	2025-05-02 06:09:41 +02:00
Georgi Gerganov	0778b6ff5f	talk-llama : sync llama.cpp ggml-ci	2025-05-01 13:29:02 +03:00
Daniel Bevenius	25efcfe3ed	server : add --no-gpu option to print usage output (#3098 ) This commit adds the the command line option `--no-gpu` to the server examples print usage function. The motivation for this is that this options is available and can be set but it is not displayed in the usage message. Refs: https://github.com/ggml-org/whisper.cpp/issues/3095	2025-05-01 09:15:12 +03:00
Sacha Arbonel	f0171f0616	examples : expose language detection probabilities to server example (#3044 ) * feat: expose language detection probabilities to server.cpp * feat: enhance language detection output in server.cpp * Remove empty spaces.	2025-04-28 18:25:45 +02:00
Georgi Gerganov	f3c42399a3	talk-llama : sync llama.cpp (#3084 ) ggml-ci	2025-04-28 16:40:23 +03:00
Pedro	f9b2dfdd8c	examples : fix deprecated FFmpeg functions (#3073 ) * Fix deprecated FFmpeg functions and free packet * avcodec_free_context	2025-04-28 06:16:50 +02:00
Daniel Bevenius	3a88f1e504	examples : add HEAPU8 to exported runtime methods (#3062 ) This commit adds `HEAPU8` to the list of exported methods. The motivation for this commit is that currently this is causing an error on Window systems where HEAPU8 in undefined, which results in the following error message in the web console: ```console main.js:1 Uncaught TypeError: Cannot read properties of undefined (reading 'buffer') at __emval_get_property (main.js:1:1363125) at 003a453a:0xc4a47 at 003a453a:0xc51cd at Object.full_default (eval at craftInvokerFunction (main.js:1:1347011), <anonymous>:9:10) at whisper.cpp/:647:42 ``` Resolves: https://github.com/ggml-org/whisper.cpp/issues/3059	2025-04-20 19:40:25 +02:00
Sacha Arbonel	170b2faf75	whisper : add no_context parameter to whisper_params (#3045 )	2025-04-16 06:24:38 +02:00
Fujimoto Seiji	f8a3509b6d	examples : add FFmpeg v7.0 support to ffmpeg-transcode.cpp (#3038 ) FFmpeg introduced a new channel layout API that uses `AVChannelLayout` interface in v6.0. It subsequently dropped the old bitmask-based API in v7.0. This updates decode_audio() to support the new channel layout API, so that we can compile `whisper-cli` and `whisper-server` with FFmpeg v7.0 or later. Tested on on Ubuntu 24.10 with FFmpeg v7.0.2. Signed-off-by: Fujimoto Seiji <fujimoto@ceptord.net>	2025-04-15 06:09:00 +02:00
Lin Xiaodong	e853620270	addon.node : support max_context api for addon.node (#3025 ) * feat: support max content * feat: show api in test file --------- Co-authored-by: linxiaodong <calm.lin@wukongsch.com>	2025-04-11 06:36:38 +02:00
Georgi Gerganov	2b6d0d2200	rename : ggerganov -> ggml-org (#3005 )	2025-04-04 16:11:52 +03:00
Daniel Bevenius	0b17d4507e	examples : update server.py to match github pages app [no ci] (#3004 ) This commit updates examples/server.py which is used to serve the wasm examples locally. The changes include: - Added a redirect from the root URL to /whisper.cpp. So now accessing http://localhost:8000/ will redirect to http://localhost:8000/whisper.cpp/ which matches the url for the app deployed to github pages. - Custom handling for coi-serviceworker.js to serve it to avoid and error in the console. This file is not strictly necessary for the local server to work as the headers are provided already but it is nice to not have an error in the console. - Fixed the shutdown of the server to ensure it exits cleanly on Ctrl+C. Previously it would continue to hang onto the port even after the processed had exited.	2025-04-04 10:23:53 +02:00
Daniel Bevenius	77e0c86ab6	whisper.wasm : fix unknown language issue (#3000 ) * whisper.wasm : fix unknown language issue This commit addresses an issue with whisper.wasm where the following error was being displayed when running the application in github pages: ``` whisper_lang_id: unknown language 'д=␙c' ``` This turned out to be a memory corruption issue and further details can be found in the reference issue below. Refs: https://github.com/ggerganov/whisper.cpp/issues/2998	2025-04-03 19:50:47 +02:00
Georgi Gerganov	eac1bc9c47	examples : add new sources ggml-ci	2025-04-03 10:30:16 +03:00
Daniel Bevenius	854c0518bc	examples : clarify Core ML encoder model usage [no ci] (#2987 ) This commit clarifies the usage of the Core ML encoder model in the whisper.obj and whisper.swiftui examples. Refs: https://github.com/ggerganov/whisper.cpp/issues/2783	2025-04-02 08:32:14 +02:00
Daniel Bevenius	b358de2458	whisper.objc : fix typo in README.md [no ci] (#2985 ) This commit fixes a typo in the README.md file of the whisper.objc example. Resolves: https://github.com/ggerganov/whisper.cpp/issues/2984	2025-04-02 08:26:57 +02:00
Daniel Bevenius	e153b8eaa2	android.java : re-add ggml source updates (#2975 ) This commit updates the ggml source to include the new unary and binary operations. I merged https://github.com/ggerganov/whisper.cpp/pull/2958 which seems to have overwritten the changes to the ggml source which were added in https://github.com/ggerganov/whisper.cpp/pull/2972. Sorry about this.	2025-03-31 16:14:33 +02:00
Georgi Gerganov	0a40ae9728	android : add new ggml source files ggml-ci	2025-03-31 14:56:53 +03:00
Daniel Bevenius	2d8e40e2a0	examples : update README links to point to pages deployment (#2971 ) This commit updates the README links to point to the pages deployment instead of whisper.ggerganov.com.	2025-03-31 12:32:27 +02:00
Daniel Bevenius	e17af6524f	ci : add github pages workflow for wasm examples (#2969 ) * ci : add github pages workflow for wasm examples This commit adds a github workflow to build and deploy the wasm examples to github pages. The whisper.wasm example is deployed as the main page. This workflow is trigged by a push to master and will deploy the examples to: https://ggerganov.github.io/whisper.cpp/. This requires that the repository has enabled github actions in `Settings` -> `Pages` -> `Build and deployment` -> `Source` be set to `GitHub Actions`. One thing to note is that this commit removes the `talk` example as I'm not sure how this example is built yet. Refs: https://github.com/ggerganov/whisper.cpp/issues/2784	2025-03-31 11:34:40 +02:00
Sacha Arbonel	88d13a17a7	feat: add health check endpoint to server (#2968 )	2025-03-31 11:03:41 +03:00
Lin Xiaodong	1279f0d0bc	examples : support progress_callback API for addon.node (#2941 ) * feat: progress supported * fix: missing params * style: Format the code to improve readability Unified code indentation ensures consistent coding style, enhancing code readability and maintainability. * feat: support prompt api --------- Co-authored-by: linxiaodong <calm.lin@wukongsch.com>	2025-03-28 06:34:26 +01:00
Daniel Bevenius	996581c5e2	whisper.android : add GGML_USE_CPU compile definition (#2945 ) This commit add GGML_USE_CPU to built target library to enable CPU backend. The motivation for this that without the compile definition the CPU backend is not enabled and the app will crash when trying to use it.	2025-03-25 18:01:18 +01:00
Daniel Bevenius	226d344f56	whisper.android.java : update build with ggml source changes (#2942 ) * whisper.android.java : update build with ggml source changes This commit updates the whisper.android.java build to include the new ggml source files and directories. The gradle build configuration is also updated to include the aliyun maven repository.	2025-03-25 16:01:59 +01:00
Daniel Bevenius	30cf30ca82	examples : reduce initial memory to 512MB (#2939 ) * examples : reduce initial memory to 512MB This commit reduces the initial memory size to 512MB. This is done to to avoid WebAssembly memory allocation issues on some platforms. It also adds a flag to allow the memory to grow dynamically (up to the maximum). The motivation for this change is that currently the initial memory is set to 2GB which might be to large for some platforms. This will lead to an error being thrown from the JavaScript code generated by Emscripten when trying to allocate memory. More details can be found in the referenced issue below. * examples : set MAXIMUM_MEMORY instead of TOTAL_MEMORY This commit sets MAXIMUM_MEMORY instead of TOTAL_MEMORY in the whisper.wasm example. The motivation for this is that TOTAL_MEMORY and INITIAL_MEMORY are actually the same thing. Instead we want to set MAXIMUM_MEMORY to 2GB. Refs: https://github.com/ggerganov/whisper.cpp/issues/2920 Refs: https://emscripten.org/docs/tools_reference/settings_reference.html#initial-memory	2025-03-24 14:42:12 +01:00
Daniel Bevenius	ee6286c35d	examples : fix nthread parsing in whisper.wasm (#2938 ) This commit fixes the nthread parsing in the whisper.wasm example when using the `Threads` slider to change the number of threads to be used. Currently this results in the following error: ```console main.js:5597 Uncaught TypeError: Cannot convert "5" to int at checkAssertions (main.js:5597:21) at Object.toWireType (main.js:5611:15) at Object.full_default (eval at new_ (main.js:5292:27), <anonymous>:10:26) at whisper.wasm/:649:42 ```	2025-03-24 14:40:00 +01:00

1 2 3 4 5 ...

518 Commits