whisper.cpp

mirror of https://github.com/ggerganov/whisper.cpp.git synced 2025-08-03 09:09:32 +02:00

Author	SHA1	Message	Date
Ewan Crawford	783cf0309f	SYCL: Bump oneMath commit (llama/14152) Update oneMath commit to merged PR https://github.com/uxlfoundation/oneMath/pull/669 which adds SYCL-Graph support for recording CUDA BLAS commands. With this change the `MUL_MAT` tests now pass on DPC++ CUDA backends with SYCL-Graph enabled. Prior to this change, an error would be thrown. ``` $ GGML_SYCL_DISABLE_GRAPH=0 ./bin/test-backend-ops -b SYCL0 -o MUL_MAT -p type_a=f16,type_b=f32,m=16,n=1,k=256,bs=\\[1,1\\],nr=\\[2 UR CUDA ERROR: Value: 700 Name: CUDA_ERROR_ILLEGAL_ADDRESS Description: an illegal memory access was encountered Function: operator() Source Location: $HOME/dpcpp/unified-runtime/source/adapters/cuda/queue.cpp:154 Native API failed. Native API returns: 2147483646 (UR_RESULT_ERROR_UNKNOWN) Exception caught at file:$HOME/llama.cpp/ggml/src/ggml-sycl/ggml-sycl.cpp, line:3598, func:operator() SYCL error: CHECK_TRY_ERROR((stream)->wait()): Meet error in this line code! in function ggml_backend_sycl_synchronize at $HOME/llama.cpp/ggml/src/ggml-sycl/ggml-sycl.cpp:3598 $HOME/llama.cpp/ggml/src/ggml-sycl/../ggml-sycl/common.hpp:118: SYCL error Could not attach to process. If your uid matches the uid of the target process, check the setting of /proc/sys/kernel/yama/ptrace_scope, or try again as the root user. For more details, see /etc/sysctl.d/10-ptrace.conf ptrace: Operation not permitted. No stack. The program is not being run. ```	2025-06-18 12:40:34 +03:00
Anton Mitkov	0097eaf839	sycl: Remove not needed copy f16->f32 for dnnl mul mat (llama/14125)	2025-06-18 12:40:34 +03:00
Georgi Gerganov	a96a880f7b	cmake : handle whitepsaces in path during metal build (llama/14126) * cmake : handle whitepsaces in path during metal build ggml-ci * cont : proper fix ggml-ci --------- Co-authored-by: Daniel Bevenius <daniel.bevenius@gmail.com>	2025-06-18 12:40:34 +03:00
Christian Kastner	26c16ad6bd	Implement GGML_CPU_ALL_VARIANTS for ARM (llama/14080) * ggml-cpu: Factor out feature detection build from x86 * ggml-cpu: Add ARM feature detection and scoring This is analogous to cpu-feats-x86.cpp. However, to detect compile-time activation of features, we rely on GGML_USE_<FEAT> which need to be set in cmake, instead of GGML_<FEAT> that users would set for x86. This is because on ARM, users specify features with GGML_CPU_ARM_ARCH, rather than with individual flags. * ggml-cpu: Implement GGML_CPU_ALL_VARIANTS for ARM Like x86, however to pass around arch flags within cmake, we use GGML_INTERNAL_<FEAT> as we don't have GGML_<FEAT>. Some features are optional, so we may need to build multiple backends per arch version (armv8.2_1, armv8.2_2, ...), and let the scoring function sort out which one can be used. * ggml-cpu: Limit ARM GGML_CPU_ALL_VARIANTS to Linux for now The other platforms will need their own specific variants. This also fixes the bug that the the variant-building branch was always being executed as the else-branch of GGML_NATIVE=OFF. The branch is moved to an elseif-branch which restores the previous behavior.	2025-06-18 12:40:34 +03:00
Jeff Bolz	40d0d47cf1	vulkan: Better thread-safety for command pools/buffers (llama/14116) This change moves the command pool/buffer tracking into a vk_command_pool structure. There are two instances per context (for compute+transfer) and two instances per device for operations that don't go through a context. This should prevent separate contexts from stomping on each other.	2025-06-18 12:40:34 +03:00
Jeff Bolz	40c6525517	vulkan: Track descriptor pools/sets per-context (llama/14109) Use the same descriptor set layout for all pipelines (MAX_PARAMETER_COUNT == 8) and move it to the vk_device. Move all the descriptor pool and set tracking to the context - none of it is specific to pipelines anymore. It has a single vector of pools and vector of sets, and a single counter to track requests and a single counter to track use.	2025-06-18 12:40:34 +03:00
lhez	74c68067dc	opencl: add `mul_mv_id_q4_0_f32_8x_flat` (llama/14003)	2025-06-18 12:40:34 +03:00
0cc4m	794bf23994	Vulkan: Don't default to CPU device (like llvmpipe), even if no other device is available, to allow fallback to CPU backend (llama/14099)	2025-06-18 12:40:34 +03:00
Isaac McFadyen	26dcc196c7	rpc : nicer error messages for RPC server crash (llama/14076)	2025-06-18 12:40:34 +03:00
Daniel Bevenius	ffe5400d1b	ggml : disable warnings for tests when using MSVC (ggml/1273) * ggml : disable warnings for tests when using MSVC This commit disables warnings for tests on windows when using MSVC. The motivation for this is that this brings the build output more inline with what Linux/MacOS systems produce. There is still one warning generated for the tests which is: ```console Building Custom Rule C:/ggml/tests/CMakeLists.txt cl : command line warning D9025: overriding '/DNDEBUG' with '/UNDEBUG' [C:\ggml\build\tests\test-arange.vcxproj] test-arange.cpp test-arange.vcxproj -> C:\ggml\build\bin\Release\test-arange.exe ``` * ggml : fix typo in tests disable list	2025-06-18 12:40:34 +03:00
Daniel Bevenius	1b01c0cc4e	ggml : remove unused ggml_context_container (ggml/1272) This commit removes the unused `ggml_context_container` structure from the ggml library. It looks like the usage of this struct was removed in Commit 4757fe18d56ec11bf9c07feaca6e9d5b5357e7f4 ("ggml : alloc ggml_contexts on the heap (whisper/2525)"). The motivation for this changes is to improve code clarity/readability.	2025-06-18 12:40:34 +03:00
Daniel Bevenius	db30f46761	examples : include examples in msvc disable warn (ggml/1270) This commit adds the examples in the "list" of targets to ignore MSVC warnings. The motivation for this is that currently the examples generate a number of warnings that are ignore/disabled for the core ggml project. This makes for a cleaner output when building.	2025-06-18 12:40:34 +03:00
Daniel Bevenius	1591558ccc	whisper : clear result_all if vad_samples is empty (#3262 ) This commit clears the results_all vector no VAD segments are found. The motivation for this is that this would normally be done by `whisper_full_with_state` but when no VAD segments are detected this current implementation does not call that function and hence the vector does not get reset. This can lead to issues in applications like the server example where it will incorrectly process the old results. Resolves: https://github.com/ggml-org/whisper.cpp/issues/3250	2025-06-18 11:30:29 +02:00
Daniel Bevenius	f3ff80ea8d	examples : set the C++ standard to C++17 for server (#3261 ) This commit updates the server example to use C++17 as the standard. The motivation for this change is that currently the ci-run `ggml-100-mac-m4` is failing when compiling the server example on macOS. The `talk-llama` example also has this setting so it looks like an alright change to make. ggml-ci Refs: https://github.com/ggml-org/ci/tree/results/whisper.cpp/2a/4d6db7d90899aff3d58d70996916968e4e0d27/ggml-100-mac-m4	2025-06-17 11:29:48 +02:00
w1redch4d	2a4d6db7d9	examples : update usage/help in yt-wsp.sh (#3251 ) This commit updates the usage/help message to be more readable and include the environment variables available to set options.	2025-06-16 12:21:16 +02:00
Sacha Arbonel	107c303e69	server : graceful shutdown, atomic server state, and health endpoint Improvements (#3243 ) * feat(server): implement graceful shutdown and server state management * refactor(server): use lambda capture by reference in server.cpp	2025-06-16 10:14:26 +02:00
Daniel Bevenius	705db0f728	whisper : fix VAD processing for skipped audio segments (#3230 ) This commit addresses an issue with token timestamps when audio segments are skipped, in `whisper_exp_compute_token_level_timestamps` related to the VAD processing and the energy levels. The motivation for this is that the token timestamps exceed the energy array bounds due to segment timing misalignment: ```console (skipped introduction) ↓ Audio segment: [2600ms → 5600ms] (3 seconds of actual audio) Energy array: [0 → 480652] (samples for 3 seconds) Token timestamps: [3266ms → 3408ms] (absolute timestamps) ``` So both `s0` and `t1` get clamped to the maximum sample index (480652) which causes the start/end timestamps to be the same for all the tokens after a certain point. This is addressed by using segment-relative timestamps in the `timestamp_to_sample` and `sample_to_timestamp`.	2025-06-13 17:35:52 +02:00
Daniel Bevenius	0a4d85cf8a	server : add Voice Activity Detection (VAD) support (#3246 ) * server : add Voice Activity Detection (VAD) support This commit adds support for Voice Activity Detection (VAD) in the server example. The motivation for this is to enable VAD processing when using whisper-server. Resolves: https://github.com/ggml-org/whisper.cpp/issues/3089 * server : add VAD parameters to usage in README.md [no ci] This commit also adds a few missing parameters. * server : fix conflicting short options [no ci]	2025-06-13 13:24:03 +02:00
Daniel Bevenius	9df8d54bcb	cli : fix short name conflict for vad options [no ci] (#3247 ) This commit fixes a short name conflict whisper-cli for `--vad-min-speech-duration-ms` and `--vad-min-silence-duration-ms` which currently have the same short name `-vsd`. Refs: https://github.com/ggml-org/whisper.cpp/pull/3246#pullrequestreview-2923800114	2025-06-13 10:25:25 +02:00
Daniel Bevenius	20d203aacf	ruby : add .gitignore entries for ext directory (#3245 ) This commit adds entries to `.gitignore` for directories in the `ext` directory. The motivation for this is that currently after building locally these following files are reported by git as untracked: ```console Untracked files: (use "git add <file>..." to include in what will be committed) ext/examples/ ext/ggml/ ext/include/ ext/scripts/ ext/src/ ```	2025-06-13 10:04:20 +02:00
Daniel Bevenius	ebbc874e85	ci : update windows runner to windows-2022 (#3242 ) * ci : update windows runner to windows-2022 This commit changes the windows-2019 runner to windows-2022. The motiation for this is that the windows-2019 runner is scheduled for deprection and will be removed 2025-06-30. There are currently "burnout" periods that started 2025-06-01 and during these times jobs with windows-2019 will fail which has happened lately on our CI. Refs: https://github.com/actions/runner-images/issues/12045	2025-06-11 13:53:16 +02:00
Daniel Bevenius	2679bec6e0	ruby : add cleaning of library names in dependencies (#3241 ) * ruby : add cleaning of library names in dependencies This commit adds a cleaning step to the library names in the `Dependencies` class of the Ruby bindings. The motivation for this is that with the introduction of a library name alias for ggml in Commit (`b933d17c30` "Add in-build ggml::ggml ALIAS library (ggml/1260)) causes the Makefile generation to break: ```console $ sed -n '165,170p' ext/Makefile CLEANOBJS = $(OBJS) .bak TARGET_SO_DIR_TIMESTAMP = $(TIMESTAMP_DIR)/.sitearchdir.time $(TARGET_SO): libcommon.a libwhisper.a libggml\n(ggml::ggml).a libggml-cpu.a libggml-base.a libcommon.a libwhisper.a libggml\n(ggml::ggml).a libggml-cpu.a libggml-base.a: cmake-targets cmake-targets: /usr/bin/cmake -S sources -B build -D BUILD_SHARED_LIBS=OFF -D CMAKE_ARCHIVE_OUTPUT_DIRECTORY=/home/danbev/work/ai/whisper.cpp/bindings/ruby/ext -D CMAKE_POSITION_INDEPENDENT_CODE=ON ``` squash! ruby : add cleaning of library names in dependencies Apply PR review feedback.	2025-06-10 15:06:40 +02:00
Georgi Gerganov	93d543905e	ggml : fix weak alias win32 (#0 ) ggml-ci	2025-06-10 12:40:33 +03:00
Georgi Gerganov	962361bd79	android : fix builds (#0 ) ggml-ci	2025-06-10 12:40:33 +03:00
Georgi Gerganov	dbe81c1042	sync : ggml ggml-ci	2025-06-10 12:40:33 +03:00
Georgi Gerganov	175e7e4f1a	files : remove old sources (part 2)	2025-06-10 12:40:33 +03:00
Georgi Gerganov	56475d01dc	sync : ggml ggml-ci	2025-06-10 12:40:33 +03:00
Georgi Gerganov	38347a7dda	files : remove old sources	2025-06-10 12:40:33 +03:00
Georgi Gerganov	db264d6220	talk-llama : sync llama.cpp ggml-ci	2025-06-10 12:40:33 +03:00
Georgi Gerganov	96eaf46ec6	sync : ggml ggml-ci	2025-06-10 12:40:33 +03:00
Georgi Gerganov	7a675807a2	metal : use less stack memory in FA kernel (llama/14088) * metal : use less stack memory in FA kernel ggml-ci * cont : fix BF16 variant	2025-06-10 12:40:33 +03:00
xctan	8cbc889f85	ggml-cpu : split arch-specific implementations (llama/13892) * move ggml-cpu-aarch64 to repack * split quantize_row_q8_0/1 * split helper functions * split ggml_vec_dot_q4_0_q8_0 * split ggml_vec_dot_q4_1_q8_1 * split ggml_vec_dot_q5_0_q8_0 * split ggml_vec_dot_q5_1_q8_1 * split ggml_vec_dot_q8_0_q8_0 * split ggml_vec_dot_tq1_0_q8_K * split ggml_vec_dot_tq2_0_q8_K * split ggml_vec_dot_q2_K_q8_K * split ggml_vec_dot_q3_K_q8_K * split ggml_vec_dot_q4_K_q8_K * split ggml_vec_dot_q5_K_q8_K * split ggml_vec_dot_q6_K_q8_K * split ggml_vec_dot_iq2_xxs_q8_K * split ggml_vec_dot_iq2_xs_q8_K * split ggml_vec_dot_iq2_s_q8_K * split ggml_vec_dot_iq3_xxs_q8_K * split ggml_vec_dot_iq3_s_q8_K * split ggml_vec_dot_iq1_s_q8_K * split ggml_vec_dot_iq1_m_q8_K * split ggml_vec_dot_iq4_nl_q8_0 * split ggml_vec_dot_iq4_xs_q8_K * fix typos * fix missing prototypes * rename ggml-cpu-quants.c * rename ggml-cpu-traits * rename arm folder * move cpu-feats-x86.cpp * rename ggml-cpu-hbm * update arm detection macro in quants.c * move iq quant tables * split ggml_quantize_mat_q8_0/K * split ggml_gemv_* * split ggml_gemm_* * rename namespace aarch64 to repack * use weak aliases to replace test macros * rename GGML_CPU_AARCH64 to GGML_CPU_REPACK * rename more aarch64 to repack * clean up rebase leftover * fix compilation errors * remove trailing spaces * try to fix clang compilation errors * try to fix clang compilation errors again * try to fix clang compilation errors, 3rd attempt * try to fix clang compilation errors, 4th attempt * try to fix clang compilation errors, 5th attempt * try to fix clang compilation errors, 6th attempt * try to fix clang compilation errors, 7th attempt * try to fix clang compilation errors, 8th attempt * try to fix clang compilation errors, 9th attempt * more cleanup * fix compilation errors * fix apple targets * fix a typo in arm version of ggml_vec_dot_q4_K_q8_K Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-06-10 12:40:33 +03:00
Diego Devesa	e16a84cd95	cuda : fix device sync on buffer clear (llama/14033)	2025-06-10 12:40:33 +03:00
Xinpeng Dou	26282282fa	CANN: Simplify the environment variable setting(#13104 ) * Simplify the environment variable setting to specify the memory pool type. * Adjust the GGML_CANN_ASYNC_MODE setting to accept yes, enable, 1, or on (case-insensitive) as valid options. * update * fix CI * update * delete whitespace * fix according to review * update CANN.md * update CANN.md	2025-06-10 12:40:33 +03:00
Nicolò Scipione	4737a8c780	sycl: Add reorder to Q6_K mmvq implementation (llama/13885) * Add Reorder to Q6_K mmvq implementation * Address PR comments: clean up comments * Remove unused parameter after refactoring q4_k * Adding inline to function and removing unnecessary reference to int --------- Signed-off-by: nscipione <nicolo.scipione@codeplay.com>	2025-06-10 12:40:33 +03:00
Diego Devesa	8a70f4d18b	cuda : fix buffer type check with integrated GPUs (llama/14069)	2025-06-10 12:40:33 +03:00
Akarshan Biswas	489dc158a6	SYCL: Implement few same quantized type copy kernels (llama/13739) * SYCL: Implement few same quantized type copy kernels * Use memcpy for copying contiguous tensors ggml-ci * feat(sycl): add contiguous tensor copy support and device checks Adds a memcpy path for contiguous tensors of the same type to optimize data transfer. Updates device support checks to recognize contiguous tensor operations, improving compatibility and performance. * refactor: replace specific block copy functions with template The changes replace multiple redundant block copy functions (e.g., cpy_block_q8_0_q8_0, cpy_block_q5_0_q5_0) with a single templated function cpy_blck_q_q. This reduces code duplication by using a generic template that works for any block type, improving maintainability while preserving the same functionality. The template is instantiated with specific block types (e.g., block_q8_0) where needed. * Exclude BF16 support for COPY tensors for now ggml-ci * perf: adjust SYCL copy kernel block sizes for efficiency Use ceil_div to ensure full element coverage and update nd_range parameters to better align with SYCL block sizes, improving parallelism and device utilization in copy operations.	2025-06-10 12:40:33 +03:00
Masato Nakasaka	f0f5a9f7fb	vulkan: Enable VK_KHR_cooperative_matrix extension for Intel Xe2 GPUs (llama/14001) * allowing B580 and U9-288V * experimenting code to detect Xe2 * allowing coopmat only for Xe2 GPUs * fixed comment wording * fixed comment wording * removed unnecessary driver check	2025-06-10 12:40:33 +03:00
Diego Devesa	13a03c5d33	llama : allow using mmap without PrefetchVirtualMemory, apply GGML_WIN_VER to llama.cpp sources (llama/14013)	2025-06-10 12:40:33 +03:00
Jeff Bolz	6dd91d4f7e	vulkan: automatically deduce size of push constants (llama/13936)	2025-06-10 12:40:33 +03:00
Ervin Áron Tasnádi	5171b24f70	ggml-vulkan: adds support for op CONV_TRANSPOSE_1D (llama/13813) * * ggml-vulkan: adds op CONV_TRANSPOSE_1D * test-backend-ops: adds more spohisticated tests for CONV_TRANSPOSE_1D * Missing barrier added to shader. Number of additional tests reduced to 108. * * Fixes typo in variable name. * Removes extra whitespaces. * Adds int64->int32 casts to prevent possible warnings. * Problem size reduced in tests to pass tests with llvmpipe. * supports_op condition moved from unintended position	2025-06-10 12:40:33 +03:00
Diego Devesa	23e2fe0682	releases : use dl backend for linux release, remove arm64 linux release (llama/13996)	2025-06-10 12:40:33 +03:00
Johannes Gäßler	7f4d110f53	CUDA: fix FTZ in FA for Gemma 3 (llama/13991)	2025-06-10 12:40:33 +03:00
Jeff Bolz	ee0ef39fee	vulkan: fix warnings in perf logger querypool code (llama/13937)	2025-06-10 12:40:33 +03:00
lhez	62791ba2e6	opencl: add `backend_synchronize` (llama/13939) * This is not needed by the normal use where the result is read using `tensor_get`, but it allows perf mode of `test-backend-ops` to properly measure performance.	2025-06-10 12:40:33 +03:00
rmatif	e16ef08884	OpenCL: Add concat, tsembd, upscale, tanh, pad and repeat (llama/13840) * add concat, pad, repeat, tsembd, tanh, upscale * small fixes	2025-06-10 12:40:33 +03:00
Georgi Gerganov	c72d3ce935	metal : use F32 accumulators in FA kernels (llama/13975) ggml-ci	2025-06-10 12:40:33 +03:00
shalinib-ibm	126aeb4a49	cmake : Handle mixed-case 'Power' strings in POWER CPU detection (llama/13966) Some systems report the CPU implementation as "Power11" instead of "POWER11". The existing CMake logic uses a case-sensitive regular expression to extract the CPU generation, which fails when the casing doesn't exactly match "POWER". This patch provides a fix by first converting the string to uppercase before applying the regex. Signed-off-by: root <root@rheldb2v.pperf.tadn.ibm.com> Co-authored-by: root <root@rheldb2v.pperf.tadn.ibm.com>	2025-06-10 12:40:33 +03:00
Atharva Dubey	ef2a79d2b8	sycl: quantize and reorder the input to q8_1 when reorder is enabled (llama/13826) * [WIP]: fuse q8 quantization and reorder * wip2: fuse q8 quantization and reorder * working q8 reorder commit * restored common.hpp * remove debug prints * remove unnecessary headers and remove trailing whitespace * Update ggml/src/ggml-sycl/ggml-sycl.cpp Co-authored-by: Alberto Cabrera Pérez <alberto.cabrera@intel.com> --------- Co-authored-by: Alberto Cabrera Pérez <alberto.cabrera@intel.com>	2025-06-10 12:40:33 +03:00
Johannes Gäßler	9589645e72	gguf: fix failure on version == 0 (llama/13956)	2025-06-10 12:40:33 +03:00

1 2 3 4 5 ...

2762 Commits