Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
56 commits
Select commit Hold shift + click to select a range
dac869b
conversion : fix Nemotron 3.5 Lightning layers (#27729)
danbev Aug 26, 2026
da9b5d6
ci : make cache bucket public (#27728)
CISC Aug 26, 2026
fc35562
cuda: unblock mmq for MoE on sm_60 (#26264)
dfriehs Aug 26, 2026
4d19b28
ci: Clean up UI builds from releases (#27706)
allozaur Aug 26, 2026
d0132a6
rpc : implement event and async backend APIs (#18626)
rgerganov Aug 26, 2026
bf94216
Implemented vulkan cross_entropy_loss and cross_entropy_loss_back (#2…
PranavUttarkar Aug 26, 2026
5e6a37c
vulkan: warptiles currently assume warp sizes <= 64, clamp to work ar…
0cc4m Aug 26, 2026
0379a19
ui: Update Dialog component styling (#27743)
allozaur Aug 26, 2026
539f245
ui: Move Settings and MCP Servers routes to dialog-based views (#27744)
allozaur Aug 26, 2026
925e117
llama: add token ID tracking to KV cell (#27762)
ngxson Aug 26, 2026
192067b
hexagon: support for multi-NPU devices (IQ9, IQ10) and fully asynchro…
max-krasnyansky Aug 27, 2026
d7a2074
models : support nanbeige4.2-3B (#27730)
zqlcode Aug 27, 2026
c5fc7e3
llama : add --n-cpu-ffn option (#26622)
John-194 Aug 27, 2026
915dc6d
metal : fix memory leaks due to missing autoreleasepools (#27758)
nikwen Aug 27, 2026
f295512
args: add --video-* CLI arguments (#24318)
ngxson Aug 27, 2026
deae5ee
model : simplify MiniMax-01 graph (#27790)
fairydreaming Aug 27, 2026
2bb9bdd
spec: Add benchmark-only synthetic speculative acceptance options (#2…
gaugarg-nv Aug 27, 2026
fe235f4
ui: Replace per-conversation MCP overrides with per-conversation tool…
allozaur Aug 27, 2026
bcb6084
convert : fix Nemotron-H LoRA GGUF conversion (#27356)
frozenblade1224 Aug 27, 2026
cae6357
ui: Improve Chat Form Actions UI/UX (models selector, add panel) (#27…
allozaur Aug 27, 2026
fac889f
llama: model_loader: add TENSOR_READ_LAZY (#27794)
ngxson Aug 27, 2026
1a946ec
pr2wt : use ssh/https remote in worktree depending on base (#27800)
CISC Aug 27, 2026
cb30059
Feature: Added LIGHTNING_INDEXER support for Deepseek V4 ops on Vulka…
shenron0101 Aug 27, 2026
732707d
quantize: cap working memory size to avoid loading big tensors onto R…
ngxson Aug 27, 2026
5854625
opencl: add bin kernels `kernel_gemm_moe_q4_0_q8_1_dp4a_bin`, `kernel…
shawngu-quic Aug 27, 2026
b10f9ca
spec : add DFlash2 support (local convolution + candidate selector) (…
ngxson Aug 27, 2026
6fdd0ac
ci : bundle HIP runtime DLLs with Windows ROCm release (#26973)
slojosic-amd Aug 27, 2026
6c84c7d
model: add Qwen3.8-Flash-Next (qwen4exp) (#27742)
danielhanchen Aug 27, 2026
3217633
ci : build only the ggml-hip backend for windows-rocm release (#27753)
harkgill-amd Aug 27, 2026
1844325
server: add ctx-per-slot (--kv-unified-per-slot) (#24124)
bartowski1182 Aug 27, 2026
83d855c
hex-unary: fix RMS_NORM_MUL weight-offset bugs for grouped/broadcast …
aparmp-quic Aug 27, 2026
e70802a
ggml-hexagon: add HTP unary ops for ABS and LOG (#27786)
cqderek Aug 27, 2026
ca3d5a3
model: add DSpark support for Nemotron3.5 (#27804)
ruixiang63 Aug 27, 2026
4e97ac8
tests : run test-save-load-state across all architectures (#27755)
ggerganov Aug 28, 2026
6d6b697
metal : add fa-vec tunings for M4 Pro (#27824)
infinitewarp Aug 28, 2026
8963a9b
metal : add fa-vec tunings for M3 Max, M5 and M5 Pro (#27863)
ggerganov Aug 28, 2026
be87620
sycl: bind the f16 KV cache in place for the oneDNN SDPA path (#27468)
Titaniumtown Aug 28, 2026
d077b4c
sycl: use TILE for quantized KV decode on BMG (#26689)
johnkarlhill Aug 28, 2026
b19cbe9
convert: prevent ndarray conversion in LazyChunkedTensor (#27869)
ngxson Aug 28, 2026
511f9c1
OpenVINO: Update OV to 2026.3.1, whisper.cpp support, Qwen3.5 on NPU,…
wine99 Aug 28, 2026
f5e85d4
metal : add fa-vec tunings for M4 (#27875)
Strongtut Aug 28, 2026
8663224
context : disable non-fused GDN and LID ops (#27877)
ggerganov Aug 28, 2026
90c26fc
Vulkan: add hoisting support for row IDs and expert count in shaders …
ravel7524 Aug 28, 2026
a43c398
ggml : fix conv_transpose_2d for multiple batches (#26132)
tekinertekin Aug 28, 2026
b387ddf
vulkan: fix missing view-alias dependencies in ggml_vk_graph_optimize…
Eric-A-Stalee Aug 28, 2026
6fe7498
model: qwen4exp: reduce number of graph splits (#27880)
ngxson Aug 28, 2026
50f068f
bench: add --tensor-read-lazy (#27881)
ngxson Aug 28, 2026
d7bd3bf
snapdragon: python SDK setup (Windows) (#27903)
kurquhar Aug 28, 2026
77f132c
vulkan: Change mul_mat_id to pad K rather than N (#27925)
jeffbolznv Aug 29, 2026
5ea1b12
metal : add fa-vec tunings for M1 Max (#27932)
jhen0409 Aug 29, 2026
c9ca51c
vulkan: combine duplicated fastdiv functions, rename the one optimizi…
jeffbolznv Aug 29, 2026
cc83d7b
sycl: make --fit respect --fit-target better (#27629)
nicois Aug 29, 2026
17252c7
metal : add remaining fa-vec tunings for M4 Pro (#27915)
nikwen Aug 29, 2026
3173a56
metal : assert shared memory padding (#27951)
ggerganov Aug 29, 2026
c841aee
opencl: use a better matmul path on two Adreno GPU generations (#27640)
wanghqc Aug 29, 2026
a3577d9
Revert the cudaMemcpy2DAsync fast path from ggml-org#25057
danielhanchen Aug 31, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 6 additions & 6 deletions .devops/openvino.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,12 +1,12 @@
ARG OPENVINO_VERSION_MAJOR=2026.3
ARG OPENVINO_VERSION_FULL=2026.3.0.22451.bd8d6542e3c
ARG OPENVINO_VERSION_MAJOR=2026.3.1
ARG OPENVINO_VERSION_FULL=2026.3.1.22476.56d9685302d
ARG UBUNTU_VERSION=24.04

# Intel GPU driver versions. https://github.com/intel/compute-runtime/releases
ARG IGC_VERSION=v2.38.2
ARG IGC_VERSION_FULL=2_2.38.2+22051
ARG COMPUTE_RUNTIME_VERSION=26.27.39122.11
ARG COMPUTE_RUNTIME_VERSION_FULL=26.27.39122.11-0
ARG IGC_VERSION=v2.40.13
ARG IGC_VERSION_FULL=2_2.40.13+22418
ARG COMPUTE_RUNTIME_VERSION=26.31.39395.13
ARG COMPUTE_RUNTIME_VERSION_FULL=26.31.39395.13-0
ARG IGDGMM_VERSION=22.10.0

# Intel NPU driver versions. https://github.com/intel/linux-npu-driver/releases
Expand Down
50 changes: 27 additions & 23 deletions .github/actions/ccache-buckets/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -58,34 +58,38 @@ runs:
if: ${{ inputs.save == 'true' }}
shell: bash
run: |
set +e -uo pipefail
source .venv-hf/bin/activate
CCACHE_DIR=$(ccache -k cache_dir)
if [[ -d "$CCACHE_DIR" ]]; then
ccache -s
if [[ -n "${{ inputs.evict-old-files }}" ]]; then
ccache --evict-older-than "${{ inputs.evict-old-files }}"
fi
DATESTAMP=$(date -u +'%Y-%m-%dT%H:%M:%SZ')
CACHEFILE="${{ inputs.key }}-$DATESTAMP.tar.gz"
if tar -czf ccache_bucket.tar.gz -C "$CCACHE_DIR" .; then
hf buckets cp ccache_bucket.tar.gz "hf://buckets/${{ inputs.hf_bucket }}/${{ inputs.folder }}/$CACHEFILE"
if [[ -n "$HF_TOKEN" ]]; then
set +e -uo pipefail
source .venv-hf/bin/activate
CCACHE_DIR=$(ccache -k cache_dir)
if [[ -d "$CCACHE_DIR" ]]; then
ccache -s
if [[ -n "${{ inputs.evict-old-files }}" ]]; then
ccache --evict-older-than "${{ inputs.evict-old-files }}"
fi
DATESTAMP=$(date -u +'%Y-%m-%dT%H:%M:%SZ')
CACHEFILE="${{ inputs.key }}-$DATESTAMP.tar.gz"
if tar -czf ccache_bucket.tar.gz -C "$CCACHE_DIR" .; then
hf buckets cp ccache_bucket.tar.gz "hf://buckets/${{ inputs.hf_bucket }}/${{ inputs.folder }}/$CACHEFILE"
fi
rm ccache_bucket.tar.gz
else
echo "'$CCACHE_DIR' not found."
fi
rm ccache_bucket.tar.gz
else
echo "'$CCACHE_DIR' not found."
fi

- name: Remove old ccache files from buckets
if: ${{ inputs.save == 'true' }}
shell: bash
run: |
set +e -uo pipefail
source .venv-hf/bin/activate
CACHE_FILES=$(hf buckets list "hf://buckets/${{ inputs.hf_bucket }}/${{ inputs.folder }}" --json | jq -r '[.[] | select(.type == "file") | select((.uploaded_at | .[:19]+"Z" | fromdateiso8601) < (now - 5 * 60)) | select(.path | startswith("${{ inputs.folder }}/${{ inputs.key }}") and endswith(".tar.gz"))] | sort_by(.path)[:-1] | .[] | [.path // ""] | @tsv')
if [[ -n "$CACHE_FILES" ]]; then
echo "Removing old ccache files..."
while IFS=$'\t' read -r CACHE_PATH; do
hf buckets rm "hf://buckets/${{ inputs.hf_bucket }}/$CACHE_PATH" -y
done <<< "$CACHE_FILES"
if [[ -n "$HF_TOKEN" ]]; then
set +e -uo pipefail
source .venv-hf/bin/activate
CACHE_FILES=$(hf buckets list "hf://buckets/${{ inputs.hf_bucket }}/${{ inputs.folder }}" --json | jq -r '[.[] | select(.type == "file") | select((.uploaded_at | .[:19]+"Z" | fromdateiso8601) < (now - 5 * 60)) | select(.path | startswith("${{ inputs.folder }}/${{ inputs.key }}") and endswith(".tar.gz"))] | sort_by(.path)[:-1] | .[] | [.path // ""] | @tsv')
if [[ -n "$CACHE_FILES" ]]; then
echo "Removing old ccache files..."
while IFS=$'\t' read -r CACHE_PATH; do
hf buckets rm "hf://buckets/${{ inputs.hf_bucket }}/$CACHE_PATH" -y
done <<< "$CACHE_FILES"
fi
fi
8 changes: 4 additions & 4 deletions .github/workflows/build-cache.yml
Original file line number Diff line number Diff line change
Expand Up @@ -41,8 +41,8 @@ jobs:

env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.3"
OPENVINO_VERSION_FULL: "2026.3.0.22451.bd8d6542e3c"
OPENVINO_VERSION_MAJOR: "2026.3.1"
OPENVINO_VERSION_FULL: "2026.3.1.22476.56d9685302d"

steps:
- name: Clone
Expand All @@ -69,8 +69,8 @@ jobs:

env:
# Sync versions in build.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.3"
OPENVINO_VERSION_FULL: "2026.3.0.22451.bd8d6542e3c"
OPENVINO_VERSION_MAJOR: "2026.3.1"
OPENVINO_VERSION_FULL: "2026.3.1.22476.56d9685302d"

steps:
- name: Clone
Expand Down
12 changes: 6 additions & 6 deletions .github/workflows/build-cuda-ubuntu.yml
Original file line number Diff line number Diff line change
Expand Up @@ -65,7 +65,7 @@ jobs:
with:
key: cuda-ubuntu-24.04-cuda
folder: llama.cpp
hf_bucket: ${{ vars.HF_BUCKET_CACHE_OUTPUT }}
hf_bucket: ggml-org/cache

- name: Build with CMake
# TODO: Remove GGML_CUDA_CUB_3DOT2 flag once CCCL 3.2 is bundled within CTK and that CTK version is used in this project
Expand All @@ -89,7 +89,7 @@ jobs:
key: cuda-ubuntu-24.04-cuda
folder: llama.cpp
evict-old-files: 1d
hf_bucket: ${{ vars.HF_BUCKET_CACHE_OUTPUT }}
hf_bucket: ggml-org/cache
save: true

hip:
Expand Down Expand Up @@ -120,7 +120,7 @@ jobs:
with:
key: cuda-ubuntu-22.04-hip
folder: llama.cpp
hf_bucket: ${{ vars.HF_BUCKET_CACHE_OUTPUT }}
hf_bucket: ggml-org/cache

- name: Build with native CMake HIP support
id: cmake_build
Expand All @@ -140,7 +140,7 @@ jobs:
key: cuda-ubuntu-22.04-hip
folder: llama.cpp
evict-old-files: 1d
hf_bucket: ${{ vars.HF_BUCKET_CACHE_OUTPUT }}
hf_bucket: ggml-org/cache
save: true

musa:
Expand Down Expand Up @@ -171,7 +171,7 @@ jobs:
with:
key: cuda-ubuntu-22.04-musa
folder: llama.cpp
hf_bucket: ${{ vars.HF_BUCKET_CACHE_OUTPUT }}
hf_bucket: ggml-org/cache

- name: Build with native CMake MUSA support
id: cmake_build
Expand All @@ -189,5 +189,5 @@ jobs:
key: cuda-ubuntu-22.04-musa
folder: llama.cpp
evict-old-files: 1d
hf_bucket: ${{ vars.HF_BUCKET_CACHE_OUTPUT }}
hf_bucket: ggml-org/cache
save: true
19 changes: 9 additions & 10 deletions .github/workflows/build-openvino.yml
Original file line number Diff line number Diff line change
Expand Up @@ -32,15 +32,17 @@ env:
LLAMA_ARG_LOG_COLORS: 1
LLAMA_ARG_LOG_PREFIX: 1
LLAMA_ARG_LOG_TIMESTAMPS: 1
# TODO: fix and re-enable the `test-llama-archs` and `test-recurrent-state-rollback`
CTEST_EXCLUDE: "test-llama-archs|^test-recurrent-state-rollback"

jobs:
ubuntu-24-openvino:
runs-on: [self-hosted, Linux, Intel, OpenVINO]

env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.3"
OPENVINO_VERSION_FULL: "2026.3.0.22451.bd8d6542e3c"
OPENVINO_VERSION_MAJOR: "2026.3.1"
OPENVINO_VERSION_FULL: "2026.3.1.22476.56d9685302d"

steps:
- name: Clone
Expand Down Expand Up @@ -78,26 +80,24 @@ jobs:
- name: Test (CPU)
id: cmake_test_cpu
# TODO: fix and re-enable the `test-llama-archs` test below
run: |
cd ${{ github.workspace }}
ctest --test-dir build/ReleaseOV -L main -E "test-llama-archs|test-recurrent-state-rollback-nemotron-h" --verbose --timeout 2000
ctest --test-dir build/ReleaseOV -L main -E "${{ env.CTEST_EXCLUDE }}" --verbose --timeout 3000
- name: Test (GPU)
id: cmake_test_gpu
# TODO: fix and re-enable the `test-llama-archs` test below
run: |
cd ${{ github.workspace }}
export GGML_OPENVINO_DEVICE=GPU
ctest --test-dir build/ReleaseOV -L main -E "test-llama-archs|test-recurrent-state-rollback-nemotron-h" --verbose --timeout 3000
ctest --test-dir build/ReleaseOV -L main -E "${{ env.CTEST_EXCLUDE }}" --verbose --timeout 3000
openvino-windows-2022:
runs-on: windows-2022

env:
# Sync versions in build-openvino.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.3"
OPENVINO_VERSION_FULL: "2026.3.0.22451.bd8d6542e3c"
OPENVINO_VERSION_MAJOR: "2026.3.1"
OPENVINO_VERSION_FULL: "2026.3.1.22476.56d9685302d"

steps:
- name: Clone
Expand Down Expand Up @@ -159,14 +159,13 @@ jobs:
- name: Test (CPU)
id: cmake_test_cpu
shell: cmd
# TODO: fix and re-enable the `test-llama-archs` test below
run: |
REM Find extracted OpenVINO folder dynamically
for /d %%i in (openvino_toolkit\*) do set OPENVINO_ROOT=%%i
call "%OPENVINO_ROOT%\setupvars.bat"
cd build
ctest --test-dir ReleaseOV -L main -E "test-llama-archs|test-recurrent-state-rollback-nemotron-h" -C Release --verbose --timeout 3000
ctest --test-dir ReleaseOV -L main -E "${{ env.CTEST_EXCLUDE }}" -C Release --verbose --timeout 3000
- name: ccache-clear
uses: ./.github/actions/ccache-clear
Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/build-self-hosted.yml
Original file line number Diff line number Diff line change
Expand Up @@ -288,8 +288,8 @@ jobs:

env:
# Sync versions in build.yml, build-self-hosted.yml, release.yml, build-cache.yml, .devops/openvino.Dockerfile
OPENVINO_VERSION_MAJOR: "2026.3"
OPENVINO_VERSION_FULL: "2026.3.0.22451.bd8d6542e3c"
OPENVINO_VERSION_MAJOR: "2026.3.1"
OPENVINO_VERSION_FULL: "2026.3.1.22476.56d9685302d"

steps:
- name: Clone
Expand Down
4 changes: 2 additions & 2 deletions .github/workflows/docker.yml
Original file line number Diff line number Diff line change
Expand Up @@ -64,7 +64,7 @@ jobs:
needs: create_tag
uses: ./.github/workflows/ui-build.yml
with:
hf_ui_version: ${{ needs.create_tag.outputs.source_tag }}
ui_version: ${{ needs.create_tag.outputs.source_tag }}

prepare_matrices:
name: Prepare Docker matrices
Expand Down Expand Up @@ -162,7 +162,7 @@ jobs:
if: ${{ matrix.config.prebuilt_ui == true }}
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8
with:
name: ui-build
name: llama-ui.zip
path: tools/ui/dist

- name: Set up QEMU
Expand Down
Loading
Loading