Skip to content

llama.cpp : bump version to 0.5.0 - #29333

Merged
ggerganov merged 1 commit into
masterfrom
llama-rc-v0.5.0
Sep 23, 2026
Merged

ggerganov merged 1 commit into
masterfrom
llama-rc-v0.5.0

Conversation

@ggerganov

Copy link
Copy Markdown
Member

Overview

This release focuses on backend performance and correctness, broader model coverage, and more robust server/router operation. It adds HRM-Text (DFM Mimir 1B) support, MiMo-V2.6 and HunyuanOCR conversion support, ggml 0.25.0 backend improvements, multi-address HTTP binding, image outputs from function calls, and several chat parser/UI fixes.

Highlights

  • Accelerate CUDA conv2d with implicit GEMM (#29135)
  • Add Metal MoE and SSM_CONV fusion optimizations (#28948)
  • Allow the server to bind to multiple addresses (#28690)

API changes

  • Add llama_adapter_lora_init_from_file_ptr() for loading LoRA from an open FILE (#28993)
  • Document llama_model_load_from_file_ptr() as reading from the current position and requiring aligned mmap (#28993)
  • Add LLAMA_VOCAB_TYPE_TEST dummy tokenizer (#29084)
  • Add input_image support to server function-call outputs (#22575)
  • Allow --host to accept comma-separated TCP addresses and UNIX sockets (#28690)

New models

  • Add HRM-Text / DFM Mimir 1B support (#27625)
  • Add MiMo-V2.6 conversion support (#29257)
  • Add DFlash support for HunyuanOCR (#28890)
  • Extend Nemotron MTP and Nemotron-H model handling (#29018, #28989)
  • Add Qwen4Exp hyper-connection ops and sparse flash attention (#28901, #28770)
  • Add --fuse-qkv support for Muse Glimmer (#29203)

Core changes

  • Add graph input/input-tensor diagnostics during scheduler reserve (#26625)
  • Enable CUDA graphs for MTP drafting (#28549)
  • Fix tensor-parallel split state/granularity for fused QKV models (#28965)
  • Fix Mamba time-step projection input contiguity (#28832)
  • Write the SWA pattern in the model saver and round-trip 15 more architectures (#29042)
  • Add environment variables for temperature, top-p, min-p and penalties (#27380)
  • Reduce the sampler backend probe size (#29285)
  • Add Ling 3.0, DeepSeek V3.2/V4, qwen3-coder, Muse Glimmer and Gemma 4 parser fixes (#28682, #29008, #28869, #29242, #29115)
  • Improve JSON Schema and PEG handling (#28518, #29127, #29161)
  • Add ufakzeka pre-tokenizer and llama-bench --version (#29033, #28971)

Multi-modality changes

  • Add sanity checks for mtmd layer indices, SAM layer counts, resize targets and graph allocation (#29276, #28149)
  • Fix SigLIP bucket buffer overrun for tall/wide images (#29276)

Server changes

  • Fix router eviction races and child process lifecycle handling (#29217)
  • Do not pass log file or API key file to router-spawned children (#29212, #28938)
  • Improve startup and model-source logging (#29125)
  • Update vendored cpp-httplib to 0.57.1 (#29239)

UI changes

  • Accept WEBM video files (#28622)
  • Add close button to UI toasts (#28246)
  • Fix mobile breakpoint and content overflow issues, including horizontal table scrolling (#29108)
  • Restore the reasoning menu in single-model desktop mode (#27985)
  • Stop re-probing a disabled /tools endpoint on every message (#28646)

ggml changes

  • Updated ggml to v0.25.0 (release)
  • The release expands hyper-connection, flash-attention, and fused MoE/SSM support across backends, with robustness, quantization, data-layout, and RPC/meta improvements.
  • API changes include gated ggml_dsv4_hc_pre_gated(), optional ggml_dsv4_hc_post() comb, and RPC protocol major v7.

@ggerganov
ggerganov merged commit 7fe450e into master Sep 23, 2026
25 checks passed
@ggerganov
ggerganov deleted the llama-rc-v0.5.0 branch September 23, 2026 17:32
@github-actions github-actions Bot added the build Compilation issues label Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

build Compilation issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant