diff --git a/CHANGELOG.md b/CHANGELOG.md index 597d7ec..8f2e564 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,9 +6,10 @@ DGX Spark (GB10) support: a silent ARM code-generation bug fixed, the ggml RPC backend wired up so a model can span two machines, and a measured runbook for both configurations in [docs/dgx-spark.md](docs/dgx-spark.md). -llama.cpp bumped to [`b10582`](https://github.com/ggml-org/llama.cpp/releases/tag/b10582) -(`e85caa81e`), by way of b10435, which brought Qwen 3.8 in under the existing -`qwen35` architecture, and MTP support for its target/sidecar split (see Added). +llama.cpp bumped to [`b10665`](https://github.com/ggml-org/llama.cpp/releases/tag/b10665) +(`ca3d5a3e1`), by way of b10435 and b10582, which brought Qwen 3.8 in under the +existing `qwen35` architecture, and MTP support for its target/sidecar split +(see Added). Verified on macOS (Metal) at `e85caa81e`, running **every** tag the suite excludes by default. Default build: **428 passed, 149 excluded** with no model; @@ -21,6 +22,10 @@ for `--include mtp_sidecar` (Qwen3.8-27B-Q4_K_M plus its `mtp-*-Q4_0` head). re-run on that build give the same **569** and **434**. On a DGX Spark (CUDA 13.0, `sm_121a`), measured at b10435: **428 passed**. +Re-verified at `ca3d5a3e1` (b10665) on macOS (Metal): default build +**428 passed, 149 excluded** with no model; **434 passed** for +`--include mtp_sidecar` (Qwen3.8-27B-Q4_K_M plus its `mtp-*-Q4_0` head). + The one tag that is not green is `:mtp_cancel`, and it moved: see Changed. ### Fixed @@ -179,6 +184,20 @@ The one tag that is not green is `:mtp_cancel`, and it moved: see Changed. tensor-split work was reverted in `f20395dae`), and the `ggml-cpu` CMake diff is OpenMP target variables, KleidiAI SME2 GEMV sources and IntelLLVM fast-math gating — nothing near the `-mcpu=native` probe. +- **llama.cpp bumped to `ca3d5a3e1`** (b10665), 83 builds past b10582, and + `LLAMA_COMMIT` moved with the submodule. One binding edit: upstream replaced + `nlohmann::ordered_json` with its own pimpl wrapper `common_json` across + `common/chat.h` and `common/json-schema-to-grammar.h` (`common/json.h`), so + `json_schema_to_grammar_nif` now parses with `common_json::parse` and the + direct `` include is gone — nlohmann is still the backing + parser, just behind upstream's wrapper, and the NIF's size/depth bounds on + untrusted schema text are unchanged. `include/llama.h` changes are additive + (`llama_tensor_read_lazy` in `llama_model_params`, session/state-seq version + bumps 9→10 / 2→3) and the NIF builds its params from + `llama_model_default_params()`, so nothing else moved. + `common/speculative.h` only gained functions (`common_speculative_n_max`, + synthetic-acceptance helpers); every `common_speculative_*` call the NIF + makes is signature-identical. - **The `:mtp_cancel` bug no longer aborts the VM — it returns an error.** The race is unchanged and unfixed: cancellation is fire-and-forget, so reusing an `%MTP{}` session immediately after halting a stream can start decoding on diff --git a/Makefile b/Makefile index 3ffe0f2..7a8518c 100644 --- a/Makefile +++ b/Makefile @@ -36,7 +36,7 @@ endif # Pinned llama.cpp commit, used when vendor/llama.cpp has to be cloned. MUST # match the vendor/llama.cpp submodule; bump both together, see # docs/release-guide.md. Override to build the NIF against another revision. -LLAMA_COMMIT ?= e85caa81ea2b65797396018c179b87ad61fa38ab +LLAMA_COMMIT ?= ca3d5a3e10d53f7ea672cb9b6178faca3e2807bc # The commit actually on disk. A submodule can be bumped without LLAMA_COMMIT # following it, and the build has to key off what is really there. diff --git a/c_src/llama_cpp_ex/llama_nif.cpp b/c_src/llama_cpp_ex/llama_nif.cpp index cbcf932..46343d2 100644 --- a/c_src/llama_cpp_ex/llama_nif.cpp +++ b/c_src/llama_cpp_ex/llama_nif.cpp @@ -2,7 +2,6 @@ #include #include #include -#include #include "json-schema-to-grammar.h" #include "speculative.h" #include @@ -2820,7 +2819,7 @@ FINE_NIF(generate, ERL_NIF_DIRTY_JOB_CPU_BOUND); // --- JSON Schema to Grammar --- -// The schema text is untrusted. `nlohmann::json::parse` and +// The schema text is untrusted. `common_json::parse` (nlohmann underneath) and // `json_schema_to_grammar` are both recursive descent, so nesting depth in the // text is C-stack depth: a deeply nested schema is a stack overflow (SIGSEGV), // which try/catch cannot recover. Bound size and depth before parsing. @@ -2834,7 +2833,7 @@ json_schema_to_grammar_nif(ErlNifEnv* env, std::string json_str) { } try { - auto schema = nlohmann::ordered_json::parse(json_str); + auto schema = common_json::parse(json_str); std::string grammar = json_schema_to_grammar(schema); return fine::Ok(grammar); } catch (const std::exception& e) { diff --git a/vendor/llama.cpp b/vendor/llama.cpp index e85caa8..ca3d5a3 160000 --- a/vendor/llama.cpp +++ b/vendor/llama.cpp @@ -1 +1 @@ -Subproject commit e85caa81ea2b65797396018c179b87ad61fa38ab +Subproject commit ca3d5a3e10d53f7ea672cb9b6178faca3e2807bc