Skip to content

Assert if to use --cache-type-k q8_0 --cache-type-v q8_0 parameters for Qwen3.5-9B-Q8_0 #4

Description

@art-den

Name and Version

D:\cpp\llama.cpp\cubetitled-ui\llama.cpp\ggml\src\ggml-cpu\ops.cpp:601: GGML_ASSERT(nb10 == sizeof(float)) failed
D:\cpp\llama.cpp\cubetitled-ui\llama.cpp\ggml\src\ggml-cpu\ops.cpp:601: GGML_ASSERT(nb10 == sizeof(float)) failed

The model loads and works correctly if these parameters are not used.

PS: master branch

Operating systems

Windows

GGML backends

CUDA

Hardware

2x RTX 5060 TI 16Gb, but only one used for small Qwen3.5-9B-Q8_0

Models

Qwen3.5-9B-Q8_0

Problem description & steps to reproduce

set LL_MODEL=Qwen3.5-9B-Q8_0

:::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::

set LLM_TEMP=--temp 1.0

set CTX_SIZE_OPTS= --ctx-size 210000 --cache-type-k q8_0 --cache-type-v q8_0

set VRAM_LOAD_OPTS= --fit-target 100 --n-gpu-layers 999
set REP_PENALITY_OPTS= --repeat-penalty 1.05 --repeat-last-n 8192
set MODEL_OPTS= -m models%LL_MODEL%%LL_MODEL%.gguf --alias "%LL_MODEL%"
set CHAT_TEMPL_OPTS= --jinja
set MISC_OPTS= --port 11112 --log-colors off --no-mmap --parallel 1 -t 12
set MTP_OPTS= --spec-type ngram-mod
set CACHE_PROMPT_OPTS= --cache-prompt --cache-ram 16384 --cache-reuse 8 -sps 0.8

:::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::

title %LL_MODEL% %LLM_TEMP% + cubetitled-ui
set CUDA_VISIBLE_DEVICES=0
set RECURRENT_D=12
set path=D:\cpp\llama.cpp\cubetitled-ui\llama.cpp\build\bin\Release;%path%
llama-server.exe %MODEL_OPTS% %CHAT_TEMPL_OPTS% %VRAM_LOAD_OPTS% %CTX_SIZE_OPTS% %LLM_TEMP% %MISC_OPTS% %MTP_OPTS% %THINK_OPTS% %GPU_OPTS% %REP_PENALITY_OPTS% %CACHE_PROMPT_OPTS%

First Bad Commit

No response

Relevant log output


Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions