Skip to content

Misc. bug: Performance: Vulkan ggml_vk_get_op_batch_size for GGML_OP_GET_ROWS always returns 0 on Apple M2 Fedora Asahi #28035

Description

@rr-it

Name and Version

$ llama-cli --version
version: 0.3.0-dev (build 10701, commit cc231cb)
built with GNU 16.2.1 for Linux aarch64

Operating systems

Linux

Which llama.cpp modules do you know to be affected?

libllama (core library)

Command line

env HK_SYSMEM=$((20 * 1 << 30)) ~/var/llama.cpp/build/bin/llama-cli \
  -m ~/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf \
  -p "Hello" \
  --gpu-layers 99 --ctx-size 32768 -b 512 -ub 256 \
  --prio 3 --threads 2 --threads-batch 2 --cpu-range 4-7 --cpu-strict 1 \
   --direct-io --split-mode none --parallel 1 \
   --cache-ram 4096 --flash-attn on \
   --cache-type-k q4_0 --cache-type-v q4_0 \
   --reasoning-preserve --chat-template-kwargs '{"reasoning_effort":"medium"}' \
   -lv 4

Problem description & steps to reproduce

Vulkan: ggml_vk_get_op_batch_size for GGML_OP_GET_ROWS always returns 0.

Thereby running an AI modell via llama results in 2 graph splits instead of only 1 - and longer prompt eval time.

Current behaviour - 2 graph splits:

Full output (click to open)
USER@fedora:~$ env HK_SYSMEM=$((20 * 1 << 30)) ~/var/llama.cpp/build/bin/llama-cli -m ~/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf -p "Hello" --gpu-layers 99 --ctx-size 32768 -b 512 -ub 256 --prio 3 --threads 2 --threads-batch 2 --cpu-range 4-7 --cpu-strict 1 --direct-io --split-mode none --parallel 1 --cache-ram 4096 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --reasoning-preserve -lv 4 --chat-template-kwargs '{"reasoning_effort":"medium"}'


Loading model... |0.00.033.192 I cmn  common_param: common_params_print_info: build 10701 (cc231cb0d) with GNU 16.2.1 for Linux aarch64
0.00.033.193 I cmn  common_param: common_params_print_info: verbosity = 4 (adjust with the `-lv N` CLI arg)
0.00.033.194 I cmn  common_param: device_info:
0.00.033.290 I cmn  common_param:   - Vulkan0 : Apple M2 (G14G B0) (20480 MiB, 20480 MiB free)
0.00.033.297 I cmn  common_param:   - CPU     : CPU (23544 MiB, 23544 MiB free)
0.00.033.360 I cmn  common_param: system_info: n_threads = 2 (n_threads_batch = 2) / 8 | CPU : NEON = 1 | ARM_FMA = 1 | FP16_VA = 1 | MATMUL_INT8 = 1 | DOTPROD = 1 | OPENMP = 1 | REPACK = 1 | 
0.00.033.394 I srv          init: running without SSL
0.00.033.429 I srv          init: using 7 threads for HTTP server
0.00.033.622 W srv  llama_server: -----------------
0.00.033.624 W srv  llama_server: CORS is set to allow all origins ('*') and no API key is set
0.00.033.624 W srv  llama_server: this can be a security risk (cross-origin attacks)
0.00.033.624 W srv  llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655
0.00.033.624 W srv  llama_server: -----------------
0.00.033.634 I srv         start: binding port with default address family
0.00.034.750 I srv    load_model: loading model '/home/USER/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf'
0.00.034.752 I srv    load_model: local path '/home/USER/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf'
0.00.034.765 I cmn  common_init_: fitting params to device memory ...
0.00.034.765 I cmn  common_init_: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
0.00.034.782 I common_params_fit_impl: getting device memory data for initial parameters:
\0.00.318.592 I common_memory_breakdown_print: | memory breakdown [MiB]           | total    free     self   model   context   compute    unaccounted |
0.00.318.594 I common_memory_breakdown_print: |   - Vulkan0 (Apple M2 (G14G B0)) | 20480 = 20480 + (15482 = 14674 +     725 +      82) +      -15482 |
0.00.318.594 I common_memory_breakdown_print: |   - Host                         |                    708 =   682 +       0 +      26                |
|0.00.345.943 I common_params_fit_impl: projected to use 15482 MiB of device memory vs. 20480 MiB of free device memory
0.00.345.946 I common_params_fit_impl: will leave 4997 >= 1024 MiB of free device memory, no changes needed
0.00.345.946 I common_fit_params: successfully fit params to free device memory
0.00.345.948 I common_fit_params: fitting params to free memory took 0,31 seconds
0.00.375.521 I llama_model_loader: loaded meta data with 50 key-value pairs and 866 tensors from /home/USER/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf (version GGUF V3 (latest))
0.00.375.548 I llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
0.00.375.551 I llama_model_loader: - kv   0:                       general.architecture str              = qwen35
0.00.375.551 I llama_model_loader: - kv   1:                               general.type str              = model
0.00.375.552 I llama_model_loader: - kv   2:                     general.sampling.top_k i32              = 20
0.00.375.554 I llama_model_loader: - kv   3:                     general.sampling.top_p f32              = 0,950000
0.00.375.555 I llama_model_loader: - kv   4:                      general.sampling.temp f32              = 1,000000
0.00.375.555 I llama_model_loader: - kv   5:                               general.name str              = Qwen3.8-27B
0.00.375.555 I llama_model_loader: - kv   6:                           general.basename str              = Qwen3.8-27B
0.00.375.556 I llama_model_loader: - kv   7:                        general.description str              = Renewal of the beloved Qwen model, de...
0.00.375.556 I llama_model_loader: - kv   8:                       general.quantized_by str              = Unsloth
0.00.375.557 I llama_model_loader: - kv   9:                         general.size_label str              = 27B
0.00.375.557 I llama_model_loader: - kv  10:                            general.license str              = apache-2.0
0.00.375.557 I llama_model_loader: - kv  11:                           general.repo_url str              = https://huggingface.co/unsloth
0.00.375.558 I llama_model_loader: - kv  12:                   general.base_model.count u32              = 1
0.00.375.558 I llama_model_loader: - kv  13:                  general.base_model.0.name str              = Qwen3.8 27B
0.00.375.558 I llama_model_loader: - kv  14:          general.base_model.0.organization str              = Qwen
0.00.375.559 I llama_model_loader: - kv  15:              general.base_model.0.repo_url str              = https://huggingface.co/Qwen/Qwen3.8-27B
0.00.375.564 I llama_model_loader: - kv  16:                               general.tags arr[str,1]       = ["unsloth"]
0.00.375.564 I llama_model_loader: - kv  17:                         qwen35.block_count u32              = 65
0.00.375.564 I llama_model_loader: - kv  18:                      qwen35.context_length u32              = 262144
0.00.375.564 I llama_model_loader: - kv  19:                    qwen35.embedding_length u32              = 5120
0.00.375.565 I llama_model_loader: - kv  20:                 qwen35.feed_forward_length u32              = 17408
0.00.375.565 I llama_model_loader: - kv  21:                qwen35.attention.head_count u32              = 24
0.00.375.565 I llama_model_loader: - kv  22:             qwen35.attention.head_count_kv u32              = 4
0.00.375.566 I llama_model_loader: - kv  23:             qwen35.rope.dimension_sections arr[i32,4]       = [11, 11, 10, 0]
0.00.375.567 I llama_model_loader: - kv  24:                      qwen35.rope.freq_base f32              = 10000000,000000
0.00.375.568 I llama_model_loader: - kv  25:    qwen35.attention.layer_norm_rms_epsilon f32              = 0,000001
0.00.375.568 I llama_model_loader: - kv  26:                qwen35.attention.key_length u32              = 256
0.00.375.568 I llama_model_loader: - kv  27:              qwen35.attention.value_length u32              = 256
0.00.375.568 I llama_model_loader: - kv  28:                qwen35.nextn_predict_layers u32              = 1
0.00.375.569 I llama_model_loader: - kv  29:                     qwen35.ssm.conv_kernel u32              = 4
0.00.375.569 I llama_model_loader: - kv  30:                      qwen35.ssm.state_size u32              = 128
0.00.375.569 I llama_model_loader: - kv  31:                     qwen35.ssm.group_count u32              = 16
0.00.375.569 I llama_model_loader: - kv  32:                  qwen35.ssm.time_step_rank u32              = 48
0.00.375.570 I llama_model_loader: - kv  33:                      qwen35.ssm.inner_size u32              = 6144
0.00.375.570 I llama_model_loader: - kv  34:             qwen35.full_attention_interval u32              = 4
0.00.375.570 I llama_model_loader: - kv  35:                qwen35.rope.dimension_count u32              = 64
0.00.375.570 I llama_model_loader: - kv  36:                       tokenizer.ggml.model str              = gpt2
0.00.375.570 I llama_model_loader: - kv  37:                         tokenizer.ggml.pre str              = qwen35
0.00.390.300 I llama_model_loader: - kv  38:                      tokenizer.ggml.tokens arr[str,248320]  = ["!", "\"", "#", "$", "%", "&", "'", ...
0.00.394.509 I llama_model_loader: - kv  39:                  tokenizer.ggml.token_type arr[i32,248320]  = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
0.00.408.941 I llama_model_loader: - kv  40:                      tokenizer.ggml.merges arr[str,247587]  = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",...
0.00.408.943 I llama_model_loader: - kv  41:                tokenizer.ggml.eos_token_id u32              = 248046
0.00.408.943 I llama_model_loader: - kv  42:            tokenizer.ggml.padding_token_id u32              = 248055
0.00.408.944 I llama_model_loader: - kv  43:                tokenizer.ggml.bos_token_id u32              = 248044
0.00.408.944 I llama_model_loader: - kv  44:               general.quantization_version u32              = 2
0.00.408.944 I llama_model_loader: - kv  45:                          general.file_type u32              = 15
0.00.408.945 I llama_model_loader: - kv  46:                      quantize.imatrix.file str              = Qwen3.8-27B-GGUF/imatrix_unsloth.gguf
0.00.408.945 I llama_model_loader: - kv  47:             quantize.imatrix.entries_count u32              = 496
0.00.408.945 I llama_model_loader: - kv  48:              quantize.imatrix.chunks_count u32              = 1251
0.00.408.947 I llama_model_loader: - kv  49:                    tokenizer.chat_template str              = {%- set image_count = namespace(value...
0.00.408.948 I llama_model_loader: - type  f32:  360 tensors
0.00.408.948 I llama_model_loader: - type q8_0:  106 tensors
0.00.408.948 I llama_model_loader: - type q3_K:    7 tensors
0.00.408.948 I llama_model_loader: - type q4_K:  104 tensors
0.00.408.948 I llama_model_loader: - type q5_K:  131 tensors
0.00.408.949 I llama_model_loader: - type q6_K:   30 tensors
0.00.408.949 I llama_model_loader: - type iq4_nl:    7 tensors
0.00.408.949 I llama_model_loader: - type iq3_s:    4 tensors
0.00.408.949 I llama_model_loader: - type iq4_xs:  117 tensors
0.00.408.950 I print_info: file format = GGUF V3 (latest)
0.00.408.950 I print_info: file type   = Q4_K - Medium
0.00.408.951 I print_info: file size   = 15,32 GiB (4,82 BPW) 
0.00.409.017 I llama_prepare_model_devices: using device Vulkan0 (Apple M2 (G14G B0)) (unknown id) - 20480 MiB free
/0.00.486.566 I load: 0 unused tokens
0.00.511.022 I load: printing all EOG tokens:
0.00.511.024 I load:   - 248044 ('<|endoftext|>')
0.00.511.025 I load:   - 248046 ('<|im_end|>')
0.00.511.025 I load:   - 248063 ('<|fim_pad|>')
0.00.511.025 I load:   - 248064 ('<|repo_name|>')
0.00.511.026 I load:   - 248065 ('<|file_sep|>')
0.00.511.246 I load: special tokens cache size = 33
-0.00.565.821 I load: token to piece cache size = 1,7581 MB
0.00.565.830 I print_info: arch                  = qwen35
0.00.565.830 I print_info: vocab_only            = 0
0.00.565.830 I print_info: no_alloc              = 0
0.00.565.830 I print_info: n_ctx_train           = 262144
0.00.565.831 I print_info: n_embd_inp            = 5120
0.00.565.831 I print_info: n_embd                = 5120
0.00.565.832 I print_info: n_embd_out            = 5120
0.00.565.832 I print_info: n_layer               = 64
0.00.565.832 I print_info: n_layer_all           = 65
0.00.565.836 I print_info: n_head                = 24
0.00.565.837 I print_info: n_head_kv             = 4
0.00.565.837 I print_info: n_rot                 = 64
0.00.565.838 I print_info: n_swa                 = 0
0.00.565.838 I print_info: is_swa_any            = 0
0.00.565.838 I print_info: n_embd_head_k         = 256
0.00.565.838 I print_info: n_embd_head_v         = 256
0.00.565.839 I print_info: n_gqa                 = 6
0.00.565.840 I print_info: n_embd_k_gqa          = 1024
0.00.565.841 I print_info: n_embd_v_gqa          = 1024
0.00.565.842 I print_info: f_norm_eps            = 0,0e+00
0.00.565.842 I print_info: f_norm_rms_eps        = 1,0e-06
0.00.565.843 I print_info: f_clamp_kqv           = 0,0e+00
0.00.565.843 I print_info: f_max_alibi_bias      = 0,0e+00
0.00.565.843 I print_info: f_logit_scale         = 0,0e+00
0.00.565.843 I print_info: f_attn_scale          = 0,0e+00
0.00.565.843 I print_info: f_attn_value_scale    = 0,0000
0.00.565.844 I print_info: n_ff                  = 17408
0.00.565.844 I print_info: n_expert              = 0
0.00.565.844 I print_info: n_expert_used         = 0
0.00.565.844 I print_info: n_expert_groups       = 0
0.00.565.845 I print_info: n_group_used          = 0
0.00.565.845 I print_info: causal attn           = 1
0.00.565.845 I print_info: pooling type          = -1
0.00.565.845 I print_info: rope type             = 40
0.00.565.845 I print_info: rope scaling          = linear
0.00.565.846 I print_info: freq_base_train       = 10000000,0
0.00.565.846 I print_info: freq_scale_train      = 1
0.00.565.846 I print_info: n_ctx_orig_yarn       = 262144
0.00.565.846 I print_info: rope_yarn_log_mul     = 0,0000
0.00.565.846 I print_info: rope_finetuned        = unknown
0.00.565.847 I print_info: mrope sections        = [11, 11, 10, 0]
0.00.565.847 I print_info: ssm_d_conv            = 4
0.00.565.847 I print_info: ssm_d_inner           = 6144
0.00.565.847 I print_info: ssm_d_state           = 128
0.00.565.847 I print_info: ssm_dt_rank           = 48
0.00.565.847 I print_info: ssm_n_group           = 16
0.00.565.847 I print_info: ssm_dt_b_c_rms        = 0
0.00.565.848 I print_info: model type            = 27B
0.00.565.848 I print_info: model params          = 27,32 B
0.00.565.848 I print_info: general.name          = Qwen3.8-27B
0.00.565.849 I print_info: vocab type            = BPE
0.00.565.849 I print_info: n_vocab               = 248320
0.00.565.850 I print_info: n_merges              = 247587
0.00.565.850 I print_info: BOS token             = 248044 '<|endoftext|>'
0.00.565.850 I print_info: EOS token             = 248046 '<|im_end|>'
0.00.565.850 I print_info: EOT token             = 248046 '<|im_end|>'
0.00.565.850 I print_info: PAD token             = 248055 '<|vision_pad|>'
0.00.565.850 I print_info: LF token              = 198 'Ċ'
0.00.565.851 I print_info: FIM PRE token         = 248060 '<|fim_prefix|>'
0.00.565.851 I print_info: FIM SUF token         = 248062 '<|fim_suffix|>'
0.00.565.851 I print_info: FIM MID token         = 248061 '<|fim_middle|>'
0.00.565.851 I print_info: FIM PAD token         = 248063 '<|fim_pad|>'
0.00.565.851 I print_info: FIM REP token         = 248064 '<|repo_name|>'
0.00.565.851 I print_info: FIM SEP token         = 248065 '<|file_sep|>'
0.00.565.851 I print_info: EOG token             = 248044 '<|endoftext|>'
0.00.565.852 I print_info: EOG token             = 248046 '<|im_end|>'
0.00.565.852 I print_info: EOG token             = 248063 '<|fim_pad|>'
0.00.565.852 I print_info: EOG token             = 248064 '<|repo_name|>'
0.00.565.852 I print_info: EOG token             = 248065 '<|file_sep|>'
0.00.565.852 I print_info: max token length      = 256
0.00.565.853 I load_tensors: loading model tensors, this can take a while... (load_mode = dio)
0.00.567.664 W model has unused tensor blk.64.attn_norm.weight (size = 20480 bytes) -- ignoring
0.00.567.667 W model has unused tensor blk.64.post_attention_norm.weight (size = 20480 bytes) -- ignoring
0.00.567.674 W model has unused tensor blk.64.attn_q.weight (size = 51609600 bytes) -- ignoring
0.00.567.677 W model has unused tensor blk.64.attn_k.weight (size = 5570560 bytes) -- ignoring
0.00.567.679 W model has unused tensor blk.64.attn_v.weight (size = 5570560 bytes) -- ignoring
0.00.567.687 W model has unused tensor blk.64.attn_output.weight (size = 25804800 bytes) -- ignoring
0.00.567.690 W model has unused tensor blk.64.attn_q_norm.weight (size = 1024 bytes) -- ignoring
0.00.567.692 W model has unused tensor blk.64.attn_k_norm.weight (size = 1024 bytes) -- ignoring
0.00.567.694 W model has unused tensor blk.64.ffn_gate.weight (size = 73113600 bytes) -- ignoring
0.00.567.696 W model has unused tensor blk.64.ffn_down.weight (size = 73113600 bytes) -- ignoring
0.00.567.699 W model has unused tensor blk.64.ffn_up.weight (size = 73113600 bytes) -- ignoring
0.00.567.702 W model has unused tensor blk.64.nextn.eh_proj.weight (size = 43008000 bytes) -- ignoring
0.00.567.704 W model has unused tensor blk.64.nextn.enorm.weight (size = 20480 bytes) -- ignoring
0.00.567.707 W model has unused tensor blk.64.nextn.hnorm.weight (size = 20480 bytes) -- ignoring
0.00.567.715 W model has unused tensor blk.64.nextn.shared_head_norm.weight (size = 20480 bytes) -- ignoring
|0.01.574.498 I load_tensors: offloading output layer to GPU
0.01.574.501 I load_tensors: offloading 64 repeating layers to GPU
0.01.574.501 I load_tensors: offloaded 66/66 layers to GPU
0.01.574.503 I load_tensors:      Vulkan0 model buffer size = 14674,45 MiB
0.01.574.504 I load_tensors:  Vulkan_Host model buffer size =   682,03 MiB
|0.07.555.179 I cmn  common_init_: added <|endoftext|> logit bias = -inf
0.07.555.182 I cmn  common_init_: added <|im_end|> logit bias = -inf
0.07.555.182 I cmn  common_init_: added <|fim_pad|> logit bias = -inf
0.07.555.183 I cmn  common_init_: added <|repo_name|> logit bias = -inf
0.07.555.183 I cmn  common_init_: added <|file_sep|> logit bias = -inf
0.07.555.236 I llama_context: constructing llama_context
0.07.555.239 I llama_context: n_seq_max             = 1
0.07.555.239 I llama_context: n_ctx                 = 32768
0.07.555.239 I llama_context: n_ctx_seq             = 32768
0.07.555.239 I llama_context: n_batch               = 512
0.07.555.239 I llama_context: n_ubatch              = 256
0.07.555.240 I llama_context: causal_attn           = 1
0.07.555.240 I llama_context: flash_attn            = enabled
0.07.555.240 I llama_context: kv_unified            = false
0.07.555.241 I llama_context: freq_base             = 10000000,0
0.07.555.242 I llama_context: freq_scale            = 1
0.07.555.242 I llama_context: n_rs_seq              = 0
0.07.555.242 I llama_context: n_outputs_max         = 1
0.07.555.242 I llama_context: n_outputs_max_per_seq = 1
0.07.555.242 I llama_context: n_ctx_seq (32768) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
0.07.555.439 I llama_context: Vulkan_Host  output buffer size =     0,95 MiB
0.07.594.849 I llama_kv_cache:    Vulkan0 KV buffer size =   576,00 MiB
0.07.613.586 I llama_kv_cache: size =  576,00 MiB ( 32768 cells,  16 layers,  1/1 seqs), K (q4_0):  288,00 MiB, V (q4_0):  288,00 MiB
0.07.613.589 I llama_kv_cache: attn_rot_k = 1, n_embd_head_k_all = 256
0.07.613.589 I llama_kv_cache: attn_rot_v = 1, n_embd_head_k_all = 256
0.07.636.665 I llama_memory_recurrent:    Vulkan0 RS buffer size =   149,62 MiB
0.07.636.669 I llama_memory_recurrent: size =  149,62 MiB (     1 cells,  64 layers,  1 seqs  0 rs_seq), R (f32):    5,62 MiB, S (f32):  144,00 MiB, P (f32):    0,00 MiB
0.07.636.673 I sched_reserve: reserving ...
/0.07.638.260 I resolve_fused_ops: resolving fused DeepSeek V4 HC support:
0.07.640.871 I resolve_fused_ops: fused DeepSeek V4 HC pre enabled
0.07.643.356 I resolve_fused_ops: fused DeepSeek V4 HC comb enabled
0.07.645.803 I resolve_fused_ops: fused DeepSeek V4 HC post enabled
0.07.667.703 I sched_reserve:    Vulkan0 compute buffer size =    82,27 MiB
0.07.667.704 I sched_reserve: Vulkan_Host compute buffer size =    26,36 MiB
0.07.667.705 I sched_reserve: graph nodes  = 3847
0.07.667.705 I sched_reserve: graph splits = 2
0.07.667.705 I sched_reserve: reserve took 31,03 ms, sched copies = 1
0.07.667.767 I cmn          init: llama threadpool init, n_threads = 2
0.07.667.788 I cmn  common_init_: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|0.08.345.425 I cmn  common_conte: the context does not support partial sequence removal
-0.08.975.157 I srv    load_model: speculative decoding will use checkpoints
0.08.975.159 I srv    load_model: initializing, n_slots = 1, n_ctx_slot = 32768, kv_unified = 'false'
0.08.975.165 I spec common_specu: no implementations specified for speculative decoding
0.08.975.177 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 32768
0.08.975.192 I srv    load_model: prompt cache is enabled, size limit: 4096 MiB
0.08.975.193 I srv    load_model: use `--cache-ram 0` to disable the prompt cache
0.08.975.193 I srv    load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
0.08.975.193 I srv    load_model: context checkpoints enabled, max = 32, min spacing = 8192
0.08.975.203 I srv          init: idle slots will be saved to prompt cache upon starting a new task
0.08.979.206 I srv          init: init: chat template, example_format: '<|im_start|>system
You are a helpful assistant<|im_end|>
<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
<think>

</think>

Hi there<|im_end|>
<|im_start|>user
How are you?<|im_end|>
<|im_start|>assistant
<think>
'
0.08.979.785 I srv          init: init: chat template, thinking = 1
0.08.979.806 I srv  llama_server: model loaded
0.08.979.807 I srv  llama_server: listening on http://127.0.0.1:37797
0.08.979.846 I srv  update_slots: all slots are idle
 

▄▄ ▄▄
██ ██
██ ██  ▀▀█▄ ███▄███▄  ▀▀█▄    ▄████ ████▄ ████▄
██ ██ ▄█▀██ ██ ██ ██ ▄█▀██    ██    ██ ██ ██ ██
██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                    ██    ██
                                    ▀▀    ▀▀

build      : b10701-cc231cb0d
model      : /home/USER/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf
ftype      : Q4_K - Medium
modalities : text

available commands:
  /exit or Ctrl+C     stop or exit
  /regen              regenerate the last response
  /clear              clear the chat history
  /read <file>        add a text file
  /glob <pattern>     add text files using globbing pattern



> Hello
0.09.046.607 I srv  server_strea: conv_id= (empty=1)
0.09.046.767 I srv    operator(): chat format: peg-native
0.09.046.888 I slot get_availabl: id  0 | task -1 |  - skipping, slot is empty
0.09.046.889 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
0.09.046.889 I srv  get_availabl: updating prompt cache
0.09.046.900 I srv          load:  - looking for better prompt, base f_keep = -1,000, f_sim = 0,000
0.09.046.903 I srv        update:  - cache state: 0 prompts, 0,000 MiB (limits: 4096,000 MiB, 32768 tokens, 4294967296 est)
0.09.046.904 I srv  get_availabl: prompt cache update took 0,01 ms
0.09.046.949 I slot launch_slot_: id  0 | task -1 | sampler chain: logits -> ?penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> top-p -> min-p -> ?xtc -> ?temp-ext -> dist 
0.09.046.954 I slot launch_slot_: id  0 | task -1 | sampler params: 
	repeat_last_n = 64, repeat_penalty = 1,000, frequency_penalty = 0,000, presence_penalty = 0,000
	dry_multiplier = 0,000, dry_base = 1,750, dry_allowed_length = 2, dry_penalty_last_n = 64
	top_k = 20, top_p = 0,950, min_p = 0,050, xtc_probability = 0,000, xtc_threshold = 0,100, typical_p = 1,000, top_n_sigma = -1,000, temp = 1,000
	mirostat = 0, mirostat_lr = 0,100, mirostat_ent = 5,000, adaptive_target = -1,000, adaptive_decay = 0,900
0.09.046.955 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
0.09.046.961 I slot   operator(): id  0 | task 0 | new prompt, n_ctx_slot = 32768, n_keep = 0, task.n_tokens = 11
0.09.046.966 I slot   operator(): id  0 | task 0 | cached n_tokens = 0, memory_seq_rm [0, end)
0.09.057.056 I slot   operator(): id  0 | task 0 | cached n_tokens = 7, memory_seq_rm [7, end)
0.09.057.090 I slot init_sampler: id  0 | task 0 | init sampler, took 0,00 ms, tokens: text = 11, total = 11
0.11.121.305 I slot create_check: id  0 | task 0 | created context checkpoint 1 of 32 (pos_min = 6, pos_max = 6, n_tokens = 7, size = 149,626 MiB)

[Start thinking]

The user has sent a simple greeting "Hello". This is a basic social interaction. I should respond warmly and naturally, keeping it concise and friendly, while opening the door for further conversation.
[End thinking]

Hello! How are you doing? Is there something I can help you with today?0.35.627.104 I slot print_timing: id  0 | task 0 | prompt eval time =    3084,32 ms /    11 tokens (  280,39 ms per token,     3,57 tokens per second)
0.35.627.106 I slot print_timing: id  0 | task 0 |        eval time =   23495,82 ms /    59 tokens (  405,10 ms per token,     2,47 tokens per second)
0.35.627.106 I slot print_timing: id  0 | task 0 |       total time =   26580,14 ms /    70 tokens
0.35.627.110 I slot print_timing: id  0 | task 0 |    graphs reused =         58
0.35.627.125 I slot      release: id  0 | task 0 | stop processing: n_tokens = 69, truncated = 0
0.35.627.128 I srv  update_slots: all slots are idle
0.07.667.703 I sched_reserve:    Vulkan0 compute buffer size =    82,27 MiB
0.07.667.704 I sched_reserve: Vulkan_Host compute buffer size =    26,36 MiB
0.07.667.705 I sched_reserve: graph nodes  = 3847
0.07.667.705 I sched_reserve: graph splits = 2

static int64_t ggml_vk_get_op_batch_size(const ggml_tensor * op) {
switch (op->op) {
case GGML_OP_GET_ROWS:
return 0;
case GGML_OP_MUL_MAT:
return op->ne[1];

With change - graph splits = 1 (with bs=256), 2 (with bs=1):

Full output (click to open)
USER@fedora:~$ env HK_SYSMEM=$((20 * 1 << 30)) ~/var/llama.cpp/build/bin/llama-cli -m ~/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf -p "Hello" --gpu-layers 99 --ctx-size 32768 -b 512 -ub 256 --prio 3 --threads 2 --threads-batch 2 --cpu-range 4-7 --cpu-strict 1 --direct-io --split-mode none --parallel 1 --cache-ram 4096 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --reasoning-preserve -lv 4 --chat-template-kwargs '{"reasoning_effort":"medium"}'


Loading model... |0.00.018.337 I cmn  common_param: common_params_print_info: build 10701 (cc231cb0d) with GNU 16.2.1 for Linux aarch64
0.00.018.339 I cmn  common_param: common_params_print_info: verbosity = 4 (adjust with the `-lv N` CLI arg)
0.00.018.340 I cmn  common_param: device_info:
0.00.018.471 I cmn  common_param:   - Vulkan0 : Apple M2 (G14G B0) (20480 MiB, 20480 MiB free)
0.00.018.478 I cmn  common_param:   - CPU     : CPU (23544 MiB, 23544 MiB free)
0.00.018.538 I cmn  common_param: system_info: n_threads = 2 (n_threads_batch = 2) / 8 | CPU : NEON = 1 | ARM_FMA = 1 | FP16_VA = 1 | MATMUL_INT8 = 1 | DOTPROD = 1 | OPENMP = 1 | REPACK = 1 | 
0.00.018.568 I srv          init: running without SSL
0.00.018.652 I srv          init: using 7 threads for HTTP server
0.00.018.986 W srv  llama_server: -----------------
0.00.018.987 W srv  llama_server: CORS is set to allow all origins ('*') and no API key is set
0.00.018.988 W srv  llama_server: this can be a security risk (cross-origin attacks)
0.00.018.988 W srv  llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655
0.00.018.994 W srv  llama_server: -----------------
0.00.019.002 I srv         start: binding port with default address family
0.00.020.157 I srv    load_model: loading model '/home/USER/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf'
0.00.020.159 I srv    load_model: local path '/home/USER/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf'
0.00.020.169 I cmn  common_init_: fitting params to device memory ...
0.00.020.170 I cmn  common_init_: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
0.00.020.188 I common_params_fit_impl: getting device memory data for initial parameters:
\0.00.302.345 I common_memory_breakdown_print: | memory breakdown [MiB]           | total    free     self   model   context   compute    unaccounted |
0.00.302.349 I common_memory_breakdown_print: |   - Vulkan0 (Apple M2 (G14G B0)) | 20480 = 20480 + (16103 = 14674 +     725 +     703) +      -16103 |
0.00.302.349 I common_memory_breakdown_print: |   - Host                         |                    703 =   682 +       0 +      21                |
|0.00.329.826 I common_params_fit_impl: projected to use 16103 MiB of device memory vs. 20480 MiB of free device memory
0.00.329.830 I common_params_fit_impl: will leave 4376 >= 1024 MiB of free device memory, no changes needed
0.00.329.830 I common_fit_params: successfully fit params to free device memory
0.00.329.832 I common_fit_params: fitting params to free memory took 0,31 seconds
0.00.359.454 I llama_model_loader: loaded meta data with 50 key-value pairs and 866 tensors from /home/USER/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf (version GGUF V3 (latest))
0.00.359.481 I llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
0.00.359.484 I llama_model_loader: - kv   0:                       general.architecture str              = qwen35
0.00.359.485 I llama_model_loader: - kv   1:                               general.type str              = model
0.00.359.486 I llama_model_loader: - kv   2:                     general.sampling.top_k i32              = 20
0.00.359.488 I llama_model_loader: - kv   3:                     general.sampling.top_p f32              = 0,950000
0.00.359.488 I llama_model_loader: - kv   4:                      general.sampling.temp f32              = 1,000000
0.00.359.489 I llama_model_loader: - kv   5:                               general.name str              = Qwen3.8-27B
0.00.359.489 I llama_model_loader: - kv   6:                           general.basename str              = Qwen3.8-27B
0.00.359.490 I llama_model_loader: - kv   7:                        general.description str              = Renewal of the beloved Qwen model, de...
0.00.359.490 I llama_model_loader: - kv   8:                       general.quantized_by str              = Unsloth
0.00.359.491 I llama_model_loader: - kv   9:                         general.size_label str              = 27B
0.00.359.491 I llama_model_loader: - kv  10:                            general.license str              = apache-2.0
0.00.359.491 I llama_model_loader: - kv  11:                           general.repo_url str              = https://huggingface.co/unsloth
0.00.359.492 I llama_model_loader: - kv  12:                   general.base_model.count u32              = 1
0.00.359.492 I llama_model_loader: - kv  13:                  general.base_model.0.name str              = Qwen3.8 27B
0.00.359.492 I llama_model_loader: - kv  14:          general.base_model.0.organization str              = Qwen
0.00.359.493 I llama_model_loader: - kv  15:              general.base_model.0.repo_url str              = https://huggingface.co/Qwen/Qwen3.8-27B
0.00.359.498 I llama_model_loader: - kv  16:                               general.tags arr[str,1]       = ["unsloth"]
0.00.359.498 I llama_model_loader: - kv  17:                         qwen35.block_count u32              = 65
0.00.359.498 I llama_model_loader: - kv  18:                      qwen35.context_length u32              = 262144
0.00.359.499 I llama_model_loader: - kv  19:                    qwen35.embedding_length u32              = 5120
0.00.359.499 I llama_model_loader: - kv  20:                 qwen35.feed_forward_length u32              = 17408
0.00.359.499 I llama_model_loader: - kv  21:                qwen35.attention.head_count u32              = 24
0.00.359.499 I llama_model_loader: - kv  22:             qwen35.attention.head_count_kv u32              = 4
0.00.359.500 I llama_model_loader: - kv  23:             qwen35.rope.dimension_sections arr[i32,4]       = [11, 11, 10, 0]
0.00.359.501 I llama_model_loader: - kv  24:                      qwen35.rope.freq_base f32              = 10000000,000000
0.00.359.501 I llama_model_loader: - kv  25:    qwen35.attention.layer_norm_rms_epsilon f32              = 0,000001
0.00.359.502 I llama_model_loader: - kv  26:                qwen35.attention.key_length u32              = 256
0.00.359.502 I llama_model_loader: - kv  27:              qwen35.attention.value_length u32              = 256
0.00.359.502 I llama_model_loader: - kv  28:                qwen35.nextn_predict_layers u32              = 1
0.00.359.503 I llama_model_loader: - kv  29:                     qwen35.ssm.conv_kernel u32              = 4
0.00.359.503 I llama_model_loader: - kv  30:                      qwen35.ssm.state_size u32              = 128
0.00.359.503 I llama_model_loader: - kv  31:                     qwen35.ssm.group_count u32              = 16
0.00.359.503 I llama_model_loader: - kv  32:                  qwen35.ssm.time_step_rank u32              = 48
0.00.359.503 I llama_model_loader: - kv  33:                      qwen35.ssm.inner_size u32              = 6144
0.00.359.504 I llama_model_loader: - kv  34:             qwen35.full_attention_interval u32              = 4
0.00.359.504 I llama_model_loader: - kv  35:                qwen35.rope.dimension_count u32              = 64
0.00.359.504 I llama_model_loader: - kv  36:                       tokenizer.ggml.model str              = gpt2
0.00.359.504 I llama_model_loader: - kv  37:                         tokenizer.ggml.pre str              = qwen35
0.00.374.463 I llama_model_loader: - kv  38:                      tokenizer.ggml.tokens arr[str,248320]  = ["!", "\"", "#", "$", "%", "&", "'", ...
0.00.378.690 I llama_model_loader: - kv  39:                  tokenizer.ggml.token_type arr[i32,248320]  = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
0.00.393.285 I llama_model_loader: - kv  40:                      tokenizer.ggml.merges arr[str,247587]  = ["Ġ Ġ", "ĠĠ ĠĠ", "i n", "Ġ t",...
0.00.393.288 I llama_model_loader: - kv  41:                tokenizer.ggml.eos_token_id u32              = 248046
0.00.393.288 I llama_model_loader: - kv  42:            tokenizer.ggml.padding_token_id u32              = 248055
0.00.393.288 I llama_model_loader: - kv  43:                tokenizer.ggml.bos_token_id u32              = 248044
0.00.393.289 I llama_model_loader: - kv  44:               general.quantization_version u32              = 2
0.00.393.289 I llama_model_loader: - kv  45:                          general.file_type u32              = 15
0.00.393.290 I llama_model_loader: - kv  46:                      quantize.imatrix.file str              = Qwen3.8-27B-GGUF/imatrix_unsloth.gguf
0.00.393.290 I llama_model_loader: - kv  47:             quantize.imatrix.entries_count u32              = 496
0.00.393.290 I llama_model_loader: - kv  48:              quantize.imatrix.chunks_count u32              = 1251
0.00.393.292 I llama_model_loader: - kv  49:                    tokenizer.chat_template str              = {%- set image_count = namespace(value...
0.00.393.292 I llama_model_loader: - type  f32:  360 tensors
0.00.393.293 I llama_model_loader: - type q8_0:  106 tensors
0.00.393.293 I llama_model_loader: - type q3_K:    7 tensors
0.00.393.293 I llama_model_loader: - type q4_K:  104 tensors
0.00.393.293 I llama_model_loader: - type q5_K:  131 tensors
0.00.393.293 I llama_model_loader: - type q6_K:   30 tensors
0.00.393.294 I llama_model_loader: - type iq4_nl:    7 tensors
0.00.393.294 I llama_model_loader: - type iq3_s:    4 tensors
0.00.393.294 I llama_model_loader: - type iq4_xs:  117 tensors
0.00.393.295 I print_info: file format = GGUF V3 (latest)
0.00.393.295 I print_info: file type   = Q4_K - Medium
0.00.393.296 I print_info: file size   = 15,32 GiB (4,82 BPW) 
0.00.393.372 I llama_prepare_model_devices: using device Vulkan0 (Apple M2 (G14G B0)) (unknown id) - 20480 MiB free
/0.00.469.975 I load: 0 unused tokens
0.00.494.127 I load: printing all EOG tokens:
0.00.494.130 I load:   - 248044 ('<|endoftext|>')
0.00.494.131 I load:   - 248046 ('<|im_end|>')
0.00.494.131 I load:   - 248063 ('<|fim_pad|>')
0.00.494.131 I load:   - 248064 ('<|repo_name|>')
0.00.494.131 I load:   - 248065 ('<|file_sep|>')
0.00.494.374 I load: special tokens cache size = 33
-0.00.549.100 I load: token to piece cache size = 1,7581 MB
0.00.549.111 I print_info: arch                  = qwen35
0.00.549.111 I print_info: vocab_only            = 0
0.00.549.111 I print_info: no_alloc              = 0
0.00.549.112 I print_info: n_ctx_train           = 262144
0.00.549.112 I print_info: n_embd_inp            = 5120
0.00.549.112 I print_info: n_embd                = 5120
0.00.549.112 I print_info: n_embd_out            = 5120
0.00.549.112 I print_info: n_layer               = 64
0.00.549.113 I print_info: n_layer_all           = 65
0.00.549.117 I print_info: n_head                = 24
0.00.549.118 I print_info: n_head_kv             = 4
0.00.549.118 I print_info: n_rot                 = 64
0.00.549.118 I print_info: n_swa                 = 0
0.00.549.119 I print_info: is_swa_any            = 0
0.00.549.119 I print_info: n_embd_head_k         = 256
0.00.549.119 I print_info: n_embd_head_v         = 256
0.00.549.120 I print_info: n_gqa                 = 6
0.00.549.121 I print_info: n_embd_k_gqa          = 1024
0.00.549.121 I print_info: n_embd_v_gqa          = 1024
0.00.549.122 I print_info: f_norm_eps            = 0,0e+00
0.00.549.122 I print_info: f_norm_rms_eps        = 1,0e-06
0.00.549.123 I print_info: f_clamp_kqv           = 0,0e+00
0.00.549.123 I print_info: f_max_alibi_bias      = 0,0e+00
0.00.549.123 I print_info: f_logit_scale         = 0,0e+00
0.00.549.123 I print_info: f_attn_scale          = 0,0e+00
0.00.549.123 I print_info: f_attn_value_scale    = 0,0000
0.00.549.124 I print_info: n_ff                  = 17408
0.00.549.124 I print_info: n_expert              = 0
0.00.549.124 I print_info: n_expert_used         = 0
0.00.549.124 I print_info: n_expert_groups       = 0
0.00.549.124 I print_info: n_group_used          = 0
0.00.549.124 I print_info: causal attn           = 1
0.00.549.125 I print_info: pooling type          = -1
0.00.549.125 I print_info: rope type             = 40
0.00.549.125 I print_info: rope scaling          = linear
0.00.549.125 I print_info: freq_base_train       = 10000000,0
0.00.549.126 I print_info: freq_scale_train      = 1
0.00.549.126 I print_info: n_ctx_orig_yarn       = 262144
0.00.549.126 I print_info: rope_yarn_log_mul     = 0,0000
0.00.549.126 I print_info: rope_finetuned        = unknown
0.00.549.126 I print_info: mrope sections        = [11, 11, 10, 0]
0.00.549.126 I print_info: ssm_d_conv            = 4
0.00.549.127 I print_info: ssm_d_inner           = 6144
0.00.549.127 I print_info: ssm_d_state           = 128
0.00.549.127 I print_info: ssm_dt_rank           = 48
0.00.549.127 I print_info: ssm_n_group           = 16
0.00.549.127 I print_info: ssm_dt_b_c_rms        = 0
0.00.549.127 I print_info: model type            = 27B
0.00.549.128 I print_info: model params          = 27,32 B
0.00.549.128 I print_info: general.name          = Qwen3.8-27B
0.00.549.129 I print_info: vocab type            = BPE
0.00.549.129 I print_info: n_vocab               = 248320
0.00.549.129 I print_info: n_merges              = 247587
0.00.549.129 I print_info: BOS token             = 248044 '<|endoftext|>'
0.00.549.130 I print_info: EOS token             = 248046 '<|im_end|>'
0.00.549.130 I print_info: EOT token             = 248046 '<|im_end|>'
0.00.549.130 I print_info: PAD token             = 248055 '<|vision_pad|>'
0.00.549.130 I print_info: LF token              = 198 'Ċ'
0.00.549.130 I print_info: FIM PRE token         = 248060 '<|fim_prefix|>'
0.00.549.130 I print_info: FIM SUF token         = 248062 '<|fim_suffix|>'
0.00.549.130 I print_info: FIM MID token         = 248061 '<|fim_middle|>'
0.00.549.131 I print_info: FIM PAD token         = 248063 '<|fim_pad|>'
0.00.549.131 I print_info: FIM REP token         = 248064 '<|repo_name|>'
0.00.549.131 I print_info: FIM SEP token         = 248065 '<|file_sep|>'
0.00.549.131 I print_info: EOG token             = 248044 '<|endoftext|>'
0.00.549.131 I print_info: EOG token             = 248046 '<|im_end|>'
0.00.549.132 I print_info: EOG token             = 248063 '<|fim_pad|>'
0.00.549.132 I print_info: EOG token             = 248064 '<|repo_name|>'
0.00.549.132 I print_info: EOG token             = 248065 '<|file_sep|>'
0.00.549.132 I print_info: max token length      = 256
0.00.549.133 I load_tensors: loading model tensors, this can take a while... (load_mode = dio)
0.00.550.975 W model has unused tensor blk.64.attn_norm.weight (size = 20480 bytes) -- ignoring
0.00.550.979 W model has unused tensor blk.64.post_attention_norm.weight (size = 20480 bytes) -- ignoring
0.00.550.985 W model has unused tensor blk.64.attn_q.weight (size = 51609600 bytes) -- ignoring
0.00.550.987 W model has unused tensor blk.64.attn_k.weight (size = 5570560 bytes) -- ignoring
0.00.550.991 W model has unused tensor blk.64.attn_v.weight (size = 5570560 bytes) -- ignoring
0.00.550.999 W model has unused tensor blk.64.attn_output.weight (size = 25804800 bytes) -- ignoring
0.00.551.002 W model has unused tensor blk.64.attn_q_norm.weight (size = 1024 bytes) -- ignoring
0.00.551.004 W model has unused tensor blk.64.attn_k_norm.weight (size = 1024 bytes) -- ignoring
0.00.551.006 W model has unused tensor blk.64.ffn_gate.weight (size = 73113600 bytes) -- ignoring
0.00.551.009 W model has unused tensor blk.64.ffn_down.weight (size = 73113600 bytes) -- ignoring
0.00.551.011 W model has unused tensor blk.64.ffn_up.weight (size = 73113600 bytes) -- ignoring
0.00.551.014 W model has unused tensor blk.64.nextn.eh_proj.weight (size = 43008000 bytes) -- ignoring
0.00.551.026 W model has unused tensor blk.64.nextn.enorm.weight (size = 20480 bytes) -- ignoring
0.00.551.028 W model has unused tensor blk.64.nextn.hnorm.weight (size = 20480 bytes) -- ignoring
0.00.551.035 W model has unused tensor blk.64.nextn.shared_head_norm.weight (size = 20480 bytes) -- ignoring
|0.01.557.001 I load_tensors: offloading output layer to GPU
0.01.557.004 I load_tensors: offloading 64 repeating layers to GPU
0.01.557.004 I load_tensors: offloaded 66/66 layers to GPU
0.01.557.006 I load_tensors:      Vulkan0 model buffer size = 14674,45 MiB
0.01.557.007 I load_tensors:  Vulkan_Host model buffer size =   682,03 MiB
|0.07.532.720 I cmn  common_init_: added <|endoftext|> logit bias = -inf
0.07.532.724 I cmn  common_init_: added <|im_end|> logit bias = -inf
0.07.532.725 I cmn  common_init_: added <|fim_pad|> logit bias = -inf
0.07.532.725 I cmn  common_init_: added <|repo_name|> logit bias = -inf
0.07.532.725 I cmn  common_init_: added <|file_sep|> logit bias = -inf
0.07.532.769 I llama_context: constructing llama_context
0.07.532.771 I llama_context: n_seq_max             = 1
0.07.532.771 I llama_context: n_ctx                 = 32768
0.07.532.771 I llama_context: n_ctx_seq             = 32768
0.07.532.771 I llama_context: n_batch               = 512
0.07.532.771 I llama_context: n_ubatch              = 256
0.07.532.771 I llama_context: causal_attn           = 1
0.07.532.772 I llama_context: flash_attn            = enabled
0.07.532.772 I llama_context: kv_unified            = false
0.07.532.773 I llama_context: freq_base             = 10000000,0
0.07.532.774 I llama_context: freq_scale            = 1
0.07.532.774 I llama_context: n_rs_seq              = 0
0.07.532.774 I llama_context: n_outputs_max         = 1
0.07.532.774 I llama_context: n_outputs_max_per_seq = 1
0.07.532.774 I llama_context: n_ctx_seq (32768) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
0.07.532.995 I llama_context: Vulkan_Host  output buffer size =     0,95 MiB
0.07.571.864 I llama_kv_cache:    Vulkan0 KV buffer size =   576,00 MiB
0.07.590.256 I llama_kv_cache: size =  576,00 MiB ( 32768 cells,  16 layers,  1/1 seqs), K (q4_0):  288,00 MiB, V (q4_0):  288,00 MiB
0.07.590.259 I llama_kv_cache: attn_rot_k = 1, n_embd_head_k_all = 256
0.07.590.259 I llama_kv_cache: attn_rot_v = 1, n_embd_head_k_all = 256
0.07.618.249 I llama_memory_recurrent:    Vulkan0 RS buffer size =   149,62 MiB
0.07.618.254 I llama_memory_recurrent: size =  149,62 MiB (     1 cells,  64 layers,  1 seqs  0 rs_seq), R (f32):    5,62 MiB, S (f32):  144,00 MiB, P (f32):    0,00 MiB
0.07.618.257 I sched_reserve: reserving ...
0.07.619.791 I resolve_fused_ops: resolving fused DeepSeek V4 HC support:
0.07.622.389 I resolve_fused_ops: fused DeepSeek V4 HC pre enabled
/0.07.624.906 I resolve_fused_ops: fused DeepSeek V4 HC comb enabled
0.07.627.365 I resolve_fused_ops: fused DeepSeek V4 HC post enabled
0.07.699.832 I sched_reserve:    Vulkan0 compute buffer size =   703,31 MiB
0.07.699.836 I sched_reserve: Vulkan_Host compute buffer size =    21,36 MiB
0.07.699.836 I sched_reserve: graph nodes  = 3847
0.07.699.837 I sched_reserve: graph splits = 1 (with bs=256), 2 (with bs=1)
0.07.699.837 I sched_reserve: reserve took 81,58 ms, sched copies = 1
0.07.699.899 I cmn          init: llama threadpool init, n_threads = 2
0.07.699.931 I cmn  common_init_: warming up the model with an empty run - please wait ... (--no-warmup to disable)
|0.08.398.728 I cmn  common_conte: the context does not support partial sequence removal
\0.09.035.558 I srv    load_model: speculative decoding will use checkpoints
0.09.035.561 I srv    load_model: initializing, n_slots = 1, n_ctx_slot = 32768, kv_unified = 'false'
0.09.035.567 I spec common_specu: no implementations specified for speculative decoding
0.09.035.573 I slot   load_model: id  0 | task -1 | new slot, n_ctx = 32768
0.09.035.590 I srv    load_model: prompt cache is enabled, size limit: 4096 MiB
0.09.035.593 I srv    load_model: use `--cache-ram 0` to disable the prompt cache
0.09.035.593 I srv    load_model: for more info see https://github.com/ggml-org/llama.cpp/pull/16391
0.09.035.593 I srv    load_model: context checkpoints enabled, max = 32, min spacing = 8192
0.09.035.602 I srv          init: idle slots will be saved to prompt cache upon starting a new task
0.09.039.573 I srv          init: init: chat template, example_format: '<|im_start|>system
You are a helpful assistant<|im_end|>
<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
<think>

</think>

Hi there<|im_end|>
<|im_start|>user
How are you?<|im_end|>
<|im_start|>assistant
<think>
'
0.09.040.143 I srv          init: init: chat template, thinking = 1
0.09.040.165 I srv  llama_server: model loaded
0.09.040.167 I srv  llama_server: listening on http://127.0.0.1:50529
0.09.040.209 I srv  update_slots: all slots are idle
 

▄▄ ▄▄
██ ██
██ ██  ▀▀█▄ ███▄███▄  ▀▀█▄    ▄████ ████▄ ████▄
██ ██ ▄█▀██ ██ ██ ██ ▄█▀██    ██    ██ ██ ██ ██
██ ██ ▀█▄██ ██ ██ ██ ▀█▄██ ██ ▀████ ████▀ ████▀
                                    ██    ██
                                    ▀▀    ▀▀

build      : b10701-cc231cb0d
model      : /home/USER/.local/share/ramalama/store/huggingface/unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_M.gguf/snapshots/sha256-322e194ff79741c7baa497c240f677f54b201b0efab44ca8e50f122b39123482/Qwen3.8-27B-UD-Q4_K_M.gguf
ftype      : Q4_K - Medium
modalities : text

available commands:
  /exit or Ctrl+C     stop or exit
  /regen              regenerate the last response
  /clear              clear the chat history
  /read <file>        add a text file
  /glob <pattern>     add text files using globbing pattern



> Hello
0.09.231.598 I srv  server_strea: conv_id= (empty=1)
0.09.231.758 I srv    operator(): chat format: peg-native
0.09.231.914 I slot get_availabl: id  0 | task -1 |  - skipping, slot is empty
0.09.231.915 I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = -1
0.09.231.915 I srv  get_availabl: updating prompt cache
0.09.231.927 I srv          load:  - looking for better prompt, base f_keep = -1,000, f_sim = 0,000
0.09.231.930 I srv        update:  - cache state: 0 prompts, 0,000 MiB (limits: 4096,000 MiB, 32768 tokens, 4294967296 est)
0.09.231.932 I srv  get_availabl: prompt cache update took 0,02 ms
0.09.231.967 I slot launch_slot_: id  0 | task -1 | sampler chain: logits -> ?penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> top-p -> min-p -> ?xtc -> ?temp-ext -> dist 
0.09.231.971 I slot launch_slot_: id  0 | task -1 | sampler params: 
	repeat_last_n = 64, repeat_penalty = 1,000, frequency_penalty = 0,000, presence_penalty = 0,000
	dry_multiplier = 0,000, dry_base = 1,750, dry_allowed_length = 2, dry_penalty_last_n = 64
	top_k = 20, top_p = 0,950, min_p = 0,050, xtc_probability = 0,000, xtc_threshold = 0,100, typical_p = 1,000, top_n_sigma = -1,000, temp = 1,000
	mirostat = 0, mirostat_lr = 0,100, mirostat_ent = 5,000, adaptive_target = -1,000, adaptive_decay = 0,900
0.09.231.972 I slot launch_slot_: id  0 | task 0 | processing task, is_child = 0
0.09.231.979 I slot   operator(): id  0 | task 0 | new prompt, n_ctx_slot = 32768, n_keep = 0, task.n_tokens = 11
0.09.231.985 I slot   operator(): id  0 | task 0 | cached n_tokens = 0, memory_seq_rm [0, end)
0.09.243.575 I slot   operator(): id  0 | task 0 | cached n_tokens = 7, memory_seq_rm [7, end)
0.09.243.608 I slot init_sampler: id  0 | task 0 | init sampler, took 0,01 ms, tokens: text = 11, total = 11
0.11.287.276 I slot create_check: id  0 | task 0 | created context checkpoint 1 of 32 (pos_min = 6, pos_max = 6, n_tokens = 7, size = 149,626 MiB)

[Start thinking]

The user is simply saying "Hello" - a basic greeting. I should respond warmly and naturally, keeping it simple and friendly. No need for anything elaborate here.
[End thinking]

Hello! How are you doing? Is there something I can help you with today?0.33.690.472 I slot print_timing: id  0 | task 0 | prompt eval time =    3066,70 ms /    11 tokens (  278,79 ms per token,     3,59 tokens per second)
0.33.690.475 I slot print_timing: id  0 | task 0 |        eval time =   21391,78 ms /    54 tokens (  403,62 ms per token,     2,48 tokens per second)
0.33.690.475 I slot print_timing: id  0 | task 0 |       total time =   24458,49 ms /    65 tokens
0.33.690.478 I slot print_timing: id  0 | task 0 |    graphs reused =         53
0.33.690.495 I slot      release: id  0 | task 0 | stop processing: n_tokens = 64, truncated = 0
0.33.690.499 I srv  update_slots: all slots are idle
0.07.699.832 I sched_reserve:    Vulkan0 compute buffer size =   703,31 MiB
0.07.699.836 I sched_reserve: Vulkan_Host compute buffer size =    21,36 MiB
0.07.699.836 I sched_reserve: graph nodes  = 3847
0.07.699.837 I sched_reserve: graph splits = 1 (with bs=256), 2 (with bs=1)
static int64_t ggml_vk_get_op_batch_size(const ggml_tensor * op) {
    switch (op->op) {
        case GGML_OP_GET_ROWS:
        case GGML_OP_MUL_MAT:
            return op->ne[1];

llama compiled with GGML_VULKAN=ON

sudo dnf install cmake gcc-c++ vulkan-loader-devel vulkan-headers spirv-headers-devel openssl-devel glsl

mkdir ~/var
cd ~/var/
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp/
rm -rf build
cmake -B build -DGGML_VULKAN=ON
cmake --build build --config Release --parallel
sudo cmake --install build

sudo sh -c "echo '/home/$USER/var/llama.cpp/build/bin' > /etc/ld.so.conf.d/llama-cpp.conf"
sudo ldconfig

Further information

This feature was introduced with ggml-org/ggml@162fd7c

Environment

  • Macbook Air M2 - 24 GB RAM
  • Fedora Asahi Linux
  • Vulkan0 : Apple M2 (G14G B0)

First Bad Commit

No response

Relevant log output

No response

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions