Repository navigation
[Bug] ERNIE - white blank (diffusion-fa and gen-size affected) #1447
Description
Activity
I am also still getting this issue with
z-image-Q4_K_M.ggufI am piping this tool into Home-Assistant and it returns a white image
And the logs I get
[INFO ] stable-diffusion.cpp:776 - Using flash attention in the diffusion model |====================> | 453/1095 - 4.71GB/s |======================================> | 851/1095 - 3.25GB/s |==================================================| 1095/1095 - 3.04GB/s [INFO ] model.cpp:1006 - loading tensors completed, taking 2.37s (process: 0.00s, read: 0.49s, memcpy: 0.00s, convert: 0.15s, copy_to_backend: 0.55s) [INFO ] stable-diffusion.cpp:897 - total params memory size = 8549.12MB (VRAM 8549.12MB, RAM 0.00MB): text_encoders 3555.38MB(VRAM), diffusion_model 4833.74MB(VRAM), vae 160.00MB(VRAM), controlnet 0.00MB(VRAM), pmid 0.00MB(VRAM) [INFO ] stable-diffusion.cpp:981 - running in FLOW mode [INFO ] main.cpp:142 - listening on: 0.0.0.0:1234 [INFO ] stable-diffusion.cpp:3160 - generate_image 512x512 [INFO ] denoiser.hpp:499 - get_sigmas with discrete scheduler [INFO ] stable-diffusion.cpp:2736 - sampling using Euler method [INFO ] stable-diffusion.cpp:3090 - get_learned_condition completed, taking 17.68s [INFO ] stable-diffusion.cpp:3194 - generating image: 1/1 - seed 42 |==================================================| 20/20 - 1.28s/it [INFO ] stable-diffusion.cpp:3225 - sampling completed, taking 25.69s [INFO ] stable-diffusion.cpp:3243 - generating 1 latent images completed, taking 25.69s [INFO ] stable-diffusion.cpp:3114 - decoding 1 latents [INFO ] stable-diffusion.cpp:3130 - latent 1 decoded, taking 0.38s [INFO ] stable-diffusion.cpp:3134 - decode_first_stage completed, taking 0.38s [INFO ] stable-diffusion.cpp:3255 - generate_image completed in 43.77sPlease post the full verbose generation log, including the command line and backend initialization part. It's also useful to know where did the model files came from, and if it affects other model types on the same backend.
It's also useful to know if turning on/off certain flags help with the issue:
--diffusion-fa(or--fa),--offload-to-cpu,--clip-on-cpu,--vae-on-cpu.Turning off --diffusion-fa helped.
Result looks some artifact (text, cat legs), but any way high quality.

main model - https://huggingface.co/unsloth/ERNIE-Image-Turbo-GGUF/blob/main/ernie-image-turbo-Q8_0.gguf
original vae from ERNIE-Image-Turbo rep
llm - https://huggingface.co/unsloth/Ministral-3-3B-Instruct-2512-GGUF/blob/main/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf
(the same if using https://huggingface.co/Comfy-Org/ERNIE-Image/blob/main/text_encoders/ministral-3-3b.safetensors)Below will be result without diffusion-fa. Also will add in the top message log from white blank with diffusion-fa.
"../sd-cli.exe" --diffusion-model ../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf --vae ../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors --llm ../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf -p "a lovely cat holding a sign says 'ernie.cpp'" --cfg-scale 1.0 --steps 8 -H 1024 -W 1024 -v -o ../output/cli_out.jpg
[DEBUG] main.cpp:547 - version: stable-diffusion.cpp version unknown, commit 44cca3d
[DEBUG] main.cpp:548 - System Info:
SSE3 = 1 | AVX = 1 | AVX2 = 1 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | FMA = 1 | NEON = 0 | ARM_FMA = 0 | F16C = 1 | FP16_VA = 0 | WASM_SIMD = 0 | VSX = 0 |
[DEBUG] main.cpp:549 - SDCliParams {
mode: img_gen,
output_path: "../output/cli_out.jpg",
image_path: "",
metadata_format: "text",
verbose: true,
color: false,
canny_preprocess: false,
convert_name: false,
preview_method: none,
preview_interval: 1,
preview_path: "preview.png",
preview_fps: 16,
taesd_preview: false,
preview_noisy: false,
metadata_raw: false,
metadata_brief: false,
metadata_all: false
}
[DEBUG] main.cpp:550 - SDContextParams {
n_threads: 8,
model_path: "",
clip_l_path: "",
clip_g_path: "",
clip_vision_path: "",
t5xxl_path: "",
llm_path: "../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf",
llm_vision_path: "",
diffusion_model_path: "../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf",
high_noise_diffusion_model_path: "",
vae_path: "../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors",
taesd_path: "",
esrgan_path: "",
control_net_path: "",
embedding_dir: "",
embeddings: {
}
wtype: NONE,
tensor_type_rules: "",
lora_model_dir: ".",
photo_maker_path: "",
rng_type: cuda,
sampler_rng_type: NONE,
offload_params_to_cpu: false,
enable_mmap: false,
control_net_cpu: false,
clip_on_cpu: false,
vae_on_cpu: false,
flash_attn: false,
diffusion_flash_attn: false,
diffusion_conv_direct: false,
vae_conv_direct: false,
circular: false,
circular_x: false,
circular_y: false,
chroma_use_dit_mask: true,
qwen_image_zero_cond_t: false,
chroma_use_t5_mask: false,
chroma_t5_mask_pad: 1,
prediction: NONE,
lora_apply_mode: auto,
force_sdxl_vae_conv_scale: false
}
[DEBUG] main.cpp:551 - SDGenerationParams {
loras: "{
}",
high_noise_loras: "{
}",
prompt: "a lovely cat holding a sign says 'ernie.cpp'",
negative_prompt: "",
clip_skip: -1,
width: 1024,
height: 1024,
batch_count: 1,
init_image_path: "",
end_image_path: "",
mask_image_path: "",
control_image_path: "",
ref_image_paths: [],
control_video_path: "",
auto_resize_ref_image: true,
increase_ref_index: false,
pm_id_images_dir: "",
pm_id_embed_path: "",
pm_style_strength: 20,
skip_layers: [7, 8, 9],
sample_params: (txt_cfg: 1.00, img_cfg: 1.00, distilled_guidance: 3.50, slg.layer_count: 0, slg.layer_start: 0.01, slg.layer_end: 0.20, slg.scale: 0.00, scheduler: NONE, sample_method: NONE, sample_steps: 8, eta: inf, shifted_timestep: 0, flow_shift: inf),
high_noise_skip_layers: [7, 8, 9],
high_noise_sample_params: (txt_cfg: 7.00, img_cfg: 7.00, distilled_guidance: 3.50, slg.layer_count: 0, slg.layer_start: 0.01, slg.layer_end: 0.20, slg.scale: 0.00, scheduler: NONE, sample_method: NONE, sample_steps: 20, eta: inf, shifted_timestep: 0, flow_shift: inf),
custom_sigmas: [],
cache_mode: "",
cache_option: "",
cache: disabled (threshold=inf, start=0.15, end=0.95),
moe_boundary: 0.875,
video_frames: 1,
fps: 16,
vace_strength: 1,
strength: 0.75,
control_strength: 0.9,
seed: 42,
upscale_repeats: 1,
upscale_tile_size: 128,
vae_tiling_params: { 0, 0, 0, 0.5, 0, 0 },
}
[DEBUG] stable-diffusion.cpp:175 - Using CUDA backend
[INFO ] ggml_extend.hpp:81 - ggml_cuda_init: found 1 CUDA devices (Total VRAM: 12281 MiB):
[INFO ] ggml_extend.hpp:81 - Device 0: NVIDIA GeForce RTX 4070 Ti, compute capability 8.9, VMM: yes, VRAM: 12281 MiB
[INFO ] stable-diffusion.cpp:269 - loading diffusion model from '../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf'
[INFO ] model.cpp:229 - load ../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf using gguf format
[DEBUG] model.cpp:278 - init from '../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf'
[INFO ] stable-diffusion.cpp:316 - loading llm from '../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf'
[INFO ] model.cpp:229 - load ../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf using gguf format
[DEBUG] model.cpp:278 - init from '../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf'
[INFO ] stable-diffusion.cpp:330 - loading vae from '../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors'
[INFO ] model.cpp:232 - load ../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors using safetensors format
[DEBUG] model.cpp:307 - init from '../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors', prefix = 'vae.'
[INFO ] stable-diffusion.cpp:355 - Version: Ernie Image
[INFO ] stable-diffusion.cpp:383 - Weight type stat: f32: 203 | f16: 26 | q8_0: 410 | bf16: 254
[INFO ] stable-diffusion.cpp:384 - Conditioner weight type stat: f32: 53 | f16: 26 | q8_0: 157
[INFO ] stable-diffusion.cpp:385 - Diffusion model weight type stat: f32: 150 | q8_0: 253 | bf16: 6
[INFO ] stable-diffusion.cpp:386 - VAE weight type stat: bf16: 248
[DEBUG] stable-diffusion.cpp:388 - ggml tensor size = 400 bytes
[DEBUG] mistral_tokenizer.cpp:23 - vocab size: 131072
[DEBUG] mistral_tokenizer.cpp:31 - merges size 269443
[DEBUG] llm.hpp:697 - llm: num_layers = 26, vocab_size = 131072, hidden_size = 3072, intermediate_size = 9216
[INFO ] ernie_image.hpp:383 - ernie_image: layers = 36, hidden_size = 4096, heads = 32, ffn_hidden_size = 12288, in_channels = 128, out_channels = 128
[DEBUG] ggml_extend.hpp:2050 - ministral3.3b params backend buffer size = 4285.00 MB(VRAM) (236 tensors)
[DEBUG] ggml_extend.hpp:2050 - ernie_image params backend buffer size = 8292.08 MB(VRAM) (409 tensors)
[INFO ] stable-diffusion.cpp:681 - using VAE for encoding / decoding
[INFO ] auto_encoder_kl.hpp:517 - vae decoder: ch = 128
[DEBUG] ggml_extend.hpp:2050 - vae params backend buffer size = 94.72 MB(VRAM) (140 tensors)
[DEBUG] stable-diffusion.cpp:805 - loading weights
[DEBUG] model.cpp:755 - using 8 threads for model loading
[DEBUG] model.cpp:777 - loading tensors from ../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf
|======================> | 409/893 - 3.18GB/s
[DEBUG] model.cpp:777 - loading tensors from ../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf
|====================================> | 645/893 - 3.14GB/s
[DEBUG] model.cpp:777 - loading tensors from ../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors
|==================================================| 893/893 - 3.00GB/s
[INFO ] model.cpp:1012 - loading tensors completed, taking 4.12s (process: 0.00s, read: 2.45s, memcpy: 0.00s, convert: 0.03s, copy_to_backend: 1.09s)
[DEBUG] stable-diffusion.cpp:845 - finished loaded file
[INFO ] stable-diffusion.cpp:912 - total params memory size = 12671.79MB (VRAM 12671.79MB, RAM 0.00MB): text_encoders 4285.00MB(VRAM), diffusion_model 8292.08MB(VRAM), vae 94.72MB(VRAM), controlnet 0.00MB(VRAM), pmid 0.00MB(VRAM)
[INFO ] stable-diffusion.cpp:981 - running in FLOW mode
[INFO ] stable-diffusion.cpp:3160 - generate_image 1024x1024
[INFO ] denoiser.hpp:499 - get_sigmas with discrete scheduler
[INFO ] stable-diffusion.cpp:2736 - sampling using Euler method
[DEBUG] conditioner.hpp:1699 - parse 'a lovely cat holding a sign says 'ernie.cpp'' to [['a lovely cat holding a sign says 'ernie.cpp'', 1], ]
[DEBUG] bpe_tokenizer.cpp:183 - split prompt "a lovely cat holding a sign says 'ernie.cpp'" to tokens ["a", "Ġlovely", "Ġcat", "Ġholding", "Ġa", "Ġsign", "Ġsays", "Ġ'", "ern", "ie", ".cpp", "'", ]
[DEBUG] ggml_extend.hpp:1862 - ministral3.3b compute buffer size: 1.42 MB(VRAM)
[DEBUG] conditioner.hpp:1953 - computing condition graph completed, taking 87 ms
[INFO ] stable-diffusion.cpp:3090 - get_learned_condition completed, taking 0.09s
[INFO ] stable-diffusion.cpp:3194 - generating image: 1/1 - seed 42
[DEBUG] ggml_extend.hpp:1862 - ernie_image compute buffer size: 2452.20 MB(VRAM)
|==================================================| 8/8 - 3.56s/it
[INFO ] stable-diffusion.cpp:3225 - sampling completed, taking 28.96s
[INFO ] stable-diffusion.cpp:3245 - generating 1 latent images completed, taking 29.09s
[INFO ] stable-diffusion.cpp:3114 - decoding 1 latents
[DEBUG] ggml_extend.hpp:1862 - vae compute buffer size: 6658.00 MB(VRAM)
[DEBUG] vae.hpp:206 - computing vae decode graph completed, taking 0.96s
[INFO ] stable-diffusion.cpp:3130 - latent 1 decoded, taking 0.99s
[INFO ] stable-diffusion.cpp:3134 - decode_first_stage completed, taking 0.99s
[INFO ] stable-diffusion.cpp:3255 - generate_image completed in 30.66s
[INFO ] main.cpp:438 - save result image 0 to '../output/cli_out.jpg' (success)
[INFO ] main.cpp:487 - 1/1 images savedReacted by Erik Scholz- changed the title
[-][Bug] Ernie - white blank[/-][+][Bug] Ernie - white blank (with diffusion fa)[/+]on Apr 21, 2026 Lot's of paws with that cat 🤣
Addidtional way getting white blank
Without setting Sampler/Scheduler - White Blank
"sample_method": 'euler', "scheduler": 'simple' --> Normal generation
I will check sample_methods and schedulers and provide here an info
-
Unsloth Weights works only without -diffusion-fa, and some bad artifacts
-
Orig Weights working with -diffusion-fa and without it. Manual and accurate convert work as well.
-
Now Investigated that some bad generation failed several samplers and schedulers
currently - "discrete", "kl_optimal", "karras", "exponential" with 1024 - white blank, others - good
Reacted by KuhnChris-
- changed the title
[-][Bug] Ernie - white blank (with diffusion fa)[/-][+][Bug] Ernie - white blank (diffusion fa and gen-size affected)[/+]on Apr 21, 2026 - changed the title
[-][Bug] Ernie - white blank (diffusion fa and gen-size affected)[/-][+][Bug] ERNIE - white blank (diffusion-fa and gen-size affected)[/+]on Apr 21, 2026 I'll take a look at the samplers code later. Could you see if the preview images get blank right away, or at the end (or middle) of the generation?
@wbruna
sampler/scheduler - "default"
steps = 8
size = 1024x1024Vae Proj Step - 1 

Step - 2 

Step - 3 

Step - 4 

Step - 5 

Step - 6 

Step - 7 

Step - 8 

Final Image 

The problem on last steps looks same as mentioned here #1427 (comment)
This is the sigma schedule for 8 steps, omitting the first (always 1) and last (always 0) values:
scheduler 1 2 3 4 5 6 7 discrete 0.960045 0.909207 0.842338 0.750437 0.616212 0.401677 0.00398828 kl_optimal 0.798406 0.629932 0.483682 0.352475 0.231242 0.116136 0.00398804 karras 0.566517 0.305216 0.154863 0.0730525 0.0314801 0.0120888 0.00398803 exponential 0.454204 0.206302 0.0937031 0.0425604 0.0193311 0.00878027 0.00398804 bong_tangent 0.971458 0.929704 0.863435 0.745319 0.501994 0.21626 0.0733563 smoothstep 0.988892 0.955832 0.896461 0.800639 0.649923 0.426921 0.155477 simple 0.965517 0.923077 0.869565 0.8 0.705882 0.571429 0.363636 sgm_uniform 0.965555 0.923172 0.869747 0.80032 0.706436 0.572407 0.365484 The model seems to be choking on small sigma values. Happens even with the LCM sampler, which just takes the plain model output for the last step.
Reacted by Erik ScholzInteresting: Vulkan without Flash Attention works for me. With FA, the second step already produces scrambled output. Disabling FA doesn't seem to make any difference for ROCm.
Sounds like some value is sensitive to getting cast to f16.
As I investigated it's need a manual convert from original weights because some layers too sensitive to flash attention. With disabling diffusion-fa and manual tensor-type-rules I've got normal generation for all samplers/schedulers (1024 size).
below example default sampler/scheduler, size 1536x1024
on 12 gb vram I can't reach more sizes...
Reacted by Erik Scholzand manual tensor-type-rules
Can you share your rule with us?
disable diffusion-fa
+
fixed model
https://huggingface.co/Solictous/ERNIE-Image-Turbo-GGUF/blob/main/ernie-image-turbo_q4_0.gguf



Git commit
44cca3d
Operating System & Version
Win 10
GGML backends
CUDA
Command-line arguments used
chcp 65001 "../sd-cli.exe" ^ --diffusion-model ../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-UD-Q4_K_M.gguf ^ --vae ../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors ^ --llm ../../Models/Diffusers/Ministral/ministral-3-3b.safetensors ^ -p "a lovely cat holding a sign says 'ernie.cpp'" ^ --cfg-scale 1.0 ^ --steps 8 ^ --diffusion-fa ^ -H 1024 ^ -W 1024 ^ -v ^ -o ../output/cli_out.jpg pause
Steps to reproduce
Run any model of Ernie-Image
What you expected to happen
Images as on example
What actually happened
Getting white blank
Logs / error messages / stack trace
"../sd-cli.exe" --diffusion-model ../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf --vae ../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors --llm ../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf -p "a lovely cat holding a sign says 'ernie.cpp'" --cfg-scale 1.0 --steps 8 --diffusion-fa -H 1024 -W 1024 -v -o ../output/cli_out.jpg
[DEBUG] main.cpp:547 - version: stable-diffusion.cpp version unknown, commit 44cca3d
[DEBUG] main.cpp:548 - System Info:
SSE3 = 1 | AVX = 1 | AVX2 = 1 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | FMA = 1 | NEON = 0 | ARM_FMA = 0 | F16C = 1 | FP16_VA = 0 | WASM_SIMD = 0 | VSX = 0 |
[DEBUG] main.cpp:549 - SDCliParams {
mode: img_gen,
output_path: "../output/cli_out.jpg",
image_path: "",
metadata_format: "text",
verbose: true,
color: false,
canny_preprocess: false,
convert_name: false,
preview_method: none,
preview_interval: 1,
preview_path: "preview.png",
preview_fps: 16,
taesd_preview: false,
preview_noisy: false,
metadata_raw: false,
metadata_brief: false,
metadata_all: false
}
[DEBUG] main.cpp:550 - SDContextParams {
n_threads: 8,
model_path: "",
clip_l_path: "",
clip_g_path: "",
clip_vision_path: "",
t5xxl_path: "",
llm_path: "../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf",
llm_vision_path: "",
diffusion_model_path: "../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf",
high_noise_diffusion_model_path: "",
vae_path: "../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors",
taesd_path: "",
esrgan_path: "",
control_net_path: "",
embedding_dir: "",
embeddings: {
}
wtype: NONE,
tensor_type_rules: "",
lora_model_dir: ".",
photo_maker_path: "",
rng_type: cuda,
sampler_rng_type: NONE,
offload_params_to_cpu: false,
enable_mmap: false,
control_net_cpu: false,
clip_on_cpu: false,
vae_on_cpu: false,
flash_attn: false,
diffusion_flash_attn: true,
diffusion_conv_direct: false,
vae_conv_direct: false,
circular: false,
circular_x: false,
circular_y: false,
chroma_use_dit_mask: true,
qwen_image_zero_cond_t: false,
chroma_use_t5_mask: false,
chroma_t5_mask_pad: 1,
prediction: NONE,
lora_apply_mode: auto,
force_sdxl_vae_conv_scale: false
}
[DEBUG] main.cpp:551 - SDGenerationParams {
loras: "{
}",
high_noise_loras: "{
}",
prompt: "a lovely cat holding a sign says 'ernie.cpp'",
negative_prompt: "",
clip_skip: -1,
width: 1024,
height: 1024,
batch_count: 1,
init_image_path: "",
end_image_path: "",
mask_image_path: "",
control_image_path: "",
ref_image_paths: [],
control_video_path: "",
auto_resize_ref_image: true,
increase_ref_index: false,
pm_id_images_dir: "",
pm_id_embed_path: "",
pm_style_strength: 20,
skip_layers: [7, 8, 9],
sample_params: (txt_cfg: 1.00, img_cfg: 1.00, distilled_guidance: 3.50, slg.layer_count: 0, slg.layer_start: 0.01, slg.layer_end: 0.20, slg.scale: 0.00, scheduler: NONE, sample_method: NONE, sample_steps: 8, eta: inf, shifted_timestep: 0, flow_shift: inf),
high_noise_skip_layers: [7, 8, 9],
high_noise_sample_params: (txt_cfg: 7.00, img_cfg: 7.00, distilled_guidance: 3.50, slg.layer_count: 0, slg.layer_start: 0.01, slg.layer_end: 0.20, slg.scale: 0.00, scheduler: NONE, sample_method: NONE, sample_steps: 20, eta: inf, shifted_timestep: 0, flow_shift: inf),
custom_sigmas: [],
cache_mode: "",
cache_option: "",
cache: disabled (threshold=inf, start=0.15, end=0.95),
moe_boundary: 0.875,
video_frames: 1,
fps: 16,
vace_strength: 1,
strength: 0.75,
control_strength: 0.9,
seed: 42,
upscale_repeats: 1,
upscale_tile_size: 128,
vae_tiling_params: { 0, 0, 0, 0.5, 0, 0 },
}
[DEBUG] stable-diffusion.cpp:175 - Using CUDA backend
[INFO ] ggml_extend.hpp:81 - ggml_cuda_init: found 1 CUDA devices (Total VRAM: 12281 MiB):
[INFO ] ggml_extend.hpp:81 - Device 0: NVIDIA GeForce RTX 4070 Ti, compute capability 8.9, VMM: yes, VRAM: 12281 MiB
[INFO ] stable-diffusion.cpp:269 - loading diffusion model from '../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf'
[INFO ] model.cpp:229 - load ../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf using gguf format
[DEBUG] model.cpp:278 - init from '../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf'
[INFO ] stable-diffusion.cpp:316 - loading llm from '../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf'
[INFO ] model.cpp:229 - load ../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf using gguf format
[DEBUG] model.cpp:278 - init from '../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf'
[INFO ] stable-diffusion.cpp:330 - loading vae from '../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors'
[INFO ] model.cpp:232 - load ../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors using safetensors format
[DEBUG] model.cpp:307 - init from '../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors', prefix = 'vae.'
[INFO ] stable-diffusion.cpp:355 - Version: Ernie Image
[INFO ] stable-diffusion.cpp:383 - Weight type stat: f32: 203 | f16: 26 | q8_0: 410 | bf16: 254
[INFO ] stable-diffusion.cpp:384 - Conditioner weight type stat: f32: 53 | f16: 26 | q8_0: 157
[INFO ] stable-diffusion.cpp:385 - Diffusion model weight type stat: f32: 150 | q8_0: 253 | bf16: 6
[INFO ] stable-diffusion.cpp:386 - VAE weight type stat: bf16: 248
[DEBUG] stable-diffusion.cpp:388 - ggml tensor size = 400 bytes
[DEBUG] mistral_tokenizer.cpp:23 - vocab size: 131072
[DEBUG] mistral_tokenizer.cpp:31 - merges size 269443
[DEBUG] llm.hpp:697 - llm: num_layers = 26, vocab_size = 131072, hidden_size = 3072, intermediate_size = 9216
[INFO ] ernie_image.hpp:383 - ernie_image: layers = 36, hidden_size = 4096, heads = 32, ffn_hidden_size = 12288, in_channels = 128, out_channels = 128
[DEBUG] ggml_extend.hpp:2050 - ministral3.3b params backend buffer size = 4285.00 MB(VRAM) (236 tensors)
[DEBUG] ggml_extend.hpp:2050 - ernie_image params backend buffer size = 8292.08 MB(VRAM) (409 tensors)
[INFO ] stable-diffusion.cpp:681 - using VAE for encoding / decoding
[INFO ] auto_encoder_kl.hpp:517 - vae decoder: ch = 128
[DEBUG] ggml_extend.hpp:2050 - vae params backend buffer size = 94.72 MB(VRAM) (140 tensors)
[INFO ] stable-diffusion.cpp:776 - Using flash attention in the diffusion model
[DEBUG] stable-diffusion.cpp:805 - loading weights
[DEBUG] model.cpp:755 - using 8 threads for model loading
[DEBUG] model.cpp:777 - loading tensors from ../../Models/Diffusers/ERNIE-Image/ernie-image-turbo-Q8_0.gguf
|======================> | 409/893 - 3.23GB/s
[DEBUG] model.cpp:777 - loading tensors from ../../Models/Diffusers/Ministral/Ministral-3-3B-Instruct-2512-UD-Q8_K_XL.gguf
|====================================> | 645/893 - 3.24GB/s
[DEBUG] model.cpp:777 - loading tensors from ../../Models/Diffusers/Vae/vae_ernie_image_turbo.safetensors
|==================================================| 893/893 - 3.09GB/s
[INFO ] model.cpp:1012 - loading tensors completed, taking 4.01s (process: 0.00s, read: 2.39s, memcpy: 0.00s, convert: 0.04s, copy_to_backend: 1.03s)
[DEBUG] stable-diffusion.cpp:845 - finished loaded file
[INFO ] stable-diffusion.cpp:912 - total params memory size = 12671.79MB (VRAM 12671.79MB, RAM 0.00MB): text_encoders 4285.00MB(VRAM), diffusion_model 8292.08MB(VRAM), vae 94.72MB(VRAM), controlnet 0.00MB(VRAM), pmid 0.00MB(VRAM)
[INFO ] stable-diffusion.cpp:981 - running in FLOW mode
[INFO ] stable-diffusion.cpp:3160 - generate_image 1024x1024
[INFO ] denoiser.hpp:499 - get_sigmas with discrete scheduler
[INFO ] stable-diffusion.cpp:2736 - sampling using Euler method
[DEBUG] conditioner.hpp:1699 - parse 'a lovely cat holding a sign says 'ernie.cpp'' to [['a lovely cat holding a sign says 'ernie.cpp'', 1], ]
[DEBUG] bpe_tokenizer.cpp:183 - split prompt "a lovely cat holding a sign says 'ernie.cpp'" to tokens ["a", "Ġlovely", "Ġcat", "Ġholding", "Ġa", "Ġsign", "Ġsays", "Ġ'", "ern", "ie", ".cpp", "'", ]
[DEBUG] ggml_extend.hpp:1862 - ministral3.3b compute buffer size: 1.42 MB(VRAM)
[DEBUG] conditioner.hpp:1953 - computing condition graph completed, taking 85 ms
[INFO ] stable-diffusion.cpp:3090 - get_learned_condition completed, taking 0.09s
[INFO ] stable-diffusion.cpp:3194 - generating image: 1/1 - seed 42
[DEBUG] ggml_extend.hpp:1862 - ernie_image compute buffer size: 647.99 MB(VRAM)
|==================================================| 8/8 - 2.14s/it
[INFO ] stable-diffusion.cpp:3225 - sampling completed, taking 17.57s
[INFO ] stable-diffusion.cpp:3245 - generating 1 latent images completed, taking 17.69s
[INFO ] stable-diffusion.cpp:3114 - decoding 1 latents
[DEBUG] ggml_extend.hpp:1862 - vae compute buffer size: 6658.00 MB(VRAM)
[DEBUG] vae.hpp:206 - computing vae decode graph completed, taking 0.95s
[INFO ] stable-diffusion.cpp:3130 - latent 1 decoded, taking 0.99s
[INFO ] stable-diffusion.cpp:3134 - decode_first_stage completed, taking 0.99s
[INFO ] stable-diffusion.cpp:3255 - generate_image completed in 19.23s
[INFO ] main.cpp:438 - save result image 0 to '../output/cli_out.jpg' (success)
[INFO ] main.cpp:487 - 1/1 images saved
Additional context / environment details
No response