Skip to content

Fix GGUF tensor alignment and validate model config - #20

Closed
dipeshbabu wants to merge 3 commits into
RightNow-AI:mainfrom
dipeshbabu:fix/gguf-alignment-and-config-validation
Closed

dipeshbabu wants to merge 3 commits into
RightNow-AI:mainfrom
dipeshbabu:fix/gguf-alignment-and-config-validation

Conversation

@dipeshbabu

@dipeshbabu dipeshbabu commented Feb 28, 2026 •

Copy link
Copy Markdown

GGUF loading now aligns tensor data for the declared alignment and rejects unsupported layer counts, missing feed-forward dimensions, and incompatible attention/KV-head ratios before inference. Missing KV-head metadata defaults to ordinary multi-head attention.

Attention accumulation uses the existing output buffer, allowing head dimensions above 256 without a fixed-size stack array or another allocation. Failed loads release the mapped file.

Validation: make -C picolm native passed on Linux. Synthetic GGUF load/forward tests passed under AddressSanitizer and UndefinedBehaviorSanitizer, covering two alignments, a 512-wide head, and four invalid dimension configurations. No model download was needed. The upstream Build PicoLLM workflow currently requires maintainer approval.

@dipeshbabu

Copy link
Copy Markdown
Author

This also fixes for avoiding per-token malloc/leak by using mmap string views.

@dipeshbabu dipeshbabu closed this by deleting the head repository Sep 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant