Skip to content

Add AVX2 kernels for Q4_K and Q6_K - #34

Open
giankev wants to merge 1 commit into
RightNow-AI:mainfrom
giankev:experiments/avx2-kernels-clean
Open

giankev wants to merge 1 commit into
RightNow-AI:mainfrom
giankev:experiments/avx2-kernels-clean

Conversation

@giankev

@giankev giankev commented Sep 14, 2026

Copy link
Copy Markdown

What does this PR do?

Adds AVX2 + FMA implementations for the Q4_K and Q6_K fused dot-product kernels on x86.

The existing NEON and scalar implementations are preserved as fallbacks. AVX2 is enabled at compile time when both __AVX2__ and __FMA__ are available.

On an Intel Core i7-10510U with TinyLlama 1.1B Q4_K_M, generation improved from approximately 1.4 tok/s to 4.2–4.3 tok/s, around a 3x end-to-end speedup.

Type of change

  • Bug fix
  • New feature
  • Performance improvement
  • Documentation update
  • Refactoring (no behavior change)

Testing

  • Tested on x86-64 (Linux/macOS/Windows)
  • Tested on ARM64 (Raspberry Pi)
  • Tested with TinyLlama 1.1B Q4_K_M

Test command:

./picolm ./models/tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf \
  -p "The capital of Italy is" \
  -n 100 \
  -t 0 \
  -j 4

Output:

Rome.

2. Cities in Italy: Florence, Rome, Venice, Milan, Naples, Turin, Bologna, Verona, Palermo, and Genoa.

...

Prefill: 6 tokens in 1.25–1.59s
Generation: 101 tokens in 23.36–24.02s (4.2–4.3 tok/s)
Total: 24.64–25.61s
Memory: 45.17 MB runtime state (FP16 KV cache)

Additional validation:

  • Greedy output matches the scalar implementation.
  • Q6_K passed 1,000 randomized scalar-vs-AVX2 dot-product comparisons.
  • Five benchmark runs were performed for both baseline and AVX2 builds.
  • Baseline generation: 70.97–71.47s (~1.4 tok/s).
  • AVX2 Q4_K + Q6_K generation: 23.36–24.02s (~4.2–4.3 tok/s).
  • perf profiling showed the expected bottleneck migration from Q4_K to Q6_K after optimizing Q4_K.

Checklist

  • Code compiles without warnings (make native)
  • No new dependencies added
  • Memory usage not increased (check stderr output)
  • Works with --json mode (if touching generation/sampling)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant