Skip to content

docs: align FP8 KV and placement claims - #2236

Open
localai-org-maint-bot wants to merge 1 commit into
mudler:mainfrom
localai-org-maint-bot:docs/current-code-audit
Open

docs: align FP8 KV and placement claims#2236
localai-org-maint-bot wants to merge 1 commit into
mudler:mainfrom
localai-org-maint-bot:docs/current-code-audit

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

What changed

  • update the README news to count all six expert-placement families
  • document the gated ROCm FP8 E4M3 KV store/read path
  • keep the CUDA device and end-to-end FP8 gates explicit
  • compact the expert-placement feature row to its current support and gaps

Why

The ROCm FP8 KV path landed in 191f64608, but the public feature and usage pages still said ROCm refused it. The dots3-note model also joined RunMoePlaced, while the README still counted five families.

Verification

  • python3 scripts/check-readme-structure.py
  • python3 tests/scripts/test_check_readme_structure.py
  • python3 scripts/check-agent-record.py
  • python3 tests/scripts/test_agent_record.py
  • git diff --check

No GPU verification is required because this PR changes documentation only and cites the existing accepted component gates.

ROCm now implements and gates the FP8 E4M3 KV store and read, while the public usage pages still described it as refused. dots3-note also joined the shared expert-placement path after the README counted five families.

Keep the unmeasured CUDA and end-to-end gates explicit, and compact the placement row to its current support and gap summary.

FOLLOWING_AGENTS_PROTOCOL

Assisted-by: Codex:gpt-5 [Codex]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant