Skip to content

docs: align GLM-5.3 public status - #2374

Open
localai-org-maint-bot wants to merge 1 commit into
mudler:mainfrom
localai-org-maint-bot:row/DOCS-CODE-ALIGNMENT
Open

docs: align GLM-5.3 public status#2374
localai-org-maint-bot wants to merge 1 commit into
mudler:mainfrom
localai-org-maint-bot:row/DOCS-CODE-ALIGNMENT

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Collaborator

What changed

  • Update the public architecture count from 43 to 44.
  • Add the current GLM-5.3 and GLM-5.3-Flash states to README News and the supported-model summary.
  • Record CUDA keep-quant support for IQ2_XS and IQ4_XS without claiming a CUDA model forward.
  • Replace the two GLM feature-table attempt logs with concise current-state rows and explicit gaps.

Why

Recent GLM-5.3 model and quantization landings made the public overview stale. The README still described the GLM-5.3-Flash forward as incomplete, and docs/FEATURES.md still said GLM-5.3 loaded and forwarded nothing.

Verification

  • python3 scripts/check-readme-structure.py
  • python3 tests/scripts/test_check_readme_structure.py
  • python3 scripts/check-supported-models.py
  • python3 tests/scripts/test_check_supported_models.py
  • python3 scripts/check-agent-record.py
  • python3 scripts/check-model-checklist.py
  • python3 scripts/check-env-doc.py
  • python3 scripts/check-commit-trailers.py --range upstream/main..HEAD
  • git diff --check

scripts/agent-preflight.sh passes all documentation and record gates. Two host-dependent checks remain unavailable: release-workflow validation needs PyYAML, and test-registration configuration aborts in the host toolchain. This documentation-only change does not modify either surface.

Claims

No new speed or oracle-parity claim is made. GLM-5.3-Flash is described as coherent CPU output, and GLM-5.3 as a synthetic first-token forward only.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:gpt-5.6-sol [Codex]

The GLM-5.3 model landings moved the registry to 44 architectures, but
the public overview still reported 43. The GLM feature rows also kept
superseded refusal history after both forwards landed.

Report the current CPU-only states and name the remaining gates. Record
the CUDA keep-quant support without making a speed or parity claim.

FOLLOWING_AGENTS_PROTOCOL

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: Codex:gpt-5.6-sol [Codex]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant