Repository navigation
feat(novita-ai): refresh and audit model catalog - #9243
Merged
Merged
Conversation
Contributor
Action items
|
Collaborator
Author
|
Addressed the actionable findings in 91f2a57 and updated the PR description:
Retained evidence-backed values:
Validation: |
Contributor
Action items
|
Contributor
Action items
|
Contributor
|
No actionable findings. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Data-only Novita AI refresh, independently researched against the current public model library, authenticated model inventory, canonical lab/model cards, and direct inference requests. This separates catalog accuracy from the automation proposed in #7050; no sync code or workflow changes are included here.
base_modeldefinitions.404 MODEL_NOT_FOUNDresponses, including the subsequently retired baredeepseek/deepseek-v4-flashroute. Remove two additional obsolete Sao10K ID spellings that also returnedMODEL_NOT_FOUND; their correctly cased replacement offerings are deliberately deferred to avoid a case-insensitive filesystem collision.minimax/minimax-m2.7-highspeedroute and preserve its curated pricing. Retain routes returning transient 429/503 responses.The resulting provider catalog has 84 offerings, covering usable public catalog rows except the two deliberately deferred Sao10K offerings, while excluding directly verified retired routes. The public library includes four non-chat offerings omitted from the authenticated
/modelsresponse. Two Ming image placeholder rows have zero context/output limits and insufficient model/pricing information, so they are not imported. Unadvertised private/internal inventory aliases are not blindly published.Sources
Checked on 2026-10-09:
Canonical lab entries and provider overrides contain leading source comments. Prices use Novita's explicit USD-per-million decimal fields rather than treating its scaled integer fields as USD. The account-scoped inventory is not an authoritative deletion source.
Reasoning audit
Novita is a multi-model host/relay; controls are matched to the underlying model and actual Novita request surface, not to a universal OpenAI-compatible effort enum.
thinking.type=enabled|disabledprobes verify the toggles on the supported routes by inspectingreasoning_content.low|high|maxbaseline, with live Novita acceptance oflow. Fabricatedminimal/medium/xhighsets are removed.none, because Novita supports disabling reasoning through that effort value. No redundant toggle is authored.low|high|max, not a false toggle.thinking_budget=64. Harder-prompt checks verify 16/64-token budgets on Qwen 3.8 Flash, 27B, and 2.4T; Qwen 3.8 Max rejects budget requests. The open 2.4T route also rejects disabling thinking. Qwen 3.8 effort lists follow nativelow|medium|xhighrather than every accepted alias.high|max, with Novita's verifiednoneoff mode. Macaron Tall uses an effort-onlynone|highdefinition for its verifiedreasoning_effort=none|highoff/on values, nottoggleorthinking.type; other accepted labels were not demonstrated as distinct depth controls.<think>content on the tested routes. No unsupported interleaved side channel is invented.reasoning_options=[]. Novita rejects the tested enable/effort controls and exposes no separate reasoning on the tested prompts, but this is not sufficient evidence to declare the checkpoint non-reasoning or silently re-identify it as Instruct. Its text-only output has no audio-output price; the curated input-audio rate is preserved because inventory metadata does not return it.nonedid not suppress reasoning.Known operational/source limitations
chat_template_kwargs.thinking_mode=on|offtoggle baseline. Default, on/off, and other control probes returned overload 429s; the host toggle is not meaningfully live-verified. Missing inventory features are not negative capability evidence.deepseek/deepseek-v4-flashroute initially failed only with explicithigh, but follow-up checks now return404 MODEL_NOT_FOUNDfor default, low, high, and max, and the route is absent from the current authenticated inventory. It is now removed on direct retirement evidence, not merely on an unsupported effort setting.Validation
bun validate— passed after each catalog batch.bun test packages/core/test/schema.test.ts packages/core/test/meta.test.ts packages/core/test/generate-v2.test.ts packages/core/test/filter.test.ts— 24 passed, 0 failed.generate.test.tsgives 34 passed, 1 pre-existing failure: the repository-wide open-weight-link invariant reports the same 41 existing unlinked lab entries onorigin/devand this branch; none of the 27 new lab entries adds a failure.git diff --check— passed.The follow-up sync should preserve working unlisted routes, supplement
/modelswith the public non-chat catalog, retain curated metadata and leading wire comments, and avoid deletions based on inventory absence or transient/parameter-specific failures. It should skip both deferred Sao10K offerings until exact API IDs can be represented without case-colliding directories.Review follow-up
Corrected the DeepSeek effort baseline, Macaron Tall's effort-vs-toggle shape, the unsupported MiniMax highspeed
last_updatedoverride, the two insufficiently justifiedreasoning=falseoverrides, and the three non-generative output-price fields. Stronger-prompt follow-up probes for Baichuan and Qwen Omni are reflected above. Subsequent retirement checks removed the bare DeepSeek V4 Flash route; MiniMax M2.5 now has an explicit leading comment documenting the host's exact 131100 limit.Live-verified GLM 5.3/Kimi K3/Tencent suppression behavior and Novita's explicitly reported MiniMax M2.5 output limit of 131,100 are retained. Neither is normalized to the lab's API behavior or a guessed round number.
The two correctly cased Sao10K offerings (
Sao10K/L3-8B-Stheno-v3.2andsao10k/l3-70b-euryale-v2.1) and their two newly added, otherwise unused lab entries have been left out of this PR by request. No case-colliding Novita directories remain in the Git tree, and no schema/loader changes were introduced. Their upstream IDs cannot simply be lowercased: future support needs a portable filename/exact API ID representation before importing or syncing them.