Skip to content

feat(novita-ai): refresh and audit model catalog - #9243

Merged
rekram1-node merged 4 commits into
devfrom
novita-sync
Oct 9, 2026
Merged

rekram1-node merged 4 commits into
devfrom
novita-sync

Conversation

@rekram1-node

@rekram1-node rekram1-node commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Data-only Novita AI refresh, independently researched against the current public model library, authenticated model inventory, canonical lab/model cards, and direct inference requests. This separates catalog accuracy from the automation proposed in #7050; no sync code or workflow changes are included here.

  • Add 34 missing provider offerings, including the newer GLM, MiMo, DeepSeek, Qwen, Tencent, MiniMax, Ling, Macaron, StepFun, and Nemotron routes, both embedding models, both rerankers, and two inventory-listed legacy routes.
  • Add 27 complete canonical lab entries and migrate all retained Novita offerings to override-only base_model definitions.
  • Remove 55 retired routes only after exact 404 MODEL_NOT_FOUND responses, including the subsequently retired bare deepseek/deepseek-v4-flash route. Remove two additional obsolete Sao10K ID spellings that also returned MODEL_NOT_FOUND; their correctly cased replacement offerings are deliberately deferred to avoid a case-insensitive filesystem collision.
  • Keep the working but unlisted minimax/minimax-m2.7-highspeed route and preserve its curated pricing. Retain routes returning transient 429/503 responses.
  • Refresh host limits, modalities, pricing, cache-read rates, and advertised positive capabilities; correct stale lab release dates, descriptions, open-weight flags, and links through canonical inheritance.
  • Add MiniMax M3's context pricing tier at 524,288 tokens, remove the redundant MiMo V2.5 Pro tier, and correct reasoning controls per route.

The resulting provider catalog has 84 offerings, covering usable public catalog rows except the two deliberately deferred Sao10K offerings, while excluding directly verified retired routes. The public library includes four non-chat offerings omitted from the authenticated /models response. Two Ming image placeholder rows have zero context/output limits and insufficient model/pricing information, so they are not imported. Unadvertised private/internal inventory aliases are not blindly published.

Sources

Checked on 2026-10-09:

Canonical lab entries and provider overrides contain leading source comments. Prices use Novita's explicit USD-per-million decimal fields rather than treating its scaled integer fields as USD. The account-scoped inventory is not an authoritative deletion source.

Reasoning audit

Novita is a multi-model host/relay; controls are matched to the underlying model and actual Novita request surface, not to a universal OpenAI-compatible effort enum.

  • Live thinking.type=enabled|disabled probes verify the toggles on the supported routes by inspecting reasoning_content.
  • DeepSeek V4/dated V4/ V4.1 routes use the current first-party/same-surface low|high|max baseline, with live Novita acceptance of low. Fabricated minimal/medium/xhigh sets are removed.
  • GLM 5.2 and 5.3 and Kimi K3 use effort-only definitions containing none, because Novita supports disabling reasoning through that effort value. No redundant toggle is authored.
  • GLM 5.3 Flash ignores the tested disable controls; it has graded low|high|max, not a false toggle.
  • Qwen 3.5/3.6/3.7 honor thinking_budget=64. Harder-prompt checks verify 16/64-token budgets on Qwen 3.8 Flash, 27B, and 2.4T; Qwen 3.8 Max rejects budget requests. The open 2.4T route also rejects disabling thinking. Qwen 3.8 effort lists follow native low|medium|xhigh rather than every accepted alias.
  • Macaron Venti follows native high|max, with Novita's verified none off mode. Macaron Tall uses an effort-only none|high definition for its verified reasoning_effort=none|high off/on values, not toggle or thinking.type; other accepted labels were not demonstrated as distinct depth controls.
  • MiniMax M2.5/M2.7/highspeed and Kimi K2.7 Code remain always-on after disable-control probes; MiniMax M3 genuinely toggles.
  • R1 0528, R1 Turbo, and MiniMax M1 emit inline <think> content on the tested routes. No unsupported interleaved side channel is invented.
  • Qwen3 Omni Thinking retains the advertised thinking checkpoint's inherited reasoning capability with reasoning_options=[]. Novita rejects the tested enable/effort controls and exposes no separate reasoning on the tested prompts, but this is not sufficient evidence to declare the checkpoint non-reasoning or silently re-identify it as Instruct. Its text-only output has no audio-output price; the curated input-audio rate is preserved because inventory metadata does not return it.
  • GPT OSS 120B stays text-only despite Novita's inventory image label. Its native low/medium/high effort baseline is retained; accepted none did not suppress reasoning.

Known operational/source limitations

  • Baichuan M2 32B retains its lab reasoning capability and native OpenAI-compatible chat_template_kwargs.thinking_mode=on|off toggle baseline. Default, on/off, and other control probes returned overload 429s; the host toggle is not meaningfully live-verified. Missing inventory features are not negative capability evidence.
  • Ling 3.1 Flash and CoBuddy returned overload 429s, so their documented/peer controls are retained and explicitly marked as not meaningfully live-verified.
  • R1 0528 Qwen3 8B and Llama 3.2 1B Instruct returned 503 and are retained. The correctly cased Euryale route also returned 503; its deferral is for filesystem portability, not a claim that it is retired.
  • The bare deepseek/deepseek-v4-flash route initially failed only with explicit high, but follow-up checks now return 404 MODEL_NOT_FOUND for default, low, high, and max, and the route is absent from the current authenticated inventory. It is now removed on direct retirement evidence, not merely on an unsupported effort setting.
  • Both embedding routes answered (Qwen: 4096 dimensions; BGE: 1024). BGE reranking answered; Qwen reranking timed out. The latter's public mirror URL points to a different VL checkpoint, but its advertised ID/description/modalities identify the text Qwen3 Reranker 8B, which is what this entry represents.

Validation

  • bun validate — passed after each catalog batch.
  • bun test packages/core/test/schema.test.ts packages/core/test/meta.test.ts packages/core/test/generate-v2.test.ts packages/core/test/filter.test.ts — 24 passed, 0 failed.
  • Adding generate.test.ts gives 34 passed, 1 pre-existing failure: the repository-wide open-weight-link invariant reports the same 41 existing unlinked lab entries on origin/dev and this branch; none of the 27 new lab entries adds a failure.
  • Public-catalog reconciliation — only the two intentionally deferred Sao10K offerings are missing among non-retired positive-context rows; no chat context/output or input/output/cache-read pricing mismatches, and no identical inherited provider overrides. Non-generative embedding/reranking output cost is normalized to zero instead of copying the generic catalog output-price fields; their input rates remain the published USD/MTok prices.
  • git diff --check — passed.

The follow-up sync should preserve working unlisted routes, supplement /models with the public non-chat catalog, retain curated metadata and leading wire comments, and avoid deletions based on inventory absence or transient/parameter-specific failures. It should skip both deferred Sao10K offerings until exact API IDs can be represented without case-colliding directories.

Review follow-up

Corrected the DeepSeek effort baseline, Macaron Tall's effort-vs-toggle shape, the unsupported MiniMax highspeed last_updated override, the two insufficiently justified reasoning=false overrides, and the three non-generative output-price fields. Stronger-prompt follow-up probes for Baichuan and Qwen Omni are reflected above. Subsequent retirement checks removed the bare DeepSeek V4 Flash route; MiniMax M2.5 now has an explicit leading comment documenting the host's exact 131100 limit.

Live-verified GLM 5.3/Kimi K3/Tencent suppression behavior and Novita's explicitly reported MiniMax M2.5 output limit of 131,100 are retained. Neither is normalized to the lab's API behavior or a guessed round number.

The two correctly cased Sao10K offerings (Sao10K/L3-8B-Stheno-v3.2 and sao10k/l3-70b-euryale-v2.1) and their two newly added, otherwise unused lab entries have been left out of this PR by request. No case-colliding Novita directories remain in the Git tree, and no schema/loader changes were introduced. Their upstream IDs cannot simply be lowercased: future support needs a portable filename/exact API ID representation before importing or syncing them.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/novita-ai/models/deepseek/deepseek-v4-flash.toml:5 - Check: Relay effort baseline must copy lab low|high|max. Why: providers/deepseek/models/deepseek-v4-flash.toml:5 and deepseek-v4-pro.toml:2 are toggle + low|high|max, but this PR narrows to high|max while its own comment says effort=high returns MODEL_NOT_FOUND and low/max work. Action: Change to low|high|max or cite Novita docs proving high works and low does not; apply same fix to deepseek-v4-flash-0731.toml:3 and deepseek-v4-pro-0813.toml:3.
  • [high] [violation] providers/novita-ai/models/mindai/macaron-v1-tall.toml:6 - Check: toggle vs effort with none per AGENTS.md. Why: Wire comment is reasoning_effort = none|high, which is an effort control with none and must be effort only with no toggle; sibling macaron-v1-venti.toml correctly uses effort with none. Action: Replace [{type="toggle"}] with [{type="effort", values=["none","high"]}] or matching verified graded set.
  • [high] [violation] providers/novita-ai/models/baichuan/baichuan-m2-32b.toml:3 - Check: Override-only files must not flip reasoning to avoid reasoning_options. Why: New models/baichuan/baichuan-m2-32b.toml is reasoning=true, provider sets reasoning=false with only an assertion that Novita does not advertise reasoning and no wire test. Action: Restore inherited reasoning=true with verified reasoning_options, or provide live Novita evidence that this route never reasons.
  • [high] [violation] providers/novita-ai/models/qwen/qwen3-omni-30b-a3b-thinking.toml:6 - Check: Thinking checkpoint must not be silenced via reasoning=false. Why: Base models/alibaba/qwen3-omni-30b-a3b-thinking.toml is reasoning=true, provider sets reasoning=false while still pointing at the thinking checkpoint because enable controls were rejected. Action: Keep reasoning=true with [] if no caller control, or re-point base_model to the instruct checkpoint with evidence.
  • [high] [violation] providers/novita-ai/models/Sao10K/L3-8B-Stheno-v3.2.toml:1 - Check: Provider model IDs come from filenames with consistent casing. Why: Diff creates Sao10K/L3-8B-Stheno-v3.2.toml alongside sao10k/l3-70b-euryale-v2.1.toml after deleting sao10K/..., yielding three casings for the same lab and risking case-insensitive collision. Action: Normalize to lowercase sao10k/... paths or prove bun validate requires mixed-case IDs and document why the two corrected IDs use different casing.
  • [high] [possible mistake] providers/novita-ai/models/zai-org/glm-5.3.toml:4 - Check: Do not invent none beyond lab/peer baseline. Why: providers/zhipuai/models/glm-5.3.toml:7 is always-reasons low|high|max with no none; Novita flash in this PR matches that, but non-flash adds none contradicting native. Action: Remove none or cite Novita docs/live proof that none suppresses reasoning on this route.
  • [high] [possible mistake] providers/novita-ai/models/moonshotai/kimi-k3.toml:3 - Check: Do not invent none beyond lab/peer baseline. Why: providers/moonshotai/models/kimi-k3.toml:11 is always-thinks low|high|max with no none; PR adds none as Novita-specific with no citation. Action: Remove none or cite Novita docs/live proof.
  • [medium] [possible mistake] providers/novita-ai/models/tencent/hy3.toml:2 - Check: Effort set must come from lab + same-surface peers. Why: No providers/tencent baseline exists and lab models/tencent/hy3.toml exposes no controls, yet PR invents none|high with only a native-baseline assertion. Action: Cite Tencent/Novita docs or live reasoning_effort tests, or correct to verified controls; same for hy4-preview.toml:2.
  • [medium] [possible mistake] providers/novita-ai/models/baai/bge-m3.toml:6 - Check: Embedding costs must be USD/MTok and consistent for non-generative output. Why: Sets output=0.01 while new qwen/qwen3-embedding-8b.toml sets output=0 with a comment that output is not a generative limit. Action: Set output=0 or cite Novita billing showing output tokens are charged for embeddings; check bge-reranker-v2-m3 and qwen3-reranker-8b similarly.
  • [medium] [violation] providers/novita-ai/models/minimax/minimax-m2.7-highspeed.toml:6 - Check: Provider base_model files are override-only. Why: Overrides last_updated="2026-05-27" while models/minimax/MiniMax-M2.7-highspeed.toml:5 is 2026-03-18; host last_updated must inherit from lab, and this retains a stale value. Action: Drop the last_updated override to inherit.
  • [low] [possible mistake] providers/novita-ai/models/minimax/minimax-m2.5.toml:15 - Check: Provider limit deltas must be supported. Why: Sets output=131_100 vs lab MiniMax-M2.5.toml:14 131_072, a 28-token difference with no source. Action: Verify against Novita catalog and fix typo if needed.

@rekram1-node

Copy link
Copy Markdown
Collaborator Author

Addressed the actionable findings in 91f2a57 and updated the PR description:

  • DeepSeek V4 and the dated V4 routes now use the current low|high|max baseline. The parameter-specific V4 Flash high routing failure remains documented rather than causing model deletion.
  • Macaron Tall now has effort-only none|high, not a toggle.
  • MiniMax M2.7 highspeed inherits the canonical last_updated date.
  • Removed the two insufficiently justified reasoning=false overrides. Baichuan inherits reasoning and its native OpenAI-compatible toggle baseline, explicitly marked unverified on Novita after all nine follow-up probes returned overload 429s. Qwen Omni keeps the advertised Thinking checkpoint, reasoning_options=[], and no invented interleaved side channel; rejection of enable/effort controls and absence of separate reasoning are not proof that the checkpoint lacks reasoning capability.
  • Normalized output cost to zero for BGE embeddings and both rerankers because those endpoints return vectors/scores rather than generated completions. Published input rates are unchanged. This is a non-generative metadata normalization, not a claim that the generic catalog's output-price fields were billing-tested.

Retained evidence-backed values:

  • GLM 5.3: effort none produced no reasoning; high produced reasoning. Kimi K3: the same paired behavior was observed. These are Novita-specific capabilities, not assumptions copied from the lab's API.
  • Tencent: the native entries are under providers/tencent-tokenhub, and both HY3/HY4 live probes showed no reasoning with none and reasoning with high.
  • Sao10K casing: the exact public/API IDs are Sao10K/L3-8B-Stheno-v3.2 and sao10k/l3-70b-euryale-v2.1. The previous spellings returned MODEL_NOT_FOUND. The corrected Stheno route answered; Euryale returned a transient 503. Lowercasing the IDs would not be a valid normalization. Case-insensitive checkout portability is a separate concern.
  • MiniMax M2.5 output limit: both the public catalog and authenticated inventory explicitly report 131100, so the host override is retained.

Validation: bun validate passed; schema/meta/generate-v2/filter tests 24 passed, 0 failed; git diff --check passed.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Action items

  • [high] [violation] providers/novita-ai/models/Sao10K/L3-8B-Stheno-v3.2.toml:1 - Check: Provider model IDs must remain usable on case-insensitive checkouts. Why: This PR creates providers/novita-ai/models/Sao10K/ alongside providers/novita-ai/models/sao10k/ (l3-70b-euryale-v2.1.toml), two directories differing only by case that collide on macOS/Windows and risk one route overwriting the other. Action: Normalize to a single casing for the provider folder and preserve the exact upstream ID only in documentation/comments if needed, or prove the split is intentional and safe.
  • [medium] [possible mistake] providers/novita-ai/models/deepseek/deepseek-v4-flash.toml:6 - Check: reasoning_options effort values must reflect controls that work on this host. Why: The file lists effort values ["low", "high", "max"] while its own leading comment states effort=high currently returns MODEL_NOT_FOUND and only default/low/max work — self-contradictory evidence that high is advertised but fails. Action: Verify live reasoning_effort=high on this route and either remove high from values or provide evidence it is now accepted.
  • [low] [possible mistake] providers/novita-ai/models/minimax/minimax-m2.5.toml:12 - Check: Provider [limit] overrides must match host-advertised limits, not typos of inherited values. Why: Lab minimax/MiniMax-M2.5 sets output = 131_072 but the provider override sets output = 131_100 (+28 tokens) with no wire comment or mapped source explaining the non-power-of-two delta. Action: Verify the Novita-advertised output limit for this route and correct to 131_072 if 131_100 is a transcription error.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Action items

  • [medium] [possible mistake] providers/novita-ai/models/qwen/qwen3.8-27b.toml:17 - Check: Provider limit deltas must reflect this host's advertised limits. Why: Lab models/alibaba/qwen3.8-27b.toml is context = 262_144 / output = 32_768, but this new relay claims context = 1_000_000 / output = 131_072 with no limit source comment, matching the flagship qwen3.8-flash/max 1M family rather than the dense-27B card. Action: Verify Novita's advertised context/output for qwen-qwen3.8-27b and correct or cite the host delta; if 1M is wrong inherit the lab limit.
  • [medium] [possible mistake] providers/novita-ai/models/google/gemma-4-26b-a4b-it.toml:14 - Check: Provider output limits must not exceed the underlying checkpoint without host evidence. Why: Lab models/google/gemma-4-26b-a4b-it.toml (and gemma-4-31b-it.toml) is output = 32_768, but Novita retains output = 131_072 (4x) with only a model-detail URL. Hosts normally serve equal or smaller windows, not 4x larger outputs. Action: Verify Novita's max_output/inventory for both Gemma 4 routes and correct to the lab value or add a leading limit citation.
  • [medium] [possible mistake] providers/novita-ai/models/tencent/hy3.toml:17 - Check: Relay limits and reasoning controls must copy the lab baseline plus verified host deltas. Why: Lab models/tencent/hy3.toml is context = 256_000 / output = 128_000 text-only with no documented none|high effort control and no providers/tencent peer, but Novita claims context = 262_144 / output = 262_144 (output doubled to equal context) plus effort = ["none","high"] without host docs. Action: Verify Novita's advertised limits and that reasoning_effort = none|high actually toggles reasoning on this route; correct limits/options or cite the host evidence.
  • [medium] [possible mistake] providers/novita-ai/models/moonshotai/kimi-k3.toml:4 - Check: Do not add effort levels beyond lab/peers without host proof. Why: First-party providers/moonshotai/models/kimi-k3.toml states always thinks with effort = ["low","high","max"] only, but Novita adds none (["none","low","high","max"]) while noting it differs from the native API. Action: Verify that reasoning_effort = none truly suppresses reasoning on Novita's moonshotai-kimi-k3 route or drop none to match the native baseline.
  • [medium] [possible mistake] providers/novita-ai/models/zai-org/glm-5.3.toml:4 - Check: Do not add effort levels beyond lab/peers without host proof. Why: First-party providers/zhipuai/models/glm-5.3.toml states always reasons (thinking cannot be disabled) with effort = ["low","high","max"], but Novita adds none (["none","low","high","max"]). Action: Verify that none truly suppresses reasoning on Novita's zai-org-glm-5.3 route or drop none to match the Zhipu baseline.
  • [medium] [possible mistake] providers/novita-ai/models/baichuan/baichuan-m2-32b.toml:6 - Check: Every toggle needs the exact wire path verified on this host. Why: File claims Toggle: chat_template_kwargs.thinking_mode = on|off (native baseline) but admits Novita probes returned server_overload ... not live-verified, while Novita's documented generic controls are thinking.type/enable_thinking. The native path may not exist on Novita. Action: Verify the toggle on Novita's baichuan-baichuan-m2-32b route and correct the wire comment to the working Novita field, or remove reasoning_options if no caller control was demonstrated.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

No actionable findings.

@github-actions github-actions Bot added the reviewer: ready Automated review found no actionable items label Oct 9, 2026
@rekram1-node
rekram1-node merged commit 3e1c13d into dev Oct 9, 2026
2 checks passed
@rekram1-node
rekram1-node deleted the novita-sync branch October 9, 2026 17:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

reviewer: ready Automated review found no actionable items

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant