Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
99 changes: 99 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -194,6 +194,105 @@ The script loads your prediction file, makes API calls using the models specifie
> [!NOTE]
> - For robustness evaluation, we only measure the model-selection flip ratio after adding noise to the original prompt, so no additional LLM inference is required for this stage.

### Kilo Auto Efficient: combined routing and generation

The standalone driver uses Kilo's public OpenAI-compatible endpoint with
`model: "kilo-auto/efficient"`. It does not train or tune on RouterArena.
Prepare the official datasets from the repository root first:

```bash
python -m pip install 'datasets>=4.0.0' 'pandas>=2.3.2'
python scripts/process_datasets/prep_datasets.py
```

The preparation script downloads RouterArena and LiveCodeBench data. The driver
itself uses only Python's standard library and this repository's model registry.
Set `KILO_API_KEY` securely in the environment; do not put it in a command, config,
or prediction file. To bill authorized organization credits, optionally set
`KILO_ORGANIZATION_ID` securely in the environment; it is sent only as the
`x-kilocode-organizationid` request header, not recorded in submission metadata.
The endpoint is also supplied through the environment:

```bash
export KILO_API_URL=https://api.kilo.ai/api/gateway/chat/completions
python scripts/submit_kilo_auto_efficient.py full --concurrency 4
python scripts/submit_kilo_auto_efficient.py robustness --concurrency 4
```

Each request sends the prepared prompt unchanged as one user message.
Streaming transport avoids the gateway's non-streaming header timeout.
The driver does not override temperature, token limits, seeds, or reasoning settings.
Routing decisions use the response's concrete `model`, not the router alias.
Unknown or out-of-pool models stop the run; the driver does not guess aliases.
Full submissions contain the returned answer and authoritative token usage.

Robustness streams close after the first concrete model appears.
Partial inference can still incur costs; the submission omits `generated_result`.
Completion duration is not routing-only latency.
The stock evaluator excludes classifier overhead, robustness calls, and discarded retry attempts.

Provider errors, empty answers, and missing usable usage remain `success: false` rows.
The driver records each failure and does not omit the row.
If final usage is missing or unusable, `token_usage` stays null.
The driver does not invent tokens; the stock evaluator cannot account for those costs.

Raw completions, stream events, and early selection chunks are checkpointed in
`cached_results/kilo-auto-efficient/<split>.jsonl`. Rerun the same command to
resume without repeating checkpointed calls. Do not edit the dataset or pool
during a run. A request without a valid model selection stops new work.
In-flight calls are drained and checkpointed. Preserve the journal with
the submission for auditing. A forced process/VM termination can lose an
in-flight response and therefore cause another paid request on resume.

After the full first pass, explicitly retry provider errors:

```bash
python scripts/submit_kilo_auto_efficient.py full --retry-provider-errors --concurrency 4
```

Each pass makes one additional attempt per eligible query, up to three captured
attempts in total. Only explicit timeouts and HTTP 429/500/502/503/504 errors qualify.
Empty answers and missing usage alone do not qualify. Correctness scores are never read.
Successful generations are never retried, even when their answers are incorrect.
Every explicit retry remains in the journal, including transport errors.
Final rows retain the first success or the last captured generation failure.

To explicitly retry every failed generation once, including empty answers, use:

```bash
python scripts/submit_kilo_auto_efficient.py full --retry-failed-generations --concurrency 4
```

Each invocation makes one additional attempt per failed query, beyond the provider-only cap.
Successful generations remain unchanged, including incorrect answers.
The journal marks these attempts with `retry_policy: "failed_generation"`.
This flag does not establish that an empty answer was a provider error.

Use `--limit 1 --concurrency 1` for a one-request smoke run; it produces an
explicitly incomplete file and that response is reused on the subsequent full
run. Exported JSON is rebuilt from the journal on resume, so evaluate only after
generation finishes. Do not run the stock LLM inference script afterward:
it would bypass Kilo's router and generate again.

All eleven concrete models also use `ModelInference`'s Kilo provider.
For separate model inference, set `KILO_API_KEY` and optionally `KILO_ORGANIZATION_ID`.
That path calls the concrete model without Auto Efficient's selected reasoning settings.
This submission does not use that path.

Validate the completed files:

```bash
python router_inference/check_config_prediction_files.py kilo-auto-efficient full --check-generated-result
python router_inference/check_config_prediction_files.py kilo-auto-efficient robustness
```

Official submission requires 8,400 full rows and 420 robustness rows. Local
evaluation follows Step 4 below; it needs the evaluation dependencies and ground
truth, and LiveCodeBench executes generated code, so use an isolated environment.

Use `--num-workers 1` if `filelock` reports an unsafe fork during parallel grading.
Regrade the unchanged answers; do not regenerate them.

## 4. Run Router Evaluation

As the last step, run the evaluation script:
Expand Down
81 changes: 81 additions & 0 deletions llm_inference/model_inference.py
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,8 @@ def infer(
return self._call_perplexity(model_name, prompt)
elif provider == "openrouter":
return self._call_openrouter(model_name, prompt)
elif provider == "kilo":
return self._call_kilo(model_name, prompt)
elif provider == "replicate":
return self._call_replicate(model_name, prompt)
elif provider == "aws":
Expand Down Expand Up @@ -198,6 +200,18 @@ def _get_provider(self, model_name: str) -> str:
"mistralai/devstral-2512:free": "openrouter",
"meta-llama/llama-3.3-70b-instruct": "openrouter",
"meta-llama/llama-3.1-405b-instruct": "openrouter",
# Kilo Auto Efficient's concrete production pool
"anthropic/claude-sonnet-5": "kilo",
"deepseek/deepseek-v4.1-flash": "kilo",
"meta/muse-spark-1.3-contributor": "kilo",
"minimax/minimax-m3": "kilo",
"moonshotai/kimi-k3": "kilo",
"openai/gpt-5.6-luna": "kilo",
"thinkingmachines/inkling": "kilo",
"x-ai/grok-4.6": "kilo",
"xiaomi/mimo-v2.5-pro": "kilo",
"z-ai/glm-5.3": "kilo",
"z-ai/glm-5.3-flash": "kilo",
# Replicate
"meta/codellama-34b-instruct": "replicate",
# AWS Bedrock
Expand Down Expand Up @@ -313,6 +327,73 @@ def _call_replicate(self, model_name: str, prompt: str) -> Dict[str, Any]:
"provider": "replicate",
}

def _call_kilo(self, model_name: str, prompt: str) -> Dict[str, Any]:
"""Generate with a concrete Kilo model, without invoking the router."""
from universal_model_names import ModelNameManager

api_key = os.getenv("KILO_API_KEY")
if not api_key:
raise ValueError("KILO_API_KEY is required for Kilo model inference")
headers = {}
organization_id = os.getenv("KILO_ORGANIZATION_ID")
if organization_id:
headers["x-kilocode-organizationid"] = organization_id
content = []
usage = None
served_model = None
finish_reason = None
with OpenAI(
base_url="https://api.kilo.ai/api/gateway",
api_key=api_key,
timeout=600,
max_retries=0,
) as client:
stream = client.chat.completions.create(
model=model_name,
messages=[{"role": "user", "content": prompt}],
stream=True,
stream_options={"include_usage": True},
extra_headers=headers,
)
with stream:
for chunk in stream:
if chunk.model:
served_model = chunk.model
if chunk.usage is not None:
usage = chunk.usage
for choice in chunk.choices:
if choice.index != 0:
continue
if choice.delta.content is not None:
content.append(choice.delta.content)
if choice.finish_reason is not None:
finish_reason = choice.finish_reason
if not served_model:
raise ValueError("Kilo stream did not identify the served model")
answer = "".join(content)
success = (
bool(answer.strip()) and usage is not None and usage.completion_tokens > 0
)
return {
"response": answer,
"success": success,
"token_usage": {
"input_tokens": usage.prompt_tokens,
"output_tokens": usage.completion_tokens,
"total_tokens": usage.total_tokens,
}
if usage is not None
else None,
"model_used": ModelNameManager.get_universal_name(served_model),
"response_model": served_model,
"usage_raw": usage.model_dump() if usage is not None else None,
"finish_reason": finish_reason,
"provider": "kilo",
"error": None
if success
else "Provider omitted an answer or usable token accounting",
}

def _call_openrouter(self, model_name: str, prompt: str) -> Dict[str, Any]:
"""Call OpenRouter API."""
openrouter_api_key = os.getenv("OPENROUTER_API_KEY")
Expand Down
44 changes: 44 additions & 0 deletions model_cost/model_cost.json
Original file line number Diff line number Diff line change
Expand Up @@ -383,6 +383,50 @@
"input_token_price_per_million": 1.25,
"output_token_price_per_million": 10.0
},
"anthropic/claude-sonnet-5": {
"input_token_price_per_million": 2.0,
"output_token_price_per_million": 10.0
},
"deepseek/deepseek-v4.1-flash": {
"input_token_price_per_million": 0.3,
"output_token_price_per_million": 1.2
},
"meta/muse-spark-1.3-contributor": {
"input_token_price_per_million": 0.1,
"output_token_price_per_million": 0.2
},
"minimax/minimax-m3": {
"input_token_price_per_million": 0.3,
"output_token_price_per_million": 1.2
},
"moonshotai/kimi-k3": {
"input_token_price_per_million": 3.0,
"output_token_price_per_million": 15.0
},
"openai/gpt-5.6-luna": {
"input_token_price_per_million": 0.2,
"output_token_price_per_million": 1.2
},
"thinkingmachines/inkling": {
"input_token_price_per_million": 0.95,
"output_token_price_per_million": 4.05
},
"x-ai/grok-4.6": {
"input_token_price_per_million": 2.0,
"output_token_price_per_million": 6.0
},
"xiaomi/mimo-v2.5-pro": {
"input_token_price_per_million": 0.435,
"output_token_price_per_million": 0.87
},
"z-ai/glm-5.3": {
"input_token_price_per_million": 1.4,
"output_token_price_per_million": 4.4
},
"z-ai/glm-5.3-flash": {
"input_token_price_per_million": 0.15,
"output_token_price_per_million": 0.5
},
"google/gemma-4-31b-it": {
"input_token_price_per_million": 0.08,
"output_token_price_per_million": 0.35
Expand Down
33 changes: 33 additions & 0 deletions router_inference/config/kilo-auto-efficient.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
{
"pipeline_params": {
"router_name": "kilo-auto-efficient",
"models": [
"anthropic/claude-sonnet-5",
"deepseek/deepseek-v4.1-flash",
"meta/muse-spark-1.3-contributor",
"minimax/minimax-m3",
"moonshotai/kimi-k3",
"openai/gpt-5.6-luna",
"thinkingmachines/inkling",
"x-ai/grok-4.6",
"xiaomi/mimo-v2.5-pro",
"z-ai/glm-5.3",
"z-ai/glm-5.3-flash"
],
"gateway_model": "kilo-auto/efficient",
"model_pool_source": "https://api.kilo.ai/api/gateway/models",
"model_pool_snapshot_date": "2026-10-07",
"pricing_source": "https://api.kilo.ai/api/gateway/models",
"pricing_snapshot_date": "2026-10-07",
"description": "Kilo Auto Efficient, queried through the production gateway without RouterArena training or tuning. The pool preserves the gateway's provider-prefixed model IDs; predictions must record the selected model rather than substitute a cached model. Generate submissions with the standalone driver, not the stock router-class generator.",
"retry_policy": "The provider-only pass permits at most two additional captured attempts for explicit timeouts or HTTP 429/500/502/503/504 errors. An explicitly requested failed-generation pass makes one additional attempt for every failure, including empty answers, beyond that cap. Empty answers are not presumed to be provider errors. Prompts and router settings remain unchanged. Successful generations are never retried, and correctness scores are not consulted. The append-only journal retains all captured attempts and marks the explicit failed-generation pass.",
"evaluation_limitations": [
"RouterArena computes generation cost from generated_result.token_usage at the selected or actually served model's registered input/output rates. It does not include classifier or other routing overhead; these costs must be disclosed separately and must not be represented as fabricated generation tokens.",
"The registered prices are uncached USD rates per million tokens from the production catalog snapshot. RouterArena does not account for separate input-cache read/write prices or per-request and web-search charges.",
"The live smoke used BYOK and reported usage.cost=0 with a nonzero usage.market_cost. Evaluation uses registered catalog market rates rather than treating the BYOK account charge as free inference. Reported market_cost is not evidence that classifier overhead is included.",
"Reasoning settings are not part of RouterArena's generic cache identity. Preserve production reasoning settings and authoritative token usage in each generated result; do not alias a reasoning variant to an existing cached base model.",
"Streaming transport changes delivery only; the production classifier state excludes the stream flag. Existing non-streaming results are preserved. Robustness streams close after the first concrete model appears, without changing the prompt or reasoning settings.",
"Provider generation errors remain success=false rows and count as failures. If the provider omits final token usage, token_usage is null. The stock evaluator excludes discarded retry attempts and cannot account for missing failed-inference usage, so its reported cost is incomplete."
]
}
}
Loading
Loading