diff --git a/.claude/skills/benchmark-gql/SKILL.md b/.claude/skills/benchmark-gql/SKILL.md index bd339dae..bbc8bc63 100644 --- a/.claude/skills/benchmark-gql/SKILL.md +++ b/.claude/skills/benchmark-gql/SKILL.md @@ -50,7 +50,7 @@ CPU Time はリクエストと `cf-ray` で突き合わせる。`wrangler tail` 食い違っていたら、その旨をユーザーに伝えてから続けるか止めるかを決める。 -2. **ベンチを回す。** 既定は全 23 ケース × 15 反復 × 2 環境で、2〜3 分。 +2. **ベンチを回す。** 既定は `queries.json` の全ケース × 15 反復 × 2 環境で、2〜3 分。 本番側が遅いクエリを抱えているとその分伸びる。 ```bash diff --git a/.claude/skills/create-pr/SKILL.md b/.claude/skills/create-pr/SKILL.md index 96ceeef9..729859c4 100644 --- a/.claude/skills/create-pr/SKILL.md +++ b/.claude/skills/create-pr/SKILL.md @@ -35,44 +35,9 @@ description: Create a GitHub pull request for TrainLCD StationAPI that conforms ## 前提条件 - カレントディレクトリが `git rev-parse --show-toplevel` で解決できるリポジトリ内。 -- `gh` CLI が認証済み。 -- `head` ブランチが origin に push 済み。未 push の場合はユーザーに push の可否を確認する(勝手に push しない)。散文の前提で終わらせず、手順 2 でローカルと origin の commit ID を突き合わせて機械的に検出する。 -- **ref 名をシェルソースへ直接埋め込まない。** 本書の `` / `` は説明用のプレースホルダ。実際のコマンドでは値を `BASE_REF` / `HEAD_REF` に取り込み、以降は必ず `"$BASE_REF"` / `"$HEAD_REF"` で参照する。git の ref 名は `'` / `$( )` / バッククォート / `;` を含められるため、リテラルを直接置換すると構文が壊れるか、意図しないコマンドが実行される。値はコマンド出力から取り込む(ユーザー指定がある場合のみ、その値を代入する): - - ```bash - BASE_REF="$(gh repo view --json defaultBranchRef -q .defaultBranchRef.name)" - - HEAD_REF="$(git rev-parse --abbrev-ref HEAD)" - if [ "$HEAD_REF" = "HEAD" ] || [ "$HEAD_REF" = "$BASE_REF" ]; then - printf 'PR の head にできるブランチに居ない: %s\n' "$HEAD_REF" >&2 - exit 1 # 手順 1 で切り出す - fi - - # ref 名の文字種を検証する。gh --head・ファイル名 slug・シェル展開の - # すべてで安全に使える集合か。BASE_REF / HEAD_REF に同じ規則を適用する - validate_ref() { - case "$1" in - '' | *[!A-Za-z0-9._/-]*) - printf 'ref 名に想定外の文字が含まれる: %s\n' "$1" >&2; return 1 ;; - esac - } - validate_ref "$BASE_REF" || exit 1 - validate_ref "$HEAD_REF" || exit 1 - - # origin 上のブランチを commit ID へ解決する。 - # 実際の解決は fetch 後(手順 2)に一度だけ行うので、ここでは定義のみ - resolve_remote_rev() { - validate_ref "$1" || return 1 - rev="$(git rev-parse --verify --quiet "refs/remotes/origin/$1")" - if [ -z "$rev" ]; then - printf 'origin 上に存在しない: %s\n' "$1" >&2; return 1 - fi - printf '%s\n' "$rev" - } - ``` - -- **ref は `refs/remotes/origin/<名前>` / `refs/heads/<名前>` の完全形で解決する。** 短縮名を `git rev-parse` に渡すと、同名のタグやローカルブランチが優先されて意図と違うコミットを指すことがある。`--verify --quiet` を付けると、存在しない ref は終了コード 1 と空出力になるので、**解決結果が空でないことの検証が必須**。省くと未 push のブランチが空の `BASE_REV` / `HEAD_REV` として無言で通過する。 -- **`BASE_REF` / `HEAD_REF` が `^[A-Za-z0-9._/-]+$` に一致しない場合は自動で進めない。** git の ref 名には `gh --head` やファイル名 slug に使えない文字、シェルに解釈される文字が入り得る。`gh repo view` 由来の `BASE_REF` にも同じ検査を適用し(`resolve_remote_rev` の `case` がこれを担う)、一致しない値が返ったらユーザーに正しいブランチ名を確認する。 +- `gh` CLI が認証済みで、Python 3 がある(`prepare.py` は標準ライブラリしか使わない)。 +- `head` ブランチが origin に push 済み。未 push の場合はユーザーに push の可否を確認する(勝手に push しない)。未 push かどうかは `prepare.py` がローカルと origin の commit ID を突き合わせて検出する。 +- **ref の検証・解決、差分の取得、変更の種類の判定は `prepare.py` に任せ、シェルで書き直さない。** `prepare.py` は git と gh をシェルを通さずに起動する。そのため `'` / `$( )` / `;` を含む ref 名が実行されることも、zsh が `"$BASE_REF:refs/..."` の `:r` を修飾子として展開して ref を壊すこともない。スキル本文に位置引数(`$` と数字)を書くと、Claude Code がスキルを読み込む時点で呼び出し時の引数に置き換えるので、シェル断片でも使わない。 ## 手順 @@ -108,107 +73,44 @@ description: Create a GitHub pull request for TrainLCD StationAPI that conforms ``` - **`git add -A` / `git add .` は使わない。** `.gitignore` に載っていない一時ファイルまで巻き込む。`git status` の出力を読んでから、追跡済みは `git add -u`、未追跡は明示パスで追加する。関係ないファイルが入ったら `git restore --staged ` で外す。 - 変更が既にコミット済みでブランチだけが無い(`dev` の上に直接コミットした等)場合は、`git switch -c ` だけでそのコミットを新ブランチへ引き継げる。`dev` 側を元に戻す必要があれば、**作業ツリーに触れない `git branch -f dev origin/dev`** を使う(実行の可否はユーザーに確認する)。`dev` が別の worktree で checkout 済みなら git 自身がこのコマンドを拒否するので、取り違えも起きない。`git switch dev && git reset --hard origin/dev` は避ける: 追跡済みファイルの staged/unstaged 変更を問答無用で捨てるうえ、`dev` を別の worktree が持っていると `git switch` 自体が失敗する。どうしても checkout して戻すなら、`git worktree list` で `dev` の所在を確認し、その worktree で `git status --short` が空であることを確かめてから実行する(変更があれば先に WIP コミット(推奨)か名前付き stash(`git stash push -u -m ""`。stash スタックは全 worktree 共有なので `git stash pop` ではなく `git stash apply ` で戻す)へ退避する)。 - - コミット前に下記の品質チェックを通す(`CONTRIBUTING.md` ルール、手順 3 で定義する「コード本体パス」に変更が無ければ省略可): + - コミット前に `python3 .claude/skills/create-pr/prepare.py --worktree` を実行し、origin の base から作業ツリーまでの変更(未コミット・未追跡・未 push を含む)を判定する。`code_changed` が true なら下記の品質チェックを通す(`CONTRIBUTING.md` ルール): - `cargo fmt --all -- --check` - `make clippy` - `make test` - - データのみの変更(`data/*.csv` 等)を含む場合は `cargo run -p data_validator` も流す。 + - `data_changed` が true なら `cargo run -p data_validator` も流す。 - push は新規ブランチなので安全だが、実行前にユーザーへ要約(ブランチ名・含めるファイル・コミットメッセージ案)を提示して承認を取る。 以降の手順では推論後の head を使う。 2. **状態確認とモード決定(新規作成 / 更新)** - - `git fetch` で `BASE_REF` / `HEAD_REF` の remote-tracking ref を更新する。**`BASE_REV` / `HEAD_REV` の解決は必ず fetch の後に行う。** 先に解決すると古い commit ID で差分を測ることになる。前提条件で済ませておくのは ref 名の文字種検証までにとどめる。 - - **fetch 対象は refspec で明示する。** ブランチ名だけを渡す形(`git fetch origin `)は remote-tracking ref の更新が `remote.origin.fetch` の設定に依存する。別マシンで refspec が絞られていると `refs/remotes/origin/` が古いまま後続の解決と一致検査を通ってしまう。 - - **`BASE_REV` / `HEAD_REV` は origin 側の位置なので、ローカルの `HEAD_REF` がそれと一致することを機械的に確かめる。** 一致しなければ未 push のコミットがあり、そのまま進むとその分を含まない範囲で PR が組み上がる。検出したら push の可否をユーザーに確認して中断する(勝手に push しない)。 - - **ローカルに `HEAD_REF` が無い場合も中断する。** 空は「未 push が無い」証拠ではなく、単に検証できていない状態(ローカルで削除済み、あるいは origin にしか無いブランチを `head` に指定した、など)。`git switch --track "origin/<名前>"` でローカルへ取り込んでからやり直す。 - - コミットとファイル差分の**両方**を確認する。`git log` はコミットの有無しか見ないため、空コミットだけが載ったブランチが通過してしまう。 - - ```bash - # remote.origin.fetch の設定に左右されないよう refspec で明示する - git fetch origin \ - "+refs/heads/$BASE_REF:refs/remotes/origin/$BASE_REF" \ - "+refs/heads/$HEAD_REF:refs/remotes/origin/$HEAD_REF" - BASE_REV="$(resolve_remote_rev "$BASE_REF")" || exit 1 # 前提条件で定義したヘルパ。解決はここが最初 - HEAD_REV="$(resolve_remote_rev "$HEAD_REF")" || exit 1 - - # ローカルの HEAD_REF を解決し、origin の先端と一致することを確認する。 - # HEAD_REF は validate_ref で検証済み - HEAD_LOCAL_REV="$(git rev-parse --verify --quiet "refs/heads/$HEAD_REF")" - # 該当が無いと空を返すので、中身で判定する - if [ -z "$HEAD_LOCAL_REV" ]; then - printf 'ローカルに %s が無い。origin だけを見て組むと一致を検証できない。\n' "$HEAD_REF" >&2 - printf 'git switch --track "origin/%s" で取り込んでからやり直す。\n' "$HEAD_REF" >&2 - exit 1 - fi - if [ "$HEAD_LOCAL_REV" != "$HEAD_REV" ]; then - printf 'ローカル %s が origin と一致しない:\n local = %s\n origin = %s\n' \ - "$HEAD_REF" "$HEAD_LOCAL_REV" "$HEAD_REV" >&2 - exit 1 # 未 push の変更がある。push の可否をユーザーに確認してから進む - fi - - git log --oneline "$BASE_REV..$HEAD_REV" - git diff --name-only "$BASE_REV" "$HEAD_REV" - ``` - - コミット一覧が空、または `git diff --name-only` の出力が空の場合は「PR 対象の差分が無い」と報告し、**既存 PR の検索へ進まずに中断する**。 - - `gh pr list --base "$BASE_REF" --head "$HEAD_REF" --state open --json number,url,body` で既存 open PR を確認。 - - **存在しない場合**: 新規作成モード。以降、手順 5 で `gh pr create`。 - - **存在する場合**: 更新モード。既存本文を最新差分で再生成する。以降、手順 5 で `gh pr edit`。タイトルは既存を**原則尊重**(ユーザー推論より優先)。ただし手順 5 の整合性チェックで主題が大きくズレていると判断した場合のみ更新案を提示する。 -3. **変更の種類を判定** - - `origin/..origin/` のコミット件名と変更ファイルを取得: ```bash - git log --pretty=%s "$BASE_REV..$HEAD_REV" - git diff --name-only "$BASE_REV" "$HEAD_REV" + python3 .claude/skills/create-pr/prepare.py # --base / --head で上書きできる ``` - **大原則: 判定はアプリ挙動/データに対する変更かどうかで決める**。下の「コード本体パス」が一切変わっていない場合、「バグ修正」「新機能」「リファクタリング」は OFF(コミット件名に `fix` / `feat` 等の語があっても)。スキル・設定・ドキュメントのメタ変更を「新機能」と誤分類しないための安全弁。「データの修正・追加」は `data/**` の変更を独立に判定する(後述「変更ファイルパスベース」「コミット件名ベース」を参照)。 - - この大原則のもとで、各項目を独立に評価(複数該当可、大文字小文字無視・部分一致)。 - - **コード本体パス**(バグ修正 / 新機能 / リファクタリングのゲート) - - - `stationapi/src/**` - - `stationapi/proto/**` - - `data_validator/src/**` - - `tools/**` - - `docker/**` - - `Cargo.toml` / `Cargo.lock` - - `wrangler.jsonc` + `prepare.py` は次の順に確かめる。途中で止まったら理由を標準エラーに出し、終了コード 1 で終わる。その場合は理由をユーザーに伝え、指示に従う。 + - base / head の ref 名が `^[A-Za-z0-9._/-]+$` に一致するか(先頭の `-` は不可)。head が `HEAD`(detached)や base と同じなら、手順 1 で切り出す。 + - 両ブランチを refspec で明示して fetch し、`refs/remotes/origin/<名前>` の完全形で解決する。refspec を省くと remote-tracking ref の更新が `remote.origin.fetch` の設定次第になり、短縮名だと同名のタグやローカルブランチが優先される。 + - ローカルの head が origin と一致するか。一致しない、あるいはローカルに無い場合は、未 push のコミットを検証できないので止まる。 + - コミットとファイル差分がどちらも空でないか(空コミットだけのブランチを通さない)。 + - 既存の open PR を探し、「変更の種類」を判定する(手順 3)。 - **コード本体変更ありの場合 — コミット件名ベース** + 出力は JSON で、以降の手順は `base` / `head` / `commits` / `files` / `code_changed` / `data_changed` / `types` / `checklist` / `existing_pr` を使う。手順 5 のシェルでは `base` と `head` の値を `BASE_REF` / `HEAD_REF` に代入し、`"${BASE_REF}"` のように中括弧付きで引用する。 + - `existing_pr` が null: 新規作成モード。手順 5 で `gh pr create`。 + - `existing_pr` がある: 更新モード。既存本文(`existing_pr.body`)を最新差分で再生成し、手順 5 で `gh pr edit`。タイトルは既存を**原則尊重**(ユーザー推論より優先)。ただし手順 5 の整合性チェックで主題が大きくズレていると判断した場合のみ更新案を提示する。 - | 項目 | トリガ語句 | - | ---- | ---- | - | バグ修正 | `fix`, `Hotfix`, `バグ`, `修正`, `不具合` | - | 新機能 | `feat`, `add`, `新機能`, `追加`, `導入`, `対応`, `RPC` | - | リファクタリング | `refactor`, `リファクタ`, `整理`, `clean`, `tidy` | +3. **変更の種類を判定** - **変更ファイルパスベース**(コード本体変更の有無に関わらず評価) + 判定は `prepare.py` の `classify()` が行い、`types`(項目ごとの ON / OFF と根拠)と `checklist`(テンプレ順のチェック欄)を返す。規則を変えるときは `prepare.py` を直して `--self-test` を回す。この文書に判定表を書き戻さない(二か所に置くと食い違う)。 - | 項目 | パターン | - | ---- | ---- | - | データの修正・追加 | `data/**/*.csv` | - | ドキュメント | 変更が `*.md` / `docs/**` / `README*` / `.claude/**` / `AGENTS.md` / `CONTRIBUTING.md` のみ、またはそれらを主体とする | - | CI/CD | `.github/workflows/**`, `.github/**/*.yml`, `Makefile` のいずれかを含む | + 規則の骨子: + - **判定はアプリの挙動やデータに対する変更かどうかで決める。** `CODE_PATHS`(Worker 本体の `src/`、`stationapi/src/`、`preprocessor/src/` など)に変更が無ければ、コミット件名に fix / feat とあってもバグ修正・新機能・リファクタリングは OFF にする。スキルやドキュメントの手入れを「新機能」と誤分類しないため。 + - バグ修正・新機能・リファクタリングはコミット件名のトリガ語句で決める。英字の語句は単語として一致したときだけ数えるので、`line_cd` の `cd` や `data_validator` の `data` には当たらない。 + - データの修正・追加は `data/**/*.csv` の変更で決める(`data/README.md` だけならドキュメント)。 + - ドキュメントは、コードも CSV も含まず、変更がドキュメントだけか、コミット件名がドキュメントを示すときに ON。CI/CD は `.github/` のワークフローと action、`Makefile` の変更、またはコミット件名で決める。 + - どれも OFF のときだけ「その他」を ON にする。 - **コミット件名ベース(データ・ドキュメント・CI/CD)** - - | 項目 | トリガ語句 | - | ---- | ---- | - | データの修正・追加 | `データ`, `data`, `駅`, `路線`, `numbering`, `CSV`, `csv` | - | ドキュメント | `docs`, `ドキュメント`, `README`, `changelog`, `AGENTS`, `CONTRIBUTING` | - | CI/CD | `ci`, `cd`, `workflow`, `release`, `Bump version`, `labeler` | - - 判定ロジック: - - 上の「大原則」のゲートをまず適用。コード本体/データの変更が無ければバグ修正・新機能・リファクタリング・データの修正・追加は強制 OFF。 - - 「データの修正・追加」は `data/**` の変更があれば ON。`data/README.md` のみの変更なら「ドキュメント」のみ ON にする。 - - 「ドキュメント」は変更にコード本体や CSV を含まない場合に ON。混在する場合は基本 OFF(主目的が分かるならそちらを優先)。ただし `.claude/**` や `AGENTS.md` のみの変更は「ドキュメント」を ON にする(運用ドキュメント扱い)。 - - 「CI/CD」は `.github/workflows/**` 等の変更があれば独立に ON。 - - 残りの項目は、コミット件名またはファイルパスのトリガに 1 つでも当てはまれば `- [x]`、それ以外は `- [ ]`。 - - 全項目が OFF のときのみ `その他` を `- [x]` にする。他項目が ON のときは `その他` は必ず `- [ ]`。 + 判定が PR の実態と合わないと思ったら、手で書き換えずに根拠(`types[].reasons`)をユーザーに示して確認する。 4. **本文組み立て** @@ -218,10 +120,10 @@ description: Create a GitHub pull request for TrainLCD StationAPI that conforms **新規作成モード** - 「概要」節: `summary` があれば挿入。無ければテンプレのコメントだけ残す。 - - 「変更の種類」節: 手順 3 の結果で各 `- [ ]` / `- [x]` を決定。**項目順序は必ずテンプレ通り**(バグ修正 / 新機能 / データの修正・追加 / リファクタリング / ドキュメント / CI/CD / その他)。 + - 「変更の種類」節: `checklist` をそのまま使う(テンプレ通りの順序で並んでいる)。 - 「変更内容」節: コミット件名と変更ファイルから短い箇条書きを生成。`summary` があればそれを優先。データのみの PR では追加・修正した路線・駅などを箇条書きで列挙すると親切。 - 「テスト」節: - - **判定基準: 手順 3 の「コード本体パス」(`stationapi/src/**` ほか)に変更が無い場合は Step 1 の `cargo` チェックを省略したとみなし、3 項目すべて OFF**(`skip_checks` より優先)。本文末尾に「省略: コード変更なし」等の短い注記を残す。 + - **判定基準: `code_changed` が false なら Step 1 の `cargo` チェックを省略したとみなし、3 項目すべて OFF**(`skip_checks` より優先)。本文末尾に「省略: コード変更なし」等の短い注記を残す。 - 上記に該当しない場合は `skip_checks` が真なら 3 項目すべて OFF、偽なら 3 項目すべて ON。テキストはテンプレのまま(`make fmt` / `make clippy` / `make test`)。 - 「関連Issue」節: `related_issue` が指定されていればユーザー入力を最優先で出力(`#N` のみなら `Closes #N`、`Closes/Fixes/Refs #N` 形式なら接頭語を維持)。空のときに限りコミット件名から `Closes/Fixes/Refs #N` を抽出。どちらも無ければコメントのみ。 - 「スクリーンショット」節: 常にコメントのみ(API レスポンスの diff など必要なら呼び出し側が後から編集する前提)。 @@ -235,7 +137,7 @@ description: Create a GitHub pull request for TrainLCD StationAPI that conforms | 概要 | 既存内容を尊重。空欄(テンプレのコメントのみ)なら新規作成モードと同じ生成を試みる。 | | 変更の種類 | **常に手順 3 の結果で上書き**(機械的判定)。 | | 変更内容 | 冒頭の箇条書きブロック(`-` で始まる連続行)を最新差分で再生成。その下に人間が書いた散文があれば残す。 | - | テスト | **常に `skip_checks` に従う**(手順 4 の本文組み立てと同じルール)。 | + | テスト | **新規作成モードと同じルールで上書き**(`code_changed` が false なら `skip_checks` に関わらず 3 項目すべて OFF にして省略の注記を残し、それ以外は `skip_checks` に従う)。 | | 関連Issue | 既存内容を尊重。コミット件名に `Closes/Fixes/Refs #N` があり、かつ既存本文中に同じ Issue 番号 `#N` を指す表現が存在しない場合のみ追記(重複は作らない。比較時は `Closes` / `closes` / `Fixes` / `fixes` / `Refs` / `refs` を同一視し、空白・記号差は無視して `#N` 単位で照合)。 | | スクリーンショット | 既存内容を尊重。自動では触らない。 | @@ -260,7 +162,7 @@ description: Create a GitHub pull request for TrainLCD StationAPI that conforms ```bash # ref 名(ブランチ名)をファイル名として安全な集合(A-Za-z0-9._-)にスラッグ化 - REF_SLUG="$(printf '%s' "$HEAD_REF" \ + REF_SLUG="$(printf '%s' "${HEAD_REF}" \ | tr -d '\r\n' \ | tr -c 'A-Za-z0-9._-' '_' \ | sed -E 's/_+/_/g; s/^_+//; s/_+$//' \ @@ -270,8 +172,8 @@ description: Create a GitHub pull request for TrainLCD StationAPI that conforms ( trap 'rm -f "$BODY_FILE"' EXIT INT TERM gh pr create \ - --base "$BASE_REF" \ - --head "$HEAD_REF" \ + --base "${BASE_REF}" \ + --head "${HEAD_REF}" \ --title "" \ --assignee TinyKitten \ [--label "<label1>" --label "<label2>" ...] \ diff --git a/.claude/skills/create-pr/prepare.py b/.claude/skills/create-pr/prepare.py new file mode 100755 index 00000000..f6f09a1f --- /dev/null +++ b/.claude/skills/create-pr/prepare.py @@ -0,0 +1,329 @@ +#!/usr/bin/env python3 +"""create-pr スキルの下ごしらえ。ref の検証と解決、差分の取得、「変更の種類」の判定をする。 + + python3 .claude/skills/create-pr/prepare.py # base = 既定ブランチ, head = カレント + python3 .claude/skills/create-pr/prepare.py --base dev --head feature/foo + python3 .claude/skills/create-pr/prepare.py --worktree # 未 push の作業 (手順 1) を判定する + python3 .claude/skills/create-pr/prepare.py --self-test # 判定規則の自己診断 (git も gh も呼ばない) + +結果は JSON で標準出力へ、中断の理由は標準エラーへ出す。中断したときの終了コードは 1。 + +git と gh は引数リストで直接起動し、シェルを通さない。ref 名がシェルに解釈されることも、 +zsh が `$BASE_REF:r` のような修飾子を展開することも無い。依存は Python 3 標準ライブラリのみ。 +""" + +from __future__ import annotations + +import argparse +import json +import re +import subprocess +import sys +from fnmatch import fnmatchcase + +# ref 名として受け付ける集合。gh --head にもファイル名 slug にもそのまま使える文字だけに絞る。 +# 先頭の `-` はオプションと取り違えられるので弾く。 +REF_PATTERN = re.compile(r"^(?!-)[A-Za-z0-9._/-]+$") + +# アプリの挙動を変えるパス。ここに当たる変更が無ければ、コミット件名に fix / feat と +# 書いてあってもバグ修正・新機能・リファクタリングにはしない (スキルやドキュメントの +# 手入れを「新機能」と誤分類しないため)。cargo のチェックを走らせるかもこれで決まる。 +CODE_PATHS = [ + "src/*", + "stationapi/src/*", + "preprocessor/src/*", + "data_validator/src/*", + "tools/*", + "build.rs", + "schema/*", + "Cargo.toml", + "*/Cargo.toml", + "Cargo.lock", + "wrangler.jsonc", +] +DATA_PATHS = ["data/*.csv"] +DOC_PATHS = ["*.md", "docs/*", "README*", ".claude/*"] +CI_PATHS = [".github/workflows/*", ".github/actions/*", ".github/*.yml", ".github/*.yaml", "Makefile"] + +# コミット件名のトリガ語句。英字は単語として一致したときだけ数える (`cd` が `line_cd` に、 +# `data` が `data_validator` に当たらないように)。語尾の s / es / ed / d / ing は許す。 +# 日本語は部分一致。 +CODE_SUBJECT_TRIGGERS = { + "バグ修正": ["fix", "hotfix", "バグ", "修正", "不具合"], + "新機能": ["feat", "add", "新機能", "追加", "導入", "対応"], + "リファクタリング": ["refactor", "リファクタ", "整理", "clean", "tidy"], +} +DOC_SUBJECT_TRIGGERS = ["docs", "ドキュメント", "readme", "changelog", "agents", "contributing"] +CI_SUBJECT_TRIGGERS = ["ci", "cd", "workflow", "release", "bump version", "labeler"] + +# .github/pull_request_template.md の並び順。 +LABELS = ["バグ修正", "新機能", "データの修正・追加", "リファクタリング", "ドキュメント", "CI/CD", "その他"] + + +class Abort(Exception): + """ユーザーに確認してから進めるべき状態。メッセージをそのまま伝える。""" + + +# --------------------------------------------------------------------------- 判定 + + +def matches(path: str, patterns: list[str]) -> bool: + # fnmatch の `*` は `/` もまたぐので、`src/*` は src 以下すべてに当たる。 + return any(fnmatchcase(path, p) for p in patterns) + + +def find_trigger(subject: str, triggers: list[str]) -> str | None: + for token in triggers: + if token.isascii(): + pattern = rf"(?<![A-Za-z0-9_]){re.escape(token)}(?:s|es|ed|d|ing)?(?![A-Za-z0-9_])" + if re.search(pattern, subject, re.IGNORECASE): + return token + elif token in subject: + return token + return None + + +def subject_hits(subjects: list[str], triggers: list[str]) -> list[str]: + hits = [] + for subject in subjects: + token = find_trigger(subject, triggers) + if token: + hits.append(f"コミット「{subject}」の `{token}`") + return hits + + +def classify(subjects: list[str], files: list[str]) -> dict: + code_files = [f for f in files if matches(f, CODE_PATHS)] + data_files = [f for f in files if matches(f, DATA_PATHS)] + doc_files = [f for f in files if matches(f, DOC_PATHS)] + ci_files = [f for f in files if matches(f, CI_PATHS)] + + reasons: dict[str, list[str]] = {label: [] for label in LABELS} + + if code_files: + for label, triggers in CODE_SUBJECT_TRIGGERS.items(): + reasons[label] = subject_hits(subjects, triggers) + + reasons["データの修正・追加"] = [f"`{f}` の変更" for f in data_files] + + # コードや CSV と混ざった PR は、その主目的の欄で表す。 + if not code_files and not data_files and doc_files: + if len(doc_files) == len(files): + reasons["ドキュメント"] = ["変更がドキュメントだけ"] + else: + reasons["ドキュメント"] = subject_hits(subjects, DOC_SUBJECT_TRIGGERS) + + reasons["CI/CD"] = [f"`{f}` の変更" for f in ci_files] + subject_hits(subjects, CI_SUBJECT_TRIGGERS) + + if not any(reasons[label] for label in LABELS[:-1]): + reasons["その他"] = ["どの項目にも当たらない"] + + types = [{"label": label, "checked": bool(reasons[label]), "reasons": reasons[label]} for label in LABELS] + return { + "code_changed": bool(code_files), + "data_changed": bool(data_files), + "types": types, + "checklist": "\n".join(f"- [{'x' if t['checked'] else ' '}] {t['label']}" for t in types), + } + + +# --------------------------------------------------------------------------- git / gh + + +# git fetch が認証待ちで止まるなどしても、Abort として理由を返して終わらせる。 +TIMEOUT_SEC = 300 + + +def spawn(argv: tuple[str, ...]) -> subprocess.CompletedProcess: + try: + return subprocess.run(argv, capture_output=True, text=True, timeout=TIMEOUT_SEC) + except subprocess.TimeoutExpired: + raise Abort(f"`{' '.join(argv)}` が {TIMEOUT_SEC} 秒以内に終わらなかった") from None + except OSError as e: + raise Abort(f"`{argv[0]}` を起動できない: {e}") from None + + +def run(*argv: str) -> str: + result = spawn(argv) + if result.returncode != 0: + raise Abort(f"`{' '.join(argv)}` が失敗した:\n{result.stderr.strip()}") + return result.stdout + + +def lines(text: str) -> list[str]: + return [line for line in text.splitlines() if line] + + +def validate_ref(name: str, role: str) -> None: + if not REF_PATTERN.match(name): + raise Abort(f"{role} の ref 名に想定外の文字が含まれる: {name!r}。正しいブランチ名をユーザーに確認する") + + +def rev(ref: str) -> str | None: + # 短縮名だと同名のタグやローカルブランチが優先されるので、完全形で解決する。 + return spawn(("git", "rev-parse", "--verify", "--quiet", ref)).stdout.strip() or None + + +def default_base() -> str: + return run("gh", "repo", "view", "--json", "defaultBranchRef", "-q", ".defaultBranchRef.name").strip() + + +def fetch(*branches: str) -> None: + # ブランチ名だけを渡すと remote-tracking ref の更新が remote.origin.fetch の設定に + # 左右されるので、refspec で明示する。 + run("git", "fetch", "origin", *(f"+refs/heads/{b}:refs/remotes/origin/{b}" for b in branches)) + + +def prepare_range(base: str, head: str) -> dict: + validate_ref(base, "base") + if head == "HEAD" or head == base: + raise Abort(f"PR の head にできるブランチに居ない ({head})。手順 1 でブランチを切り出す") + validate_ref(head, "head") + + # 解決は必ず fetch の後。先に解決すると古い commit で差分を測る + fetch(base) + try: + fetch(head) + except Abort: + raise Abort(f"origin に {head} が無い。push の可否をユーザーに確認する") from None + base_rev = rev(f"refs/remotes/origin/{base}") + head_rev = rev(f"refs/remotes/origin/{head}") + if not base_rev or not head_rev: + raise Abort(f"fetch したはずの origin/{base} か origin/{head} を解決できない") + + local_rev = rev(f"refs/heads/{head}") + if not local_rev: + raise Abort(f"ローカルに {head} が無いので未 push のコミットを検証できない。" + f"`git switch --track origin/{head}` で取り込んでからやり直す") + if local_rev != head_rev: + raise Abort(f"ローカルの {head} が origin と一致しない (local {local_rev} / origin {head_rev})。" + "未 push のコミットがあるので、push の可否をユーザーに確認する") + + subjects = lines(run("git", "log", "--pretty=%s", f"{base_rev}..{head_rev}")) + # GitHub の PR 差分と同じく merge-base から測る。先端同士を比べると、head を切った後に + # base へ入ったコミットの変更まで混ざる。 + files = lines(run("git", "diff", "--name-only", f"{base_rev}...{head_rev}")) + # コミットだけを見ると空コミットのブランチが通るので、ファイル差分も確かめる。 + if not subjects or not files: + raise Abort("PR 対象の差分が無い") + + prs = json.loads(run("gh", "pr", "list", "--base", base, "--head", head, "--state", "open", + "--json", "number,url,body")) + return { + "base": base, "head": head, "base_rev": base_rev, "head_rev": head_rev, + "commits": subjects, "files": files, + "existing_pr": prs[0] if prs else None, + **classify(subjects, files), + } + + +def prepare_worktree(base: str) -> dict: + """手順 1 用。origin の base から作業ツリーまでの変更 (未コミット・未追跡・未 push を含む) を判定する。""" + validate_ref(base, "base") + fetch(base) + base_ref = f"refs/remotes/origin/{base}" + merge_base = run("git", "merge-base", base_ref, "HEAD").strip() + files = sorted(set(lines(run("git", "diff", "--name-only", merge_base))) + | set(lines(run("git", "ls-files", "--others", "--exclude-standard")))) + subjects = lines(run("git", "log", "--pretty=%s", f"{base_ref}..HEAD")) + if not files: + raise Abort("PR 対象の差分が無い") + return { + "base": base, "branch": run("git", "rev-parse", "--abbrev-ref", "HEAD").strip(), + "commits": subjects, "files": files, + **classify(subjects, files), + } + + +# --------------------------------------------------------------------------- 自己診断 + +# (説明, コミット件名, 変更ファイル, ON になるべき項目, code_changed) +_CASES = [ + ("Worker 本体だけの変更はコード変更として数える", + ["近傍バス停の絞り込みを件数の上限より先に行う"], + ["src/index.rs", "src/repository.rs", "docs/nearby-bus-stops.md"], {"その他"}, True), + ("識別子の一部 (line_cd / data_validator) には CI/CD もデータも当たらない", + ["feat(data_validator): line_cd が 2!lines.csv に存在するか検証する"], + ["data_validator/src/main.rs"], {"新機能"}, True), + ("ドキュメントだけなら fix と書いてあってもバグ修正にしない", + ["fix: スキルの誤記を直す"], [".claude/skills/create-pr/SKILL.md", "AGENTS.md"], {"ドキュメント"}, False), + ("gRPC の中の RPC やカタカナには当たらない", + ["エージェント向けガイドとスキルに残っていたgRPC時代の記述と誤った参照を直した"], + ["AGENTS.md", "CONTRIBUTING.md"], {"ドキュメント"}, False), + ("CSV の変更はデータ", + ["都営大江戸線 都庁前のstation_g_cd統一とe_sortの連番化"], ["data/3!stations.csv"], + {"データの修正・追加"}, False), + ("data/README.md だけならドキュメント", + ["data/README.md の表を更新"], ["data/README.md"], {"ドキュメント"}, False), + ("日本語に続く CI も語として数える", + ["CIでWorkerのユニットテストを実行する"], [".github/workflows/ci.yml"], {"CI/CD"}, False), + ("composite action の変更は CI/CD", + ["build-worker の既定値を揃える"], [".github/actions/build-worker/action.yml"], {"CI/CD"}, False), + ("composite action のスクリプトだけの変更も CI/CD", + ["build-worker の手順を直す"], [".github/actions/build-worker/entrypoint.sh"], {"CI/CD"}, False), + ("コードとドキュメントが混ざればドキュメントは付けない", + ["駅番号の照合漏れを修正", "README を更新"], + ["stationapi/src/use_case/interactor/query.rs", "README.md"], {"バグ修正"}, True), + ("語尾の活用は数える", + ["Added sortBy to connectedRoutes"], ["src/graphql/enums.rs"], {"新機能"}, True), + ("依存更新はコード変更だが種類はその他", + ["依存を上げる"], ["Cargo.lock", "preprocessor/Cargo.toml"], {"その他"}, True), +] + +_REF_CASES = [("feature/foo-bar", True), ("release/v1.2.0", True), ("a;b", False), + ("$(x)", False), ("-x", False), ("", False), ("日本語", False)] + + +def self_test() -> int: + failures = 0 + for label, subjects, files, expected, code_changed in _CASES: + result = classify(subjects, files) + got = {t["label"] for t in result["types"] if t["checked"]} + ok = got == expected and result["code_changed"] == code_changed + failures += not ok + print(f" {'ok ' if ok else 'FAIL'} {label}" + + ("" if ok else f"\n 期待 {sorted(expected)} code={code_changed}" + f" / 実際 {sorted(got)} code={result['code_changed']}"), + file=sys.stderr) + for name, valid in _REF_CASES: + ok = bool(REF_PATTERN.match(name)) == valid + failures += not ok + print(f" {'ok ' if ok else 'FAIL'} ref {name!r} を{'受け付ける' if valid else '弾く'}", file=sys.stderr) + total = len(_CASES) + len(_REF_CASES) + print(f"全 {total} 件中 {total - failures} 件 ok", file=sys.stderr) + return 1 if failures else 0 + + +# --------------------------------------------------------------------------- main + + +def main() -> int: + parser = argparse.ArgumentParser(description=__doc__, + formatter_class=argparse.RawDescriptionHelpFormatter) + parser.add_argument("--base", help="PR の base (既定はリポジトリの既定ブランチ)") + parser.add_argument("--head", help="PR の head (既定はカレントブランチ)") + parser.add_argument("--worktree", action="store_true", + help="origin の base から作業ツリーまでの変更を判定する (手順 1)") + parser.add_argument("--self-test", action="store_true", help="判定規則の自己診断だけ走らせる") + args = parser.parse_args() + + if args.self_test: + return self_test() + + try: + base = args.base or default_base() + if args.worktree: + result = prepare_worktree(base) + else: + head = args.head or run("git", "rev-parse", "--abbrev-ref", "HEAD").strip() + result = prepare_range(base, head) + except Abort as e: + print(e, file=sys.stderr) + return 1 + + print(json.dumps(result, ensure_ascii=False, indent=2)) + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/AGENTS.md b/AGENTS.md index d6316bc7..ab85947d 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -9,7 +9,7 @@ This guide explains how automation agents and human contributors should work wit - `wrangler.jsonc` – Staging and production deployment settings. - `stationapi/src/domain/` – Entity definitions and repository abstractions. `repository/` provides `async_trait`-based interfaces, and `normalize.rs` contains text normalization for search. - `stationapi/src/use_case/` – Application logic. `interactor/query.rs` implements the `QueryUseCase` contract defined in `traits/query.rs`; `dto/` converts entities into `model` types (this is where IPA and TTS segments are built). -- `stationapi/src/model.rs` – The values the API returns. Formerly generated from `.proto`; kept as the layer between entities and GraphQL types. +- `stationapi/src/model.rs` – The values the API returns; the layer between entities and GraphQL types. - `preprocessor/` – Build-time CLI that assembles `generated/*.csv` from `data/*.csv`, the GTFS feeds, and the Tokyu ODPT JSON. - `data/` – Canonical CSV datasets. Files follow the `N!table.csv` naming scheme. Detailed instructions are in `data/README.md`. - `data_validator/` – CLI that verifies cross-file constraints (`cargo run -p data_validator`). @@ -65,7 +65,7 @@ The Worker is the workspace root package. `stationapi`, `preprocessor`, and `dat - **Endpoint benchmarks** – `make bench` (or `python3 .claude/skills/benchmark-gql/bench.py`) replays every `Query` field against production (`gql.trainlcd.app`, script `stationapi`) and staging (`gql-stg.trainlcd.app`, script `stationapi-stg`) and writes a Markdown report under `benchmarks/`. Both environments embed the same data, so any difference is implementation — which makes this the way to see what a `dev`-to-`master` release will do to performance before it ships. Besides client latency it records the Worker's `cpuTime`, read from `wrangler tail --format json` and matched to each request by `cf-ray`; the tail is filtered on a per-run request header, so production's live traffic does not leak into the sample. Collecting CPU time needs the `workers_tail (read)` scope, and the run sends hundreds of real requests to production — it is not a routine check. Add a case to `.claude/skills/benchmark-gql/queries.json` whenever a `Query` field is added, and never edit an existing case's variables: the reports are meant to stay comparable across runs. ## GraphQL Query Overview -- **Stations** – `station`, `stations`, `stationGroupStations`, `stationsNearby`, `lineStations`, `stationsByName`, `lineGroupStations`, `lineListStations`, `lineGroupListStations`. `QueryInteractor` enriches stations with lines, companies, station numbers, and train types. `lineStations` resolves the line's local train-type group (rail `kind` 0/1 or a `priority > 0` type); when no such group exists — bus lines only carry `BusRoute` (`kind` 7, `priority` 0) variants — it falls back to the line's plain typeless station list so bus stop listings never return empty. `stationsByName` with `fromStationGroupId` returns the stations reachable from there: stations sharing a line group with the origin (`line_group_cd` set, `has_train_types` true), same-line stations when either side has no line group, and — for rail — stations reachable by transferring, i.e. those for which `connectedRoutes` with `viaLineId` set to the station's line returns a route (`line_group_cd` empty, `has_train_types` false). The transfer check uses `RouteTopology` (`stationapi/src/domain/route_topology.rs`), a time-free copy of the `connectedRoutes` network built straight from the index without `Station` entities or time estimates (about 20 ms instead of about 190 ms, cached in its own `OnceLock`); `RouteNetwork` holds the same topology, both share `trim_pattern` and `line_group_rows`, and a real-data test asserts the two are equal. The check is a ride-limited BFS over line groups (a few ms) plus, only for destinations that are cut vertices of the station–line-group graph, a check that arrives without stopping over at the destination group — otherwise a branch's junction station (Ishibashi-handai-mae on the Minoo Line) would be listed although reaching it on that branch means riding out and back. +- **Stations** – `station`, `stations`, `stationGroupStations`, `stationsNearby`, `lineStations`, `stationsByName`, `lineGroupStations`, `lineListStations`, `lineGroupListStations`. `QueryInteractor` enriches stations with lines, companies, station numbers, and train types. `lineStations` resolves the line's local train-type group (rail `kind` 0/1 or a `priority > 0` type; a bullet-train line, which has no local service, takes the group of its stopping train type with the smallest `types.id`, e.g. Nozomi or Hayabusa) and returns its stops carrying that group's train type, with or without `stationId`, so a client that never picked a train type still has the `lineGroupId` `trainRoute` requires; when no such group exists — bus lines only carry `BusRoute` (`kind` 7, `priority` 0) variants — it falls back to the line's plain typeless station list so bus stop listings never return empty. `stationsByName` with `fromStationGroupId` returns the stations reachable from there: stations sharing a line group with the origin (`line_group_cd` set, `has_train_types` true), same-line stations when either side has no line group, and — for rail — stations reachable by transferring, i.e. those for which `connectedRoutes` with `viaLineId` set to the station's line returns a route (`line_group_cd` empty, `has_train_types` false). The transfer check uses `RouteTopology` (`stationapi/src/domain/route_topology.rs`), a time-free copy of the `connectedRoutes` network built straight from the index without `Station` entities or time estimates (about 20 ms instead of about 190 ms, cached in its own `OnceLock`); `RouteNetwork` holds the same topology, both share `trim_pattern` and `line_group_rows`, and a real-data test asserts the two are equal. The check is a ride-limited BFS over line groups (a few ms) plus, only for destinations that are cut vertices of the station–line-group graph, a check that arrives without stopping over at the destination group — otherwise a branch's junction station (Ishibashi-handai-mae on the Minoo Line) would be listed although reaching it on that branch means riding out and back. - **Lines** – `line`, `lines`, `linesByName`. Results include company data and computed line symbols based on repository helpers. - **Routes** – `routes`, `connectedRoutes`, `estimateArrivalTimes`, `trainRoute`. Paging tokens are currently empty (pagination not implemented). - **`trainRoute`** – Takes the line group's stops from the repository *before* any enrichment, slices them to the requested `fromStationId`–`toStationId` range (reversing when the request runs backwards), and only then attaches lines, companies, station numbers, train types, and nearby bus routes. Enrichment is per-station and independent, so slicing first does not change any segment; enriching the whole line group first made a three-station request cost the same as a 250-station one. Keep the order — the cost of this query must stay proportional to the requested range, not to the line group. @@ -73,8 +73,8 @@ The Worker is the workspace root package. `stationapi`, `preprocessor`, and `dat - **Train types** – `stationTrainTypes`, `routeTypes`. Train types aggregate by line group and include related lines plus optional train type metadata. Rail variants use `TrainTypeKind::{Default, Branch, Rapid, Express, LimitedExpress, HighSpeedRapid, CommuterRapid}` (0-6); bus variants use `BusRoute` (7), which represents a `(route_id, shape_id)` operation pattern (e.g. 循環 / 短ターン / 支線) generated automatically from the configured GTFS bus feeds (Toei Bus, Seibu Bus, Keio Bus) and the converted Tokyu Bus JSON. - **Default rail train types** – `preprocessor` fills every active rail line containing at least one station with no `station_station_types` row with a deterministic, complete all-stop group. The generated rows exist only in `generated/*.csv`; canonical CSV files remain unchanged. `type_cd=100` represents 「普通」 and `type_cd=101` represents 「各駅停車」. An existing 100/101 assignment on the line takes precedence; otherwise the label is selected per line through `LOCAL_SERVICE_RAIL_LINE_IDS` in `preprocessor/src/rail.rs`. Generated `line_group_cd` values use `1,000,000,000 + line_cd`; generation fails on a collision. Bus lines are excluded and continue to use their GTFS-derived `BusRoute` groups. - **GTFS bus integration** – `preprocessor/src/gtfs/` reads the GTFS feeds into an in-memory representation and then projects them onto the shared `stations` / `lines` / `types` / `station_station_types` tables (`gtfs/integrate.rs`). Only routes, stops, trips, and stop_times are read; calendar, shapes, feed_info, and agencies do not affect the output. Every configured GTFS feed is imported, including Seibu Bus and Keio Bus (both downloaded from ODPT with `ODPT_ACCESS_TOKEN`). Tokyu Bus ordinary-route `BusroutePattern`, `BusstopPole`, and `BusTimetable` JSON are converted into the same representation; pattern IDs become `shape_id` values so route variants remain queryable as bus TrainTypes. The Tokyu-operated Ota, Shinagawa, and Meguro community buses use their official GTFS feeds and matching JSON routes are excluded to prevent duplicates. `ODPT_ACCESS_TOKEN` is required for authenticated sources; without it those feeds are skipped with a warning rather than failing the build. Stops whose Tokyu JSON records omit coordinates remain available to name and route queries but not coordinate searches. `transport_type` (0: rail, 1: bus) on both `stations` and `lines` keeps rail and bus records queryable side by side. GTFS IDs are namespaced per feed before import to avoid cross-operator collisions. `line_cd` (100,000,000+), `station_cd` / `station_g_cd` (200,000,000+), and bus `type_cd` / `line_group_cd` (100,000,000+) are all deterministic fnv1a hashes that stay clear of the rail data ranges. Disable the entire bus pipeline with `DISABLE_BUS_FEATURE=true`. -- **Bus stop translations (readings & English)** – GTFS-JP `translations.txt` layouts differ per feed, so `load_gtfs_translations` resolves columns by header name (Seibu ships 6 columns without `record_sub_id`; Keio and the Tokyu community feeds ship 7) and indexes each `stop_name` translation under both keys it may use: `record_id` (== the stop_id, Seibu — with the "-NN" pole suffix also mapped to the parent stop_id) and `field_value` (== the Japanese stop_name, Keio / Tokyu community, where `record_id` is left empty). `import_gtfs_stops` then looks a stop's translation up by stop_id first, then by name. Keying only by `record_id` (the previous behavior) silently dropped every field_value-keyed feed, leaving `station_name_k` filled with the kanji stop_name and `station_name_r` empty. Readings arriving as half-width katakana (`ニシハチオウジ`, Keio / Tokyu community) are folded to full-width via `romaji::to_fullwidth_katakana()` before storage. -- **Bus English-name fallback** – When a feed provides no English (`en`) translation for a stop — e.g. Tokyu Bus ordinary-route JSON, which carries only `dc:title` and `odpt:kana` — `src/domain/romaji.rs::romaji_display_name()` derives a modified-Hepburn romanization (with macrons for long vowels, matching the curated rail style: Tōkyō / Kyōto / Shin-Ōsaka) from the kana reading, and the GTFS reader fills `stop_name_r` with it. The fallback never overwrites a real `en` value, and a reading with no convertible kana stays `NULL` rather than emitting a partial transcription. Because `stop_name_r` is the single upstream source that fans out into the `stations` projection, `search_by_name`, and the romanized bus route/headsign names, this supplements every English-facing surface at once. When projecting into `stations`, `station_name_rn` is filled with the plain-ASCII spelling via `romaji::strip_macrons()` (Tōkyō → Tokyo), mirroring the rail dataset's `_r` (macron) / `_rn` (macron-free) column pair. +- **Bus stop translations (readings & English)** – GTFS-JP `translations.txt` layouts differ per feed, so `load_translations` (`preprocessor/src/gtfs/parse.rs`) resolves columns by header name (Seibu ships 6 columns without `record_sub_id`; Keio and the Tokyu community feeds ship 7) and indexes each `stop_name` translation under both keys it may use: `record_id` (== the stop_id, Seibu — with the "-NN" pole suffix also mapped to the parent stop_id) and `field_value` (== the Japanese stop_name, Keio / Tokyu community, where `record_id` is left empty). `load_stops` then looks a stop's translation up by stop_id first, then by name. Keying only by `record_id` would silently drop every field_value-keyed feed, leaving `station_name_k` filled with the kanji stop_name and `station_name_r` empty. Readings arriving as half-width katakana (`ニシハチオウジ`, Keio / Tokyu community) are folded to full-width via `romaji::to_fullwidth_katakana()` before storage. +- **Bus English-name fallback** – When a feed provides no English (`en`) translation for a stop — e.g. Tokyu Bus ordinary-route JSON, which carries only `dc:title` and `odpt:kana` — `stationapi/src/domain/romaji.rs::romaji_display_name()` derives a modified-Hepburn romanization (with macrons for long vowels, matching the curated rail style: Tōkyō / Kyōto / Shin-Ōsaka) from the kana reading, and the GTFS reader fills `stop_name_r` with it. The fallback never overwrites a real `en` value, and a reading with no convertible kana stays `NULL` rather than emitting a partial transcription. Because `stop_name_r` is the single upstream source that fans out into the `stations` projection, `search_by_name`, and the romanized bus route/headsign names, this supplements every English-facing surface at once. When projecting into `stations`, `station_name_rn` is filled with the plain-ASCII spelling via `romaji::strip_macrons()` (Tōkyō → Tokyo), mirroring the rail dataset's `_r` (macron) / `_rn` (macron-free) column pair. - **TTS metadata** – `Station`, `StationNested`, `Line`, `LineNested`, `TrainType`, and `TrainTypeNested` expose `name_ipa` / `name_roman_ipa` plus `name_tts_segments` for multi-segment pronunciation output. Use `name_tts_segments` when clients need per-token SSML construction for mixed-language names such as `Kasai-Rinkai Park`. - **Connected routes** – `connectedRoutes` finds transfer routes automatically, like a journey planner, using a frequency-based RAPTOR search in `stationapi/src/domain/route_search.rs`. Each rail line group is a pattern, station groups are the transfer nodes, and ride times come from `arrival_estimation`; bus lines are excluded. The cost adds a per-boarding wait by `TrainTypeKind` (limited express 15 min, express / high-speed rapid 5 min, others 3 min) and a 3-minute transfer walk — without the wait, infrequent limited expresses would beat the Yamanote Line. Rounds give the time/transfer Pareto set; alternatives come from re-searching with one leg's parallel line groups banned along that leg (at most 8 searches), and are dropped beyond 1.15 × best + 15 min or with two more transfers than the Pareto set. Alternative routes that stop at the same station group in two different legs (backtracking to re-board a banned train) are dropped; pass-through stations are not counted, and the Pareto routes of the first search are never dropped this way (otherwise a station `stationsByName` reports as reachable could get no route). Results are ranked by cost + 5 min per transfer and capped at 6. The time and transfer count stay internal (the API does not return them), so ordering happens on the server: `sortBy: ConnectedRouteSort` picks `Recommended` (the ranking above, the default when omitted), `ArrivalTime` (the estimated time `estimateArrivalTimes` also reports — excluding the first train's wait — then fewer transfers), or `TransferCount` (fewer transfers, then earlier arrival). `route_search::sort_journeys` only reorders the set `search` returned, with a stable sort so ties keep the recommended order; the set itself never depends on `sortBy`. The network (every rail line group plus its time estimates, about 190 ms natively) is built lazily into a `OnceLock` by `StationRepository::get_route_network` on the first `connectedRoutes` call, so other queries never pay for it. Each route is a list of `legs` shaped for the app's one-train-at-a-time flow: every leg carries its boarding and alighting `Station`, both on the line of the line group the search rode — so at a transfer the previous leg's alighting station and the next leg's boarding station may be different stations of one station group — `stationGroupIds`, the station groups from boarding to alighting in travel order including pass-through stations (the search's `JourneyLeg.station_group_ids` as-is — station groups rather than station IDs so the client can match them against whichever train type it picks, which may run on another line; a group appears twice on patterns such as the Oedo Line's Tochomae), and `trainTypes`, every train type usable on that leg (real `groupId`s, so the client picks one and calls `lineGroupStations`): it is exactly `routeTypes(boarding station group, alighting station group, alighting station's line)` (same dedup, same `lines`, same order — the use case calls `get_train_types`), because the search collapses parallel services such as local and rapid into one route and the app needs them to list types and default to the local. `viaLineId`, like `routeTypes`, is the line of the tapped search result and keeps only routes whose last leg arrives on that line. `estimateArrivalTimes` and `trainRoute` accept `legs: [RouteLegInput!]` (the `groupId` of the train type picked from each leg's `trainTypes`, plus the leg's `fromStation.id` and `toStation.id`) and then return values for the whole transfer route. A leg endpoint missing from the chosen line group is matched by station group (the picked local may stop at another line's station of the same group), preferring an exact `station_cd` and, among same-group candidates, the pair giving the shortest slice (through services list two stations of a junction group). ETA estimates each leg on its own line group only and chains them from the origin, adding the 3-minute walk and the next train type's wait at each transfer (the same allowance the search ranks by), returning one route with an empty `id`; `trainRoute` concatenates each leg's segments (each leg restarts at distance 0). Both slice legs with the same function, taking the shorter arc on loop lines, so their station sequences match. More than `MAX_RIDES` (6) legs — more than `connectedRoutes` ever returns — legs that do not connect, ends that differ from `fromStationId` / `toStationId`, or combining `legs` with `viaLineIds` / `directionId` / `lineGroupId` are errors. `docs/architecture.md` (乗換経路探索) has the details, and `docs/route-search.md` documents the search internals (data structures, the scan, pruning, alternatives, determinism). - Changes to the published contract require coordinated updates to `schema/public.graphql`, the async-graphql types in `src/graphql/`, and, when the shape of a value changes, `stationapi/src/model.rs` and the DTO conversions. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 8d7f7433..c35b686d 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -49,7 +49,7 @@ make dev # http://127.0.0.1:8787 | 種類 | プレフィックス | 例 | |------|------------|-----| -| 新機能 | `feature/` | `feature/add-new-rpc` | +| 新機能 | `feature/` | `feature/add-route-sort` | | バグ修正 | `fix/` | `fix/station-query-error` | | データ変更 | `data/` | `data/update-numbering` | | 雑務 | `chore/` | `chore/update-deps` | diff --git a/data/5!station_station_types.csv b/data/5!station_station_types.csv index e03614b8..399cfa3a 100644 --- a/data/5!station_station_types.csv +++ b/data/5!station_station_types.csv @@ -20299,12 +20299,33 @@ DEFAULT,1190803,102,535,1,九郎原 DEFAULT,1190804,102,535,0,城戸南蔵院前 DEFAULT,1190805,102,535,1,筑前山手 DEFAULT,1190806,102,535,0,篠栗 -DEFAULT,1190807,102,535,2,門松 +DEFAULT,1190807,102,535,1,門松 DEFAULT,1190808,102,535,0,長者原 -DEFAULT,1190809,102,535,2,原町 +DEFAULT,1190809,102,535,1,原町 DEFAULT,1190810,102,535,0,柚須 DEFAULT,1190811,102,535,0,吉塚 DEFAULT,1190812,102,535,0,博多 +DEFAULT,1191108,102,1187,0,直方 +DEFAULT,1191109,102,1187,1,勝野 +DEFAULT,1191110,102,1187,0,小竹 +DEFAULT,1191111,102,1187,0,鯰田 +DEFAULT,1191112,102,1187,0,浦田 +DEFAULT,1191113,102,1187,0,新飯塚 +DEFAULT,1191114,102,1187,0,飯塚 +DEFAULT,1191115,102,1187,0,天道 +DEFAULT,1191116,102,1187,0,桂川 +DEFAULT,1190801,102,1187,0,桂川 +DEFAULT,1190802,102,1187,0,筑前大分 +DEFAULT,1190803,102,1187,1,九郎原 +DEFAULT,1190804,102,1187,0,城戸南蔵院前 +DEFAULT,1190805,102,1187,1,筑前山手 +DEFAULT,1190806,102,1187,0,篠栗 +DEFAULT,1190807,102,1187,0,門松 +DEFAULT,1190808,102,1187,0,長者原 +DEFAULT,1190809,102,1187,0,原町 +DEFAULT,1190810,102,1187,0,柚須 +DEFAULT,1190811,102,1187,0,吉塚 +DEFAULT,1190812,102,1187,0,博多 DEFAULT,3001518,347,536,0,犬山 DEFAULT,3001519,347,536,0,犬山遊園 DEFAULT,3001520,347,536,0,新鵜沼 diff --git a/src/repository.rs b/src/repository.rs index 67d4cae7..7a84228f 100644 --- a/src/repository.rs +++ b/src/repository.rs @@ -21,7 +21,7 @@ use stationapi::domain::repository::station_repository::StationRepository; use stationapi::domain::repository::train_type_repository::TrainTypeRepository; use stationapi::domain::route_search::RouteNetwork; use stationapi::domain::route_topology::{RouteStop, RouteTopology}; -use stationapi::model::StopCondition; +use stationapi::model::{LineType, StopCondition}; use crate::index; @@ -443,7 +443,8 @@ impl StationRepository for MemStationRepository { Ok(out) } - /// 1. その路線 (station_id 指定時はその駅) に紐づく系統を priority 降順で 1 件選ぶ + /// 1. その路線 (station_id 指定時はその駅) に紐づく系統を priority 降順で 1 件選ぶ。 + /// 候補の無い新幹線は、停車する種別のうち types.id が最も若いものの系統を選ぶ /// 2. その系統の停車駅を sst.id 順で返す /// 3. 空なら路線の全駅を e_sort, station_cd 順で返す /// @@ -481,6 +482,37 @@ impl StationRepository for MemStationRepository { // priority の降順 candidates.sort_by_key(|(priority, _)| std::cmp::Reverse(*priority)); + // 新幹線には各停にあたる種別が無く、特急系だけが停車する。種別の無い全駅を + // 返すとクライアントは trainRoute に渡す lineGroupId を得られないので、 + // 停車する種別のうち id が最も若いもの (のぞみ・はやぶさ等) の系統を選ぶ + if candidates.is_empty() + && index::line_by_cd(line_id as i32) + .is_some_and(|line| line.line_type == Some(LineType::BulletTrain as i32)) + { + let mut bullet_candidates: Vec<(i32, i32)> = Vec::new(); // (types.id, line_group_cd) + for seed in index::stations_by_line(line_id as i32) { + if let Some(target) = station_id { + if seed.station_cd != target as i32 { + continue; + } + } + for sst in index::sst_by_station(seed.station_cd) { + if sst.pass == Some(1) { + continue; + } + let (Some(ty), Some(group)) = + (index::type_by_cd(sst.type_cd), sst.line_group_cd) + else { + continue; + }; + bullet_candidates.push((ty.id, group)); + } + } + if let Some(&(_, group)) = bullet_candidates.iter().min() { + candidates.push((0, group)); + } + } + if let Some(&(_, target_group)) = candidates.first() { let mut typed: Vec<(i32, &index::StationRecord, &index::SstRecord)> = Vec::new(); for sst in index::sst_by_group(target_group) { @@ -1536,4 +1568,72 @@ mod tests { } } } + + const TOKAIDO_SHINKANSEN: u32 = 1002; + const ODAWARA_SHINKANSEN: u32 = 100204; + + /// 新幹線には各停にあたる種別が無い。lineStations は種別の無い全駅ではなく、 + /// id が最も若い種別 (東海道新幹線ならのぞみ) の系統の駅を返す + #[test] + fn bullet_train_line_stations_use_the_youngest_train_type() { + let stations = + block_on(MemStationRepository.get_by_line_id(TOKAIDO_SHINKANSEN, None, None)).unwrap(); + + assert!(!stations.is_empty()); + assert!(stations + .iter() + .all(|s| s.line_group_cd == Some(1) && s.type_name.as_deref() == Some("のぞみ"))); + } + + /// 駅を指定したときは、その駅に停車する種別から選ぶ。 + /// 小田原はのぞみが通過するので、次に若いひかりになる + #[test] + fn bullet_train_line_stations_skip_train_types_passing_the_station() { + let stations = block_on(MemStationRepository.get_by_line_id( + TOKAIDO_SHINKANSEN, + Some(ODAWARA_SHINKANSEN), + None, + )) + .unwrap(); + + // 空のリストでは all() が常に真になるので、小田原が含まれることも確かめる + assert!(stations + .iter() + .any(|s| s.station_cd == ODAWARA_SHINKANSEN as i32)); + assert!(stations + .iter() + .all(|s| s.type_name.as_deref() == Some("ひかり"))); + } + + /// 種別を選ばずに取った新幹線の駅リストでも、駅に付いた種別の系統で + /// trainRoute を引け、駅の並びが一致する + #[test] + fn bullet_train_line_stations_give_the_line_group_for_train_route() { + use stationapi::domain::entity::gtfs::TransportTypeFilter; + use stationapi::use_case::traits::query::QueryUseCase; + let interactor = crate::interactor(); + + let stations = block_on(interactor.get_stations_by_line_id( + TOKAIDO_SHINKANSEN, + None, + None, + TransportTypeFilter::RailAndBus, + )) + .unwrap(); + let line_group_id = stations[0] + .train_type + .as_ref() + .and_then(|tt| tt.line_group_cd) + .expect("駅に種別が付いていない") as u32; + let ids: Vec<u32> = stations.iter().map(|s| s.station_cd as u32).collect(); + + let segments = + block_on(interactor.get_train_route(ids[0], *ids.last().unwrap(), Some(line_group_id))) + .unwrap(); + let route_ids: Vec<u32> = segments + .iter() + .filter_map(|s| s.station.as_ref().map(|st| st.id)) + .collect(); + assert_eq!(route_ids, ids); + } } diff --git a/stationapi/src/use_case/interactor/query.rs b/stationapi/src/use_case/interactor/query.rs index f6b59a36..ff6c6ab7 100644 --- a/stationapi/src/use_case/interactor/query.rs +++ b/stationapi/src/use_case/interactor/query.rs @@ -186,13 +186,19 @@ where .get_by_line_id(line_id, station_id, direction_id) .await?; - let line_group_id = if let Some(sta) = stations - .iter() - .find(|sta| sta.station_cd == station_id.unwrap_or(0) as i32) - { - sta.line_group_cd - } else { - None + // 系統を選べた路線では、全駅がその系統の停車駅として返る (sst_id あり)。 + // stationId が無くても系統は選ばれているので、その種別を付ける。付けないと + // クライアントは trainRoute に渡す lineGroupId を得られない。 + // 系統を選べない路線では種別の無い全駅が返るので、種別は付けない + let line_group_id = match station_id { + Some(station_id) => stations + .iter() + .find(|sta| sta.station_cd == station_id as i32) + .and_then(|sta| sta.line_group_cd), + None => stations + .first() + .filter(|sta| sta.sst_id.is_some()) + .and_then(|sta| sta.line_group_cd), }; let stations = self @@ -5667,6 +5673,8 @@ mod tests { /// 系統の停車駅を返し、付帯情報の付与で要求された駅グループ ID を記録する struct RecordingStationRepository { line_group_stations: Vec<Station>, + /// get_by_line_id (lineStations) が返す駅 + line_stations: Vec<Station>, calls: Calls, } @@ -5687,6 +5695,7 @@ mod tests { Ok(self .line_group_stations .iter() + .chain(self.line_stations.iter()) .filter(|s| ids.contains(&(s.station_g_cd as u32))) .cloned() .collect()) @@ -5717,7 +5726,7 @@ mod tests { _: Option<u32>, _: Option<u32>, ) -> Result<Vec<Station>, DomainError> { - Ok(vec![]) + Ok(self.line_stations.clone()) } async fn get_by_line_id_vec(&self, _: &[u32]) -> Result<Vec<Station>, DomainError> { Ok(vec![]) @@ -5937,6 +5946,7 @@ mod tests { let interactor = QueryInteractor { station_repository: RecordingStationRepository { line_group_stations: stations, + line_stations: vec![], calls: calls.clone(), }, line_repository: StubLineRepository, @@ -5946,6 +5956,76 @@ mod tests { (interactor, calls) } + /// lineStations (get_by_line_id) の駅を返す interactor + fn build_line_interactor(stations: Vec<Station>) -> TestInteractor { + QueryInteractor { + station_repository: RecordingStationRepository { + line_group_stations: vec![], + line_stations: stations, + calls: Calls::default(), + }, + line_repository: StubLineRepository, + train_type_repository: StubTrainTypeRepository, + company_repository: StubCompanyRepository, + } + } + + /// get_by_line_id が系統を選べたときの駅 (系統の停車駅として sst_id が付く) + fn build_typed_line_stations(len: i32) -> Vec<Station> { + build_line_group(len) + .into_iter() + .map(|mut station| { + station.sst_id = Some(station.station_cd); + station + }) + .collect() + } + + /// 種別を選ばずに lineStations で駅を取るクライアントは、駅の種別から + /// trainRoute の lineGroupId を得る。stationId が無くても系統の種別を付ける + #[tokio::test] + async fn line_stations_carry_the_line_group_without_station_id() { + let interactor = build_line_interactor(build_typed_line_stations(4)); + + let stations = interactor + .get_stations_by_line_id(10, None, None, TransportTypeFilter::RailAndBus) + .await + .unwrap(); + + assert_eq!(stations.len(), 4); + for station in &stations { + let train_type = station.train_type.as_ref().expect("種別が付いていない"); + assert_eq!(train_type.line_group_cd, Some(1000)); + } + } + + /// 系統を選べない路線では、駅が別の系統に属していても種別を付けない + #[tokio::test] + async fn typeless_line_stations_stay_typeless_without_station_id() { + // build_line_group の駅は line_group_cd を持つが、系統の停車駅ではない (sst_id なし) + let interactor = build_line_interactor(build_line_group(4)); + + let stations = interactor + .get_stations_by_line_id(10, None, None, TransportTypeFilter::RailAndBus) + .await + .unwrap(); + + assert_eq!(stations.len(), 4); + assert!(stations.iter().all(|s| s.train_type.is_none())); + } + + #[tokio::test] + async fn line_stations_carry_the_line_group_of_the_station_id() { + let interactor = build_line_interactor(build_typed_line_stations(4)); + + let stations = interactor + .get_stations_by_line_id(10, Some(1001), None, TransportTypeFilter::RailAndBus) + .await + .unwrap(); + + assert!(stations.iter().all(|s| s.train_type.is_some())); + } + #[tokio::test] async fn enriches_only_the_requested_range() { let (interactor, calls) = build_interactor(build_line_group(20));