Skip to content
2 changes: 1 addition & 1 deletion .claude/skills/benchmark-gql/queries.json
Original file line number Diff line number Diff line change
Expand Up @@ -176,7 +176,7 @@
"name": "connectedRoutes",
"weight": "heavy",
"uses": ["TrainTypeTop", "StationCore", "LineCore", "TrainTypeCore"],
"query": "query Bench_connectedRoutes($from: Int!, $to: Int!) { connectedRoutes(fromStationGroupId: $from, toStationGroupId: $to) { estimatedMinutes transferCount legs { trainType { ...TrainTypeTop } fromStation { ...StationCore } toStation { ...StationCore } } } }",
"query": "query Bench_connectedRoutes($from: Int!, $to: Int!) { connectedRoutes(fromStationGroupId: $from, toStationGroupId: $to) { legs { trainTypes { ...TrainTypeTop } fromStation { ...StationCore } toStation { ...StationCore } } } }",
"variables": { "from": 1130205, "to": 1130212 },
"note": "渋谷 → 池袋。乗換を含む経路探索 (RAPTOR、初回は系統網の構築を含む)"
},
Expand Down
4 changes: 2 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,7 +65,7 @@ The Worker is the workspace root package. `stationapi`, `preprocessor`, and `dat
- **Endpoint benchmarks** – `make bench` (or `python3 .claude/skills/benchmark-gql/bench.py`) replays every `Query` field against production (`gql.trainlcd.app`, script `stationapi`) and staging (`gql-stg.trainlcd.app`, script `stationapi-stg`) and writes a Markdown report under `benchmarks/`. Both environments embed the same data, so any difference is implementation — which makes this the way to see what a `dev`-to-`master` release will do to performance before it ships. Besides client latency it records the Worker's `cpuTime`, read from `wrangler tail --format json` and matched to each request by `cf-ray`; the tail is filtered on a per-run request header, so production's live traffic does not leak into the sample. Collecting CPU time needs the `workers_tail (read)` scope, and the run sends hundreds of real requests to production — it is not a routine check. Add a case to `.claude/skills/benchmark-gql/queries.json` whenever a `Query` field is added, and never edit an existing case's variables: the reports are meant to stay comparable across runs.

## GraphQL Query Overview
- **Stations** – `station`, `stations`, `stationGroupStations`, `stationsNearby`, `lineStations`, `stationsByName`, `lineGroupStations`, `lineListStations`, `lineGroupListStations`. `QueryInteractor` enriches stations with lines, companies, station numbers, and train types. `lineStations` resolves the line's local train-type group (rail `kind` 0/1 or a `priority > 0` type); when no such group exists — bus lines only carry `BusRoute` (`kind` 7, `priority` 0) variants — it falls back to the line's plain typeless station list so bus stop listings never return empty.
- **Stations** – `station`, `stations`, `stationGroupStations`, `stationsNearby`, `lineStations`, `stationsByName`, `lineGroupStations`, `lineListStations`, `lineGroupListStations`. `QueryInteractor` enriches stations with lines, companies, station numbers, and train types. `lineStations` resolves the line's local train-type group (rail `kind` 0/1 or a `priority > 0` type); when no such group exists — bus lines only carry `BusRoute` (`kind` 7, `priority` 0) variants — it falls back to the line's plain typeless station list so bus stop listings never return empty. `stationsByName` with `fromStationGroupId` returns the stations reachable from there: stations sharing a line group with the origin (`line_group_cd` set, `has_train_types` true), same-line stations when either side has no line group, and — for rail — stations reachable by transferring, i.e. those for which `connectedRoutes` with `viaLineId` set to the station's line returns a route (`line_group_cd` empty, `has_train_types` false). The transfer check uses `RouteTopology` (`stationapi/src/domain/route_topology.rs`), a time-free copy of the `connectedRoutes` network built straight from the index without `Station` entities or time estimates (about 20 ms instead of about 190 ms, cached in its own `OnceLock`); `RouteNetwork` holds the same topology, both share `trim_pattern` and `line_group_rows`, and a real-data test asserts the two are equal. The check is a ride-limited BFS over line groups (a few ms) plus, only for destinations that are cut vertices of the station–line-group graph, a check that arrives without stopping over at the destination group — otherwise a branch's junction station (Ishibashi-handai-mae on the Minoo Line) would be listed although reaching it on that branch means riding out and back.
- **Lines** – `line`, `lines`, `linesByName`. Results include company data and computed line symbols based on repository helpers.
- **Routes** – `routes`, `connectedRoutes`, `estimateArrivalTimes`, `trainRoute`. Paging tokens are currently empty (pagination not implemented).
- **`trainRoute`** – Takes the line group's stops from the repository *before* any enrichment, slices them to the requested `fromStationId`–`toStationId` range (reversing when the request runs backwards), and only then attaches lines, companies, station numbers, train types, and nearby bus routes. Enrichment is per-station and independent, so slicing first does not change any segment; enriching the whole line group first made a three-station request cost the same as a 250-station one. Keep the order — the cost of this query must stay proportional to the requested range, not to the line group.
Expand All @@ -76,7 +76,7 @@ The Worker is the workspace root package. `stationapi`, `preprocessor`, and `dat
- **Bus stop translations (readings & English)** – GTFS-JP `translations.txt` layouts differ per feed, so `load_gtfs_translations` resolves columns by header name (Seibu ships 6 columns without `record_sub_id`; Keio and the Tokyu community feeds ship 7) and indexes each `stop_name` translation under both keys it may use: `record_id` (== the stop_id, Seibu — with the "-NN" pole suffix also mapped to the parent stop_id) and `field_value` (== the Japanese stop_name, Keio / Tokyu community, where `record_id` is left empty). `import_gtfs_stops` then looks a stop's translation up by stop_id first, then by name. Keying only by `record_id` (the previous behavior) silently dropped every field_value-keyed feed, leaving `station_name_k` filled with the kanji stop_name and `station_name_r` empty. Readings arriving as half-width katakana (`ニシハチオウジ`, Keio / Tokyu community) are folded to full-width via `romaji::to_fullwidth_katakana()` before storage.
- **Bus English-name fallback** – When a feed provides no English (`en`) translation for a stop — e.g. Tokyu Bus ordinary-route JSON, which carries only `dc:title` and `odpt:kana` — `src/domain/romaji.rs::romaji_display_name()` derives a modified-Hepburn romanization (with macrons for long vowels, matching the curated rail style: Tōkyō / Kyōto / Shin-Ōsaka) from the kana reading, and the GTFS reader fills `stop_name_r` with it. The fallback never overwrites a real `en` value, and a reading with no convertible kana stays `NULL` rather than emitting a partial transcription. Because `stop_name_r` is the single upstream source that fans out into the `stations` projection, `search_by_name`, and the romanized bus route/headsign names, this supplements every English-facing surface at once. When projecting into `stations`, `station_name_rn` is filled with the plain-ASCII spelling via `romaji::strip_macrons()` (Tōkyō → Tokyo), mirroring the rail dataset's `_r` (macron) / `_rn` (macron-free) column pair.
- **TTS metadata** – `Station`, `StationNested`, `Line`, `LineNested`, `TrainType`, and `TrainTypeNested` expose `name_ipa` / `name_roman_ipa` plus `name_tts_segments` for multi-segment pronunciation output. Use `name_tts_segments` when clients need per-token SSML construction for mixed-language names such as `Kasai-Rinkai Park`.
- **Connected routes** – `connectedRoutes` finds transfer routes automatically, like a journey planner, using a frequency-based RAPTOR search in `stationapi/src/domain/route_search.rs`. Each rail line group is a pattern, station groups are the transfer nodes, and ride times come from `arrival_estimation`; bus lines are excluded. The cost adds a per-boarding wait by `TrainTypeKind` (limited express 15 min, express / high-speed rapid 5 min, others 3 min) and a 3-minute transfer walk — without the wait, infrequent limited expresses would beat the Yamanote Line. Rounds give the time/transfer Pareto set; alternatives come from re-searching with one leg's parallel line groups banned along that leg (at most 8 searches), and are dropped beyond 1.15 × best + 15 min or with two more transfers than the Pareto set. Routes that pass the same station group in two different legs (backtracking to re-board a banned train) are dropped. Results are ranked by cost + 5 min per transfer, capped at 6, and carry `estimatedMinutes` (excluding the first wait) and `transferCount`. The network (every rail line group plus its time estimates, about 190 ms natively) is built lazily into a `OnceLock` by `StationRepository::get_route_network` on the first `connectedRoutes` call, so other queries never pay for it. Each route is a list of `legs` shaped for the app's one-train-at-a-time flow: every leg carries the same `TrainType` that `routeTypes` returns (a real `groupId`, so the client can call `lineGroupStations`) plus its boarding and alighting `Station`, both on the line that leg runs on — so at a transfer the previous leg's alighting station and the next leg's boarding station may be different stations of one station group. Each leg also carries `trainTypes`, every train type usable on that leg: it is exactly `routeTypes(boarding station group, alighting station group, alighting station's line)` (same dedup, same `lines`, same order — the use case calls `get_train_types`), because the search collapses parallel services such as local and rapid into one route and the app needs them to list types and default to the local. `trainType` stays as the search's representative and may be absent from `trainTypes` after that dedup. `viaLineId`, like `routeTypes`, is the line of the tapped search result and keeps only routes whose last leg arrives on that line. `docs/architecture.md` (乗換経路探索) has the details.
- **Connected routes** – `connectedRoutes` finds transfer routes automatically, like a journey planner, using a frequency-based RAPTOR search in `stationapi/src/domain/route_search.rs`. Each rail line group is a pattern, station groups are the transfer nodes, and ride times come from `arrival_estimation`; bus lines are excluded. The cost adds a per-boarding wait by `TrainTypeKind` (limited express 15 min, express / high-speed rapid 5 min, others 3 min) and a 3-minute transfer walk — without the wait, infrequent limited expresses would beat the Yamanote Line. Rounds give the time/transfer Pareto set; alternatives come from re-searching with one leg's parallel line groups banned along that leg (at most 8 searches), and are dropped beyond 1.15 × best + 15 min or with two more transfers than the Pareto set. Alternative routes that stop at the same station group in two different legs (backtracking to re-board a banned train) are dropped; pass-through stations are not counted, and the Pareto routes of the first search are never dropped this way (otherwise a station `stationsByName` reports as reachable could get no route). Results are ranked by cost + 5 min per transfer and capped at 6. The time and transfer count stay internal (the API does not return them). The network (every rail line group plus its time estimates, about 190 ms natively) is built lazily into a `OnceLock` by `StationRepository::get_route_network` on the first `connectedRoutes` call, so other queries never pay for it. Each route is a list of `legs` shaped for the app's one-train-at-a-time flow: every leg carries its boarding and alighting `Station`, both on the line of the line group the search rode — so at a transfer the previous leg's alighting station and the next leg's boarding station may be different stations of one station group — and `trainTypes`, every train type usable on that leg (real `groupId`s, so the client picks one and calls `lineGroupStations`): it is exactly `routeTypes(boarding station group, alighting station group, alighting station's line)` (same dedup, same `lines`, same order — the use case calls `get_train_types`), because the search collapses parallel services such as local and rapid into one route and the app needs them to list types and default to the local. `viaLineId`, like `routeTypes`, is the line of the tapped search result and keeps only routes whose last leg arrives on that line. `estimateArrivalTimes` and `trainRoute` accept `legs: [RouteLegInput!]` (the `groupId` of the train type picked from each leg's `trainTypes`, plus the leg's `fromStation.id` and `toStation.id`) and then return values for the whole transfer route. A leg endpoint missing from the chosen line group is matched by station group (the picked local may stop at another line's station of the same group), preferring an exact `station_cd` and, among same-group candidates, the pair giving the shortest slice (through services list two stations of a junction group). ETA estimates each leg on its own line group only and chains them from the origin, adding the 3-minute walk and the next train type's wait at each transfer (the same allowance the search ranks by), returning one route with an empty `id`; `trainRoute` concatenates each leg's segments (each leg restarts at distance 0). Both slice legs with the same function, taking the shorter arc on loop lines, so their station sequences match. More than `MAX_RIDES` (6) legs — more than `connectedRoutes` ever returns — legs that do not connect, ends that differ from `fromStationId` / `toStationId`, or combining `legs` with `viaLineIds` / `directionId` / `lineGroupId` are errors. `docs/architecture.md` (乗換経路探索) has the details.
- Changes to the published contract require coordinated updates to `schema/public.graphql`, the async-graphql types in `src/graphql/`, and, when the shape of a value changes, `stationapi/src/model.rs` and the DTO conversions.

## Version Control (Git)
Expand Down
Loading
Loading