Skip to content

PROBE-ORACLE-FUNNEL Stage 0: pre-registration + measured results + the legend-is-not-the-grammar finding - #894

Merged
AdaWorldAPI merged 1 commit into
mainfrom
claude/x265-x266-plans-review-h9osnl
Aug 5, 2026
Merged

PROBE-ORACLE-FUNNEL Stage 0: pre-registration + measured results + the legend-is-not-the-grammar finding#894
AdaWorldAPI merged 1 commit into
mainfrom
claude/x265-x266-plans-review-h9osnl

Conversation

@AdaWorldAPI

Copy link
Copy Markdown
Owner

What this is

The consumer-side record of PROBE-ORACLE-FUNNEL Stage 0 — the first measurement over the ogar-loco funnel the wishlist delivery (OGAR #244) unlocked. Docs/board only; no Rust code changes here. The harness itself lives OGAR-side (PR #245) per the no-cross-dependency ruling in the merged handover: candidates and legends cross as data, never as a crate edge.

Contents

  • .claude/plans/oracle-funnel-probe-v1.md (new) — the staged probe plan. §1–§4 are the pre-registration (three deterministic arms R/G/W bracketing what an LLM could emit; expectations E1–E4 including the E4 discrimination KILL at a 50-point bar; the Stage 1 LLM arm and Stage 2 Gadamer-bag arm both explicitly gated), written before the run. §5 is the measured result, appended after.
  • Results: E1–E4 all met. Floor (uniform bytes) 0.5% survival; legend-constrained 2.5%; stack-aware walk 100%. Spread 99.5 points — the funnel discriminates, so the Stage-1 LLM arm is justified on this vocabulary (still gated on operator word + API).
  • EPIPHANIES.mdE-THE-LEGEND-IS-NOT-THE-GRAMMAR-1 (FINDING, measured): the mid-model arm landed at the floor, not midway. Knowing the full vocabulary table buys almost nothing; ~all of the funnel's selectivity is the stack discipline. Stage-1 consequence: the prompt must teach the postfix discipline (or constrain decoder-side) — the ~382-token legend alone is nearly worthless. Also: the per-gate refusal distribution is a diagnosis channel (Uncovered = doesn't know the vocabulary; Underflow = doesn't know the grammar), still inside W-4's raw-loopable validity half.
  • STATUS_BOARD.md — S0 row MEASURED; S1 (rig LLM arm) and Stage-2 PROBE-GADAMER-BAG rows recorded as GATED.
  • AGENT_LOG.md — main-thread run entry (pre-register → build → run → record order kept).

Cross-refs

🤖 Generated with Claude Code

https://claude.ai/code/session_01K3RyLEbuNSHxxB3NTTrGki


Generated by Claude Code

Pre-registered (E1-E4 written before the run) 3-arm funnel probe over
the ogar-loco gates; harness OGAR PR #245 per the no-cross-dep ruling.
All four expectations met; the E4 discrimination KILL passes at 99.5
points vs the 50-point bar, justifying the (still-gated) Stage-1 LLM
arm. New finding E-THE-LEGEND-IS-NOT-THE-GRAMMAR-1: legend-constrained
generation lands at the floor (2.5% vs 0.5%); the stack discipline
carries ~all the funnel's constraint - Stage-1 prompts must teach the
discipline, not just serialize the table. Board: STATUS_BOARD rows
(S0 measured; S1/S2 gated), AGENT_LOG entry.
@coderabbitai

coderabbitai Bot commented Aug 5, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@AdaWorldAPI, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 43 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 6cd45830-7ea5-4656-88d4-0e405379989f

📥 Commits

Reviewing files that changed from the base of the PR and between 75acb67 and e4967e3.

📒 Files selected for processing (4)
  • .claude/board/AGENT_LOG.md
  • .claude/board/EPIPHANIES.md
  • .claude/board/STATUS_BOARD.md
  • .claude/plans/oracle-funnel-probe-v1.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cursor

cursor Bot commented Aug 5, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_1b0282fc-3ccd-4f1b-bc69-ce2cba347067)

@AdaWorldAPI
AdaWorldAPI marked this pull request as ready for review August 5, 2026 14:30
@AdaWorldAPI
AdaWorldAPI merged commit d7a6efc into main Aug 5, 2026
1 of 2 checks passed
AdaWorldAPI pushed a commit that referenced this pull request Aug 5, 2026
Records the cross-repo wishlist round-trip (OGAR #243 handover ->
#244 delivery -> #245 harness + #894 record) and the measured Stage-0
outcome. Hygiene-only PR per the termination clause - no arc entry of
its own.
AdaWorldAPI added a commit that referenced this pull request Aug 5, 2026
…w-h9osnl

Board hygiene: post-merge arc entry for #894 + LATEST_STATE head
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants