Repository navigation
feat(corpus): expand throughput-phase corpus with validated bundles - #2382
Conversation
Promote four validator-clean DuckDB TPC-H SF1 throughput bundles (streams 3, validation passed) into results-data/bundles with maintainer manifests. Teach generate_corpus_inventory.py to extract benchmark test_type as phase and summarize by_phase, so throughput coverage is visible (4 throughput vs 253 power). Regenerate corpus-inventory.json.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 03d35ec1a6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The four throughput bundles added earlier carry run timestamps of 2026-08-11, 2026-08-12, and 2026-08-22, all before the 2026-08-23 measurement trust boundary recorded in results-data/CORPUS_NOTES.md. That note withdraws pre-boundary measurements and permits restoration only through fresh runs, so those bundles are removed. They are replaced with four genuinely generated post-boundary runs: DuckDB TPC-H SF1 throughput (load,throughput phases, 3 streams, seed 20260709, official mode), executed locally on 2026-09-25 with validation passed and 66/66 queries passing. Each carries a maintainer-run manifest with its SHA-256. The corpus inventory is regenerated (257 bundles: 253 power, 4 throughput). Validation: results-data/validate_corpus.py passes (all 42 cohorts meet the >=3-identity floor); submission validator passes 4/4 with no errors, warnings, or overrides; inventory --check is current; trust-boundary and inventory unit tests pass (24 passed).
Regenerate the corpus inventory for the merged bundle set (develop withdrawals plus the new post-boundary throughput runs); all 42 cohorts meet the identity floor.
The corpus expansion adds four DuckDB TPC-H SF1 throughput bundles that the committed seed does not cover, so the bijection and union tests fail. Regenerate the seed against the live accepted ref: the accepted bundle set is unchanged at 244, and the four new bundles enter as legacy overlay entries pending mirror acceptance.
Resolve phase via benchmark.test_type with phases-object fallback and never emit None, so explicit nulls cannot crash inventory sorting. Throughput runs move into phase-suffixed cohorts instead of polluting power cohorts, and the seed spec records the segregation rule.
The prior Develop PR run completed its test job but its aggregator stayed queued for over an hour, blocking rerun-failed-jobs. Retrigger with an empty commit; squash-merge drops it from history.
Regenerate the corpus inventory from the merged tree so the four post-boundary throughput bundles coexist with the current develop corpus.
… into fix/throughput-corpus-expansion
|
Maintainer review: refreshed onto develop (merge b98f9b3) so the visual gate compares against the current baseline set. Zero open threads. CI re-running on the refreshed head; enqueuing after green. |
|
Status note (takeover sweep): the remaining failures are Public-site visual regression (changed /results/ captures at all four viewports) and its acceptance gate. The UI change is intentional in this PR, so the visual baselines need regeneration per the repo visual-baseline policy, then re-push. No code fix needed beyond the baseline update. |
|
This PR touches soundness review paths ( |
Bundles list skipped phases as {"status": "NOT_RUN"}, which are truthy, so
a bundle without benchmark.test_type was always inferred as power. Count only
phases that ran, and give the cohort-separation test distinct platforms so a
merged cohort cannot pass.
Both phases map to the bare cohort key, so the later tuple key replaced the earlier one and dropped platforms. Aggregate members by the final key.
|
Soundness review: external agy, gemini-3.1-pro-high, read-only. Reviewed the throughput-phase corpus expansion (inventory generator and its tests, ledger seed, corpus inventory, spec note, and the four new tpch_sf1 DuckDB throughput bundles and manifests) against the merge base with develop. The review ran in three passes. First pass: four High findings. Two were fixed: (1) the phase fallback treated Second pass: one High. Power and unknown phases share the bare cohort key, so the later tuple key overwrote the earlier one and dropped platforms. Cohort members are now merged by final key (new regression test), fixed in 626057d. Third pass (626057d): no findings at any severity. Final verdict SHIP. Hosted checks and the queue candidate must still pass. |
Regenerate the corpus inventory and ledger seed for 333 bundles, and update the README counts for the new single-identity throughput cohort.
|
|
|
/gemini review |
|
Stand-in oracle review: APPROVE d51b983 |
Promote four validator-clean DuckDB TPC-H SF1 throughput bundles
(streams 3, validation passed) into results-data/bundles with
maintainer manifests. Teach generate_corpus_inventory.py to extract
benchmark test_type as phase and summarize by_phase, so throughput
coverage is visible (4 throughput vs 253 power). Regenerate
corpus-inventory.json.
Soundness review: