Skip to content

Fix token accounting observability for skill-based extraction - #3681

Open
hopstreax wants to merge 1 commit into
Graphify-Labs:v8from
hopstreax:investigate/3658-cost-token-accounting
Open

hopstreax wants to merge 1 commit into
Graphify-Labs:v8from
hopstreax:investigate/3658-cost-token-accounting

Conversation

@hopstreax

Copy link
Copy Markdown
Contributor

Summary

Fixes #3658 by distinguishing explicitly observed zero token usage from unobserved token usage across Graphify's skill-based extraction pipeline.

Previously, chunk schemas initialized missing token counts to 0, which could make unavailable usage appear as genuine zero-token measurements and produce inaccurate cost.json totals and reports.

Changes

  • Initialize unobserved input_tokens / output_tokens as null in extraction schemas.
  • Normalize token usage across common provider formats:
    • input_tokens / prompt_tokens / inputTokens
    • output_tokens / completion_tokens / outputTokens
    • nested usage / modelUsage
    • Claude cache token fields.
  • Update chunk merging to distinguish observed 0 from unobserved null.
  • Preserve token observability through semantic and extraction merges.
  • Update cost persistence to:
    • skip null values during cumulative arithmetic,
    • preserve null in individual run records,
    • track has_unrecorded_usage.
  • Make token reporting render unobserved values as unknown instead of crashing or displaying misleading zeros.
  • Apply the changes consistently across generated skills and monolith platforms.
  • Add regression coverage for report rendering, chunk merging, cost persistence, and skill generation.

Verification

  • 151 passed
  • python -m tools.skillgen --check ✅
  • python -m tools.skillgen --monolith-roundtrip ✅
  • git diff --check ✅

@graphify-labs graphify-labs Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Worth a look — the grounded gate found no coupling regressions or blocking issues, but 2 advisory finding(s) below merit a look before merge.

Formal verification. No changes could be formally verified in this run.


Graphify review — findings

Tracks token counts as None when a subagent's usage is unobserved instead of silently coercing to 0, so merge-chunks, report.generate, and the skill merge/cost-tracker steps report "unknown" rather than a fabricated zero. Sums only real numeric counts (rejecting booleans) across chunks and, in the cost tracker, flags a run with has_unrecorded_usage and skips adding None totals so cumulative figures stay honest. Updates the skill instructions to initialize chunk token fields to null, normalize varied provider usage shapes (prompt_tokens/inputTokens, nested usage/modelUsage, Claude cache tokens), and forbid writing 0 for unmeasured usage.

Worth a look

  • Non-finite token counts from untrusted chunks can crash merge — graphify/cli.py:4769 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Non-finite token costs can crash report generation — graphify/report.py:184 · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 1651 functions depend on the 1217 functions this change touches.

Health — this change adds coupling hotspots:

  • new: _rebuild_code() — 137 callers, 54 callees
  • new: dispatch_command() — 2 callers, 125 callees
  • new: generate() — 35 callers, 7 callees
  • new: run_pipeline() — 8 callers, 13 callees
  • new: make_inputs() — 17 callers, 5 callees
  • new: render() — 13 callers, 5 callees
  • new: audit_coverage() — 8 callers, 6 callees
  • new: _stale_graph_sources() — 7 callers, 6 callees
  • …and 9 more — each is listed as a finding

Verification — 1651 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 1595 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

287 of 287 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — impact, full-run-safety
  • tests/test_affected_member_seed.py — full-run-safety
  • tests/test_agents_platform.py — impact, full-run-safety
  • tests/test_analyze.py — full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — full-run-safety
  • tests/test_astro_import_ids.py — full-run-safety
  • tests/test_atomic_canvas_export.py — full-run-safety
  • tests/test_atomic_version_stamp.py — full-run-safety
  • tests/test_atomic_writes.py — full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — full-run-safety
  • tests/test_benchmark_raw_graph.py — full-run-safety
  • tests/test_build.py — full-run-safety
  • tests/test_build_merge_dedup_scope.py — full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — full-run-safety
  • tests/test_build_merge_shrink_guard.py — full-run-safety
  • tests/test_builtin_global_type_refs.py — full-run-safety
  • tests/test_cache.py — full-run-safety
  • tests/test_callflow_html.py — full-run-safety
  • tests/test_cargo_introspect.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — full-run-safety
  • tests/test_case_sensitive_resolution.py — full-run-safety
  • tests/test_charmap_encoding.py — full-run-safety
  • tests/test_chunking.py — full-run-safety
  • tests/test_cjs_module_extension.py — full-run-safety
  • tests/test_claude_cli_backend.py — full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — full-run-safety
  • tests/test_codebuddy.py — impact, full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — full-run-safety
  • tests/test_confidence.py — impact, full-run-safety
  • tests/test_corrupt_graph_json.py — full-run-safety
  • tests/test_cost_persistence.py — impact, changed-test, full-run-safety
  • tests/test_cpp_nested_and_cli.py — full-run-safety
  • tests/test_cpp_objc_cross_file_calls.py — full-run-safety
  • tests/test_cpp_preprocess.py — full-run-safety
  • tests/test_cross_extension_reexport_self_cycle.py — full-run-safety
  • tests/test_cross_language_call_resolution.py — full-run-safety
  • tests/test_cross_repo_external_call_guards.py — full-run-safety
  • tests/test_cross_repo_member_calls.py — full-run-safety
  • tests/test_cross_repo_shared_types.py — full-run-safety
  • … and 237 more

non-code file(s) changed (graphify/skill-agents.md, graphify/skill-aider.md, graphify/skill-amp.md, graphify/skill-claw.md, graphify/skill-codex.md …) → running the full suite for safety (a code graph can't see config/fixture/data deps)

changed code file(s) with no mapped test (graphify/skill-agents.md, graphify/skill-aider.md, graphify/skill-amp.md, graphify/skill-claw.md, graphify/skill-codex.md …) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

Formal verification

Could not verify: Could not verify dispatch\_command.

The verifier did not have enough to check dispatch\_command, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 23 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly SystemExit — names the real obstacle, not a sampling gap)

Could not verify: Could not verify generate.

The verifier did not have enough to check generate, so it is saying so rather than guessing. No false assurance is the whole point.

Guarantee: No guarantee either way, this is an honest abstention, not a pass.

Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 200 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly ValueError — names the real obstacle, not a sampling gap)

· 17 more finding(s) on lines outside this diff (see the check run).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant