Review of uncommitted Copilot implementation (2026-10-01)
Scope and approach
Review VS Code native sessions, Visual Studio OTLP telemetry, CLI new-layout detection, account quotas, SQLite ingestion and overview display. Preserve existing uncommitted changes; no Agents/model calls/remote usage APIs/commits/pushes.
| Confirmed issue | Correction and acceptance behavior |
|---|---|
| User turns/per-model totals counted as model_call | Store turn/modelTotals as usage_observation; count only toolCallRounds with stable IDs/valid times, limited to persisted main-loop calls |
| Last-call input plus whole-turn output derived a complete total | Retain reported input lower bound and source-cumulative output; default total_tokens unknown, with coverage notice |
| Old keys contributed after modelTotals; model reorder changed identity | Model keys exclude array indices; complete permitted-field snapshots use monotonic revisions, replace prior contributions/recompute daily summaries; corrections may reduce values |
| Snapshot updates conflicted in one lifecycle; truncated reads could publish stale prefixes | Revise after complete replay; partial lines/insufficient time or read allowance/bad lines/invalid tokens/duplicate model keys commit no new usage; source history cleanup/file disappearance retain consumed usage |
| Other extensions in generic chatSessions attributed to Copilot | Filter request agent.id by github.copilot namespace; absent ownership stays unknown |
| Empty-window paths missed; migration copies double-counted; manual file expanded to directory | Add globalStorage/emptyWindowChatSessions; workspace/empty-window roots of one installation share a source; manual files collect only that file |
| Detection checked limits only after reading first line | Read at most limit plus one byte; incomplete first line remains Pending |
| VS checked service only on first line; failed calls without usage lost | Check service per resourceSpans batch; chat CLIENT spans with trace/span identity and time count calls; absent tokens stay unknown; skip aggregate spans |
| Missing VS trace identities and doubleValue TTFT unhandled | Reject missing trace/span IDs; convert doubleValue seconds to milliseconds with upper-bound checks |
| Missing quota fields became zero; fractions rounded; refresh replaced snapshot times | Preserve unknowns; integer milli_requests stores thousandths of requests; source timestamp_utc controls day/deduplication; metadata correction at the same time may update |
| Quotas missed clearing/hard retention, disappeared without tokens, suffered stale responses/wrong timezone | Include clearing/hard retention; old caches cannot restore hard-expired snapshots; display with empty usage; invalidated responses cannot overwrite cards; per-source scheduling does not also collect account quotas; daily queries/date display use statistics timezone |
Snapshot-update counts do not establish call counts. Thinking/premium credits/account quotas/ model context limits are not converted into tokens. Per-model modelTotals in one turn remain usage totals. Invalid cache components/token values produce diagnostics without replacing negative values with zero or choosing the larger conflicting value.
VS Code toolCallRounds count a lower bound of observed main-loop calls; absent rounds do not invent one call per turn. Full auxiliary/sub-Agent/retry/inline coverage remains unverified. Last promptTokens may establish an input lower bound, rather than complete inputs of those calls. Cumulative output depends on upstream deduplication/re-emission behavior.
References
- Read-only local source rechecks independently replay mutation logs/unpack OTLP in Python; output only product identities/field names/counts/token totals, excluding bodies. Temporary scripts/databases: build/copilot-review.
- VS Code usage types distinguish last promptTokens from whole-turn modelTotals; chatModel accumulates completion usage; existing source snapshots were also reviewed.
- Copilot toolCallingLoop saves rounds after runOne returns; toolCallRound defines timestamp as round start. main identifies this research snapshot, without establishing default rules for future releases.
- OpenTelemetry token attributes define cache as an input subset/reasoning as an output subset. VS implementation containment remains unverified per version; generic conventions cannot derive VS uncached input.
- Baselines retain earlier local records: VS Code 1.140.0/Copilot Chat 0.68.0, VS 18.10.1197. This batch rechecks native data; running Code processes/standard extension paths did not supply new installed-version metadata.
Validation progress
Independent comparison: VS Code ten usage-bearing turns, 217 rounds with IDs/timestamps; promptTokens lower-bound sum 3,408,279, cumulative completionTokens 320,141; no modelTotals in these samples. Visual Studio: two chat spans, input=17,470/output=219/cache_read=13,184; two invoke_agent spans excluded from calls/usage.
Full SQLite regressions cover multiple calls/decreasing final corrections/modelTotals replacement/ model reorder/copy/migration/other extensions/manual single files/partial-line continuation/ source history cleanup/failed calls without usage/quota hard retention. Unit cases cover reading limits/null Set/invalid tokens/TTFT number representations/absent quota fields/times/fractions. Browser checks cover quota units and visibility without token history.
Corrected real collection databases match independent expectations and daily_usage:
| Local source | Observed calls | Usage observations | Input | Output | Cache read | Complete total tokens |
|---|---|---|---|---|---|---|
| VS Code workspaceStorage, three logs | 217 | 10 | 3,408,279 (lower bound) | 320,141 (source cumulative) | Unknown | Unknown |
| Visual Studio traces, one log | 2 | 0 | 17,470 | 219 | 13,184 | Not reported |
All ten VS Code diagnostics describe input coverage, rather than corrupt lines. Both adapters add zero on repeats with unchanged calls/usage. VS cache_write/reasoning remain None; validation tools do not fill zero. Two native VS Code model identifiers remain separate source values, without fuzzy merging by similar names. Current quota cache: limit 1500 requests, remaining 1500, used zero; real extraction/ingestion/read-back passed. Earlier remaining=137.4 serves fractional regression only, without replacing the current snapshot.
Real validation from repository root, all exit 0:
- python build/copilot-review/audit.py: independent read-only native data, permitted-field totals.
- cargo run –quiet –locked –manifest-path desktop/src-tauri/Cargo.toml -p llm-usage-core –example real_verify_copilot_chat – <workspaceStorage> <build/copilot-review/chat-real>.
- cargo run –quiet –locked –manifest-path desktop/src-tauri/Cargo.toml -p llm-usage-core –example real_verify_vs_copilot – <traces> <build/copilot-review/vs-real>.
- python build/copilot-review/compare.py: real details/daily totals match independent expectations, PASS.
- cargo run –quiet –locked –manifest-path desktop/src-tauri/Cargo.toml -p llm-usage-core –example copilot_quota_probe: quota read-back PASS; no paths/login/bodies output.
Added seven Rust unit/six SQLite integration cases plus browser quota-unit/empty-history/statistics- timezone assertions. Final environment: local Windows, PowerShell 7.6.6, Node.js 24.21.0, Cargo 1.98.1. Root commands all exit 0:
- npm run verify: Markdown/assets/scripts, Svelte zero errors/warnings, Rust fmt/Clippy, 716 Rust tests and frontend production build.
- npm run test:browser: installed Edge interactions/new quota assertions.
- git diff –check: passed.
Ignored logs: build/copilot-review/verify-final.log, browser-final.log, real-compare.log. These are local static/test/real persisted-data checks, without replacing three-platform or actual IDE interaction acceptance.
Retained limits
Native and OTel file/receiver records for one interface still require exclusive source selection in the source UI. Shared Agent names do not deduplicate without verified cross-format call association; do not add them. Automatic cross-format deduplication remains unfinished; VS TEMP files do not promise complete history. Later configuration design confirms pause stops scanning without excluding historical contributions. Introducing OTel requires statistics-source selection; merely disabling native sources still mixes saved totals. JetBrains has static plugin checks only, without real IDE/outfile samples, and remains documentation-level. The CLI chronicle no-per-call-token conclusion applies only to checked records; unknown newer versions receive no guessed fields.
The existing prerelease database policy prompts rebuilding on version mismatch, without incremental migrations. This batch changed parser/source namespaces and moved schema to 10; old trial DBs follow that recovery flow. Comparisons use new isolated databases without modifying user application DBs or running actual JetBrains services/production systems.
Later corrections (2026-10-01): health and unknown-field display
Two locally reproduced display issues were corrected using independent real-data checks, without modifying the user application DB:
- Source notice to check one file: turn_input_incomplete describes promptTokens covering only the last call’s input. Previously any diagnostic degraded health. This intrinsic format limit is neither a bad record nor reconciliation mismatch (architecture). Only other diagnostics (bad lines/invalid tokens/duplicate keys/missing ownership, etc.) now degrade health; coverage-only notices stay active. session_log_v3 uses: all diagnostics turn_input_incomplete => active.
- Today’s model-details unknown-field count 396: tool rounds are model_call with quality_bucket=unknown/per-round tokens unknown; turn observations carry usage. Previously each round added one input/one output unknown field. recompute_day/enrich_hourly_metadata now count calls/events for records with quality_bucket=‘unknown’ and no known token fields, without adding input/output/total unknown counts, like transport_attempt exclusions. This applies across adapters: failed/no-usage calls stay visible through call_count without inflating unknown-field counts.
Real read-only real_verify_copilot_chat on local workspaceStorage, four logs: before correction source_files active×2/degraded×1; vscode-copilot-chat model details input_unknown+output_unknown=434 (round-model claude-opus-4.8 was an independent all-unknown row). After correction active×4/ degraded×0, unknown fields zero, calls=222/observations=12, repeat adds zero. Twelve turn_input_incomplete diagnostics remain visible without degrading health.
New regressions: session_log_v3 unit turn_input_incomplete_alone_keeps_source_active and integration copilot_rounds_stay_active_and_do_not_inflate_unknown_fields. Existing no-usage cases updated consistently (classification_v03/kilo/omp/pi/storage_jobs): syn-a2/failed calls’ input_unknown_count changes 1 to 0 with unchanged call counts.
Old-cursor recovery and further review (2026-10-02)
Preserve preceding uncommitted repairs/other-task changes and recheck implementation, field quality, scanner framework and current local data. Upstream IChatUsage / IChatUsageModelTotal and ToolCallRound confirm last input/whole-turn usage/identity-and-time-only rounds are distinct. Upstream main is a cross-check for this batch; real native files have version=3. Installed versions retain earlier checks; no newer version is presumed accepted.
Additional findings/fixes:
- Fully consumed unchanged files skip parsing, leaving old degraded health/unknown-field daily totals after health/SQL-only fixes. Copilot parse_context gains scan_policy_version; absent current marker permits one full replay. Preserve revision/tracked_events; advance existing monotonic revisions to recompute daily/hourly totals. Bad snapshots never mark success; complete success restores skipping. No cursor-context reset/revision decrease/history deletion; schema/parser-format version unchanged.
- QualityBucket::of omitted output_reasoning/source_total, misclassifying rows containing only those as unknown. Excluding unknown rows could consequently hide real missing fields. Add both token fields. Zero/cache components/input-only/output-only still represent valid partial usage; estimates remain excluded from reported totals.
New regressions first failed: old-cursor refresh returned unchanged instead of complete; reasoning-only hourly unknown counts zero instead of one each for input/output/total. Six new Rust unit/SQLite cases cover recovering the fixed 198-round sample’s 396 display/revision 7=>8, model rows/real missing output, bad-line degradation/recovery, no success marker for bad snapshots, and recomputing old history after source cleanup. Coverage includes day/hour/week/month, unknown filters/zeros/cache/reasoning/source totals/estimates/transport attempts.
Local Windows / PowerShell 7.6.6 / Node.js 24.21.0 / Cargo 1.98.1, read-only current app DB: five files active×3/degraded×2; 235 rounds/13 observations; input lower bound 3,832,423, cumulative output 487,818, complete total unknown; daily input/output missing-field sum 470. The live sample grew. The fixed 198-round synthetic sample exactly reproduces 396, rather than presenting it as the current measurement.
SQLite Online Backup made an isolated build/copilot-review/current-app copy. Native workspaceStorage refresh on the copy with old cursors/Asia/Shanghai produced active×5/ degraded×0 and input/output missing fields 470=>0. Calls/observations/input/output/unknown total stay unchanged; all 13 observations retain total_unknown_count=13, without inventing complete totals. Independent Python mutation-log replay matches counts/tokens; repeats add zero. real_verify_copilot_chat adds –reuse/–timezone for this check, passed an isolated DB only. User live DB/IDE settings/OTel events.jsonl were unchanged; no Agents/model requests ran.
Completed root checks, all exit 0:
- cargo test –locked –manifest-path desktop/src-tauri/Cargo.toml -p llm-usage-core –test copilot_ide_contract –test unknown_field_counts: 12 integration cases.
- cargo test –locked –manifest-path desktop/src-tauri/Cargo.toml -p llm-usage-core copilot – –nocapture: focused unit/integration regressions.
- python build/copilot-review/audit.py: independent replay of current native data.
- cargo run –quiet –locked –manifest-path desktop/src-tauri/Cargo.toml -p llm-usage-core –example real_verify_copilot_chat – <workspaceStorage> build/copilot-review/current-app –reuse –timezone=Asia/Shanghai: real old-cursor recovery/no new usage on repeat, PASS.
- python build/copilot-review/compare_current.py: native data/fixed-copy details/daily totals match; no conflicts/degraded files, PASS.
- npm run verify: 164 Markdown files/assets/three scripts/13 UI tests/Svelte zero errors or warnings/fmt/Clippy/775 Rust tests/frontend production build. Two previously ignored tests (models.dev live smoke/local configuration audit) were not run here.
- npm run test:browser: installed Edge interaction regressions.
- git diff –check: passed; initial-patch comparison preserves 16 other uncommitted files.
Ignored build/copilot-review logs: regression-before.log (two failed assertions), regression-after.log, copilot-regression.log, current-raw-audit.log, current-app-replay.log, current-app-after.log, current-compare.log, verify-current.log, browser-current.log. Documentation lint also ran after record/plan synchronization. The validated copy was removed after acceptance, retaining redacted summary logs only. No running desktop binary was installed or replaced; refreshing an updated build recovers these old states without clearing the DB.