Skip to content

Implementation and acceptance results for remaining work

Date: 2026-10-04; version 0.2.1, schema 11; cwd repository root. Windows 11 Pro x64 10.0.26300; Ryzen 9 9950X3D (16 cores/32 threads), about 125 GiB RAM; Node 24.21.0, Rust 1.98, Tauri CLI 2.12.0, WebView2 154.0.4258.53. Development-machine results do not verify the proposed four-core/16 GiB baseline. This is a completed development validation record; Plan.md maintains the active execution plan.

Implementation and requirements

  • Daily queries add rebuildable derived tables. Read-only connections use them only for valid partitions or partitions established empty by authoritative tables. Provider filters merge identical source/session identities before DISTINCT; model/Agent filters use expression indexes. Original events/revisions/cursors/unknown fields remain. Older writers invalidate derived data and remove derived identities. Regressions cover old-layout preparation, cross-connection writes, rollback, DST, retention/clearing and optional-SUM overflow. See query design.
  • Six chart types use SVG; features/themes/selections retain actual query results.
  • Windows optional close-to-tray, explicit exit, power-saving pause and directory notifications are implemented. Watch at most 32 directories; coalesce after two quiet seconds; pause releases handles. Fixed daily/weekly sources are excluded. Failures retain interval polling.
  • GUI per-source intervals use a monotonic clock while UTC deadlines remain persisted. Global pause covers startup, every source and residual system triggers. Resume performs one catch-up scan; manual collection still reads every enabled source.
  • Cooperative time limits default to 30 seconds per source and 300 seconds per automatic round. Transient I/O/SQLite locks retry at most three times; waits of 5/15/30 seconds run only when remaining time permits. Interrupted final state can persist and later resume from the cursor. Pause takes effect before settings wait for a write lock; saves run asynchronously. Retry waits and subsequent reads check pause/deadlines. Two source-instance slots run concurrently, including distinct roots of one Agent. Jobs for the same instance merge before parsing; parsing occurs outside the database lock and short writes remain serial. Bounded JSON/JSONL reading/parsing, SQLite VM/backup pagination, authoritative archives, retention and cost transactions cooperate; blocking OS calls cannot be forcibly interrupted. Default file window: 32 MiB; complete lines may commit before continuation. Interrupted/unvisited/merged jobs do not advance deadlines. Fingerprints/generations and events/cursors share a transaction; same-size replacements after failure are detected again. Prior results/revisions remain. See scheduling rules.
  • Telemetry retains explicit attribute keys only. Real HTTP checks verify bad gzip returns 400, oversized bodies/decompression bombs 413, global 120 requests/minute/four active connections, overflow 429; rejected requests are not persisted. Windows Claude/Codex logs and isolated CodeBuddy CLI 2.98.0 traces use per-source tokens/current-user system credentials. CodeBuddy requires the npm manifest beside the first PATH launcher and fixed-release generic headers with %20 encoding; other versions receive no inferred support. Previews show placeholders; background checks are read-only. Authentication precedes continued body reading; duplicate headers/wrong paths/Origin/forwarding headers are rejected. Failure cleanup, revocation, source isolation and reading system storage after restart are implemented. Real exporters, other platform stores, sampling and cross-format association remain incomplete; V22/V25 are unfinished. See authentication requirements.

Checks executed

Query-specific cargo test –offline –manifest-path desktop/src-tauri/Cargo.toml -p llm-usage-core –test query_cache exited 0: seven checks. Receiver-specific cargo test –offline –manifest-path desktop/src-tauri/Cargo.toml -p llm-usage-desktop otel_receiver::tests passed with the full suite: 17 checks, including real HTTP limits, compression and authentication. Pause checks: seven Codex regression cases and one pause-intent case, exit 0. Pause after transient failure avoids the next retry wait and adds no usage/cursor; resume reads that source successfully. Multiple pending pause-save requests take effect independently while a database write lock is held; failure/cancellation releases only its request.

Nine focused checks cover parallelism, bounded reads/SQLite/maintenance rollback and same-size replacement recovery after failure. Full regressions also cover interrupted paginated backup, ZCode authoritative-archive recovery, Codex continuation and Kilo lock-wait deadlines. npm run test:receiver exited 0: eight checks using actual release/IPC/HTTP/Windows credentials. Port conflicts retain original configuration/intent; three independent credentials apply and authenticate. CodeBuddy protobuf reaches isolated traces only, excluding logs/main spans. Revoking one leaves others usable; after revoking all, old tokens fail and original configuration returns. Owned credential remnants: zero. Installation/exporter inputs are synthetic; no Agents or model calls were started and real user configuration was untouched. This does not verify real product export. cargo test … -p llm-usage-desktop windows_credential_store_roundtrip – –ignored exited 0: one native system-store create/read/path-restriction/exact-delete check. Preview redaction/failure cleanup and HTTP missing/duplicate headers/wrong routes/forwarding/ revocation checks also passed.

Full npm run verify exited 0: Rust 889, frontend 20, scripts three, no type errors/warnings, fmt, Clippy -D warnings and frontend build passed. Six explicit/environment tests remain ignored by default. npm run build:desktop and npm run test:headless exited 0; headless: 11 checks, three events/75 tokens. npm run test:browser exited 0: Edge with simulated IPC checked grouped hour charts/tooltips/currency cards/narrow details/partial-data pie/configuration batch apply, retry/revoke/ten languages/themes/timezones/filters/pagination. Final NSIS: 3,901,754 bytes (3.72 MiB); executable: 9,856,000 bytes; frontend JS/CSS gzip: 363.77/9.50 kB. Built, not installed.

npm run test:desktop – –runs 20 exited 0: 17 actual WebView2/IPC checks; small-database first-screen P95 over 20 runs: 770.09 ms. Windows UI Automation checked control names at native DPI 144 (150%); native WM_CLOSE verified tray hiding and disabling the option restored the window. The script’s –force-renderer-accessibility flag is not a product default and does not verify Narrator/NVDA. Real file notifications imported appended records despite a global one-day interval; pause prevented automatic collection while manual refresh still read them. An ordinary-user OS minute task independently verified new usage before GUI startup.

node desktop/tests/native-incremental.mjs –runs 20 exited 0: append 1,000 records per round to a million-row database until visible card update; 20-round P95 717.98 ms, final 1,020,001 events. Token totals/cache invalidation/no new usage on repeat passed. An isolated copy disabled retention to prevent day rollover cleanup changing fixed benchmark counts; the original million-row baseline was untouched.

Queries and import

Run release bench_v20 <synthetic-directory> <1000000/10000000> –query-only –filtered from the root. Add –prepare-indexes for old databases; this mode verifies a nonempty synthetic/bench source before writing and rejects real databases. Benchmark: 366 days, 50 models, 20 Agents; 20 runs per query clear application-connection caches, without clearing OS file cache. Final measurements ran serially after full checks/build.

Query P95 / rows One million Ten million
No dimension filter 47.05 ms 133.86 ms
Model 35.83 ms 87.95 ms
Agent 50.02 ms 196.62 ms
Provider 52.83 ms 163.33 ms
Model + Agent 24.60 ms 145.96 ms
Application-cache hit 0.024 ms 0.026 ms
200-row detail page 1.15 ms 1.16 ms

These queries meet 200 ms on this development machine; the ten-million-row Agent case has little margin. They do not verify every filter combination/hour/week/month/proposed hardware. Initial old-database preparation took about 8.56/49.00 seconds including indexes/derived tables; subsequent preparation/planning checks: 286/2,269 ms. Database sizes 1,079.4/9,461.8 MiB, WAL zero; increases of 173.1/1,427.7 MiB over 906.3/8,034.1 MiB retain the indexing cost. Initial upgrade preparation is excluded from repeated native startup measurements on prepared databases.

node desktop/tests/bench-import.mjs exited 0. Current release core commit_batch imported 1,000,000 normalized events in 50,000-row batches: 184.5 seconds, 5,421 rows/second. All child processes sampled every 500 ms: private-bytes peak 53.45 MiB, working set 62.96 MiB. No WebView was started; bounded normalized writes do not verify all native formats or full GUI import peaks.

Full concurrent regression found early Winsock closure could lose rejected-connection responses. The fix sends the closing sequence, then discards remaining input within cumulative 100 ms/64 KiB, without parsing/persisting it. JSON responses consistently use CRLF/Content-Length. Closure reference: Microsoft Winsock.

Native resources and memory targets

node desktop/tests/native-scale.mjs –runs 20 –idle-seconds 600 –ui-cancel exited 0. Prepared million-event database: 20-run first-screen P95 1,260.72 ms. Two actual IPC cancellations and settings-button cancellation preserve events/revisions/cursors. Default GPU rendering sampled 601.19 seconds, 118 samples, at most seven processes including the executable and all application WebView children.

Metric Measured Target status
Whole-process private bytes, mean / peak 299.90 / 363.30 MiB Old 180 MiB unmet; new mean 350 / peak 400 MiB limits met on development machine
Whole-process working-set peak 525.79 MiB Separate metric; does not replace private bytes
Idle CPU, one logical core 0.244% Below 1%
First-screen P95 1,260.72 ms Below two seconds; excludes initial old-database preparation
Application scheduling-loop count 1,202 About twice/second; not all OS wakes
Process role Mean private bytes Role peak
Executable 17.01 MiB 17.49 MiB
WebView browser 43.38 MiB 45.38 MiB
WebView GPU 146.24 MiB 187.21 MiB
WebView renderer 71.13 MiB 92.04 MiB
WebView utility 19.21 MiB 19.79 MiB
Other WebView 2.93 MiB 3.37 MiB

Role peaks may occur in different samples and cannot be added into a whole-process peak. Each sample asserts role sum equals whole-process usage; raw command lines/paths are not saved. GPU accounts for about 48.8% of mean usage, the largest role locally. This measurement cannot separate runtime/driver overhead from page overhead. An earlier –disable-gpu diagnostic had one startup/31 seconds, mean 183.48/peak 185.54 MiB, still above target; it does not replace default configuration/20 startups/full-duration acceptance. Microsoft WebView2 documents these flags for testing/debugging, without justifying product-default changes.

On 2026-10-04 the user authorized adjusting resource limits using current features/measurements. The application has one window, full overview, SVG charts, SQLite and bounded query caches; no concurrent multiple WebViews or bundled Node/Python services. Forty-one adapters/postprocessing do not imply loading every source into idle memory. The executable uses about 17 MiB; GPU/browser/ renderer dominate. A hard 180 MiB limit does not fit the current Windows GUI. Microsoft performance guidance describes multiprocess/GPU-driver/buffer/content overhead and recommends retaining hardware acceleration, without setting a universal minimum. The local GPU footprint cannot all be labeled unoptimizable fixed overhead. Revised engineering limits: ten continuous minutes of whole-process private bytes mean <=350 MiB, sampled peak <=400 MiB, allowing about 17%/10% above measured values; the current machine meets them. Initial GUI import has a separate <=512 MiB peak. Working set, headless and tray-background results stay separate. Original measurements/old-target failures are retained. native-scale.mjs –enforce-budget checks mean/peak/full duration; software rendering or warmed blank pages do not verify product limits. Proposed four-core/16 GiB hardware, fixed runtime/driver/DPI, sustained growth and other platforms need rechecks. Relaxed limits do not waive leak/unbounded-loading investigation. The same raw samples show first-120-second mean 309.06 MiB and last-120-second mean 299.93 MiB: no sustained growth observed in ten minutes, without proving long-term absence of leaks. New assessment: build/plan-continuation/memory-budget.json; original samples remain unchanged.

The million-row first GUI import separately uses npm run test:import: empty database, 100 isolated Codex roots, 10,000 nonempty synthetic events per root, default product retention, real button through visible cards/charts, independent 500 ms process sampling. The initial 600-second limit expired at 950,000 committed rows/97 registered sources while work was still running; exit 1, not accepted as a peak check. Failure: build/plan-continuation/native-first-import/1791123652310/failure.json. After confirming ongoing commits, the complete observation window was extended to 900 seconds. This changes observation duration, preserving product per-source time limits/two-second incremental target/512 MiB peak. Script initialization also fixed retention parameters, Windows inline sampling, reloading UI language after backend save, and waiting for the empty-database refresh button. A Chinese number abbreviation misread after one million commits was a test-script failure, not product usage failure; its record remains.

Final npm run test:import exited 0: empty database to visible cards/charts in 668.55 seconds. Independent expectations of 1,000,000 calls/15,000,000 tokens match actual summary; repeat scan still one million rows. Continuous sampling: 669.86 seconds/1,031 samples/configured 500 ms interval including overhead/at most eight application processes. Private bytes mean 303.42/peak 371.09 MiB meet the 512 MiB import limit; working-set peak 586.97 MiB is separate. Executable private-bytes peak 72.57 MiB, GPU 165.85 MiB, renderer 80.96 MiB; role peaks are not added. No script exceptions or observed external page HTTP; this does not establish OS-wide process outbound behavior. Result: build/plan-continuation/native-first-import/1791125560874/result.json. Synthetic Codex JSONL with 100 input roots/default retention does not bypass parsing like normalized bulk writes. Their throughput is not directly comparable and does not verify other real Agents/versions/very large native files.

node desktop/tests/native-scale.mjs –runs 1 –idle-seconds 600 –blank-page also exited 0. After warming the full page on the same million-row database, it navigated to about:blank and requested GC with default GPU. Sampling 600.36 seconds/118 samples/at most seven processes: private bytes mean 270.63/peak 283.63 MiB; GPU mean 148.12, renderer 41.81, executable 16.26 MiB; idle CPU 0.083% of one core. Application scheduling counts cannot be read after navigation and are explicitly null. Usage still exceeds 180 MiB, showing a high warm-WebView GPU footprint; it does not establish minimum cold blank-page overhead or proposed hardware performance, and blank pages do not count as product acceptance. Fresh minimal pages/proposed-machine native rechecks remain needed. Result: build/plan-finalization/native-scale/1791121549201/result.json.

Native/browser checks observed no external page HTTP/script exceptions. Only observed WebView requests are covered; these are not whole-process outbound audits. Final results: build/plan-finalization/native-scale/1791109899513/result.json; other native/incremental results: build/plan-completion/native/1791123291549/ and build/query-next/native-incremental/1791121379123/. Final build/full-check logs: build/plan-continuation/build-final.log, verify-final.log; receiver authentication: build/plan-continuation/native-receiver/1791123187714/result.json.

Local source conditions

Read-only actual discover through 41 built-in adapters used known roots, outputting file counts and bounded format/version references without paths/session identities/bodies. Read-only Zed SQLite COUNT found zero threads; an empty DB does not verify usage. Thirty absent default roots: claude, gemini, qwen, cline, dsh, hermes, openclaw, opencode, mimo-code, zoo, aider, junie, xum, droid, amp, grok, roo, goose, crush, jcode, workbuddy, gajae-code, commandcode, continue, atomcode, kiro, antigravity, qoder, copilot, otel. Existing real manual WorkBuddy checks remain valid; this probe describes discovery/read conditions. Highest database versions, readable formats and directories do not verify per-call usage.

No Agents/model requests/WSL/containers were started; real IDE settings were not changed to manufacture samples. Nonempty real samples, other native platforms, proposed hardware baseline, installation/upgrade/uninstall/signing/publication remain unverified. Installation was skipped under earlier instructions; no push/deployment/publication ran. F1/full detail Merge/spending reminders remain deferred. Trace SQLite still needs fixed full schema/attributes and nonempty samples. Fixed GitHub/raw pages could not be retrieved; independent curl timed out at 30 seconds. Table names did not justify attribute/incremental/source-selection inference; earlier C04 documentation references do not establish complete implementation/real acceptance here.

Closing checks: npm run lint:md exited 0, 180 Markdown files; 180 local paths/anchors in 19 changed documents had zero errors. python -X utf8 …/quick_validate.py .agents/skills/ai-maintenance exited 0, Skill is valid. Description unchanged; no model-routing evaluation claimed. git diff –check exited 0. Read-only inspection of owned build namespaces found zero application/ WebView processes/scheduled tasks. User objects were untouched; git status had no temporary artifacts; other tasks’ wording edits were preserved.