Local AI usage desktop design
Current implementation and remaining acceptance work are in Plan.md. This directory maintains the current design. Public documentation, fixed-version source and actual implementation provide distinct references; sample, native desktop and CI results are recorded separately. See the adapter matrix for capabilities/limits and implementation prerequisites for the authorized read-only scope. Telemetry setup, dashboard interactions and pricing define their respective behavior. Local HTTP source authentication, credentials, failure recovery and revocation are in receiver authentication.
Product purpose and scope
Show observed local usage by Agent and model for today and historical periods: where consumption occurs, how cache usage changes, and which sources have incomplete collection. Statistics use local data, requiring no account login, cloud synchronization, external database or additional language-model calls.
Local means logs, databases, usage files and explicitly enabled telemetry produced by Agent instances on this computer. Local records of calls to cloud models qualify; models need not run offline. Local WSL/container instances require explicit roots, instance or distribution identity and deduplication; do not discover or start these environments automatically. Exclude SSH/remote gateways, network shares, cloud bills and enterprise/cross-device account reports. Downloading remote exports, syncing remote sessions or using a cloud-synced directory does not establish local origin. Local session exports must retain verifiable local attribution and native detail. Exclude unattributed records from totals and explain why. Missing data does not mean zero usage; source logs do not establish that all model requests were recorded.
Windows 11 x64 is the first desktop target; GitHub CI retains Windows/macOS/Linux build and test jobs. Local WSL 2 Linux builds, WSLg smoke tests, CI builds and real desktop acceptance have separate results. Harness Agent was identified as Hermes Agent, with A24 references. The user authorized read-only extraction of minimal, redacted native Agent data during implementation. This round also authorized F1 IDE investigation; missing installations/local formats leave active work, while limits remain in the matrix. Built-in Zed and Junie CLI belong to M8. Zed 1.22.0 has native two-model/cache samples for the specified external provider, verifying neither hosted service nor other providers. Extensions/telemetry with local field references continue in their assigned stages; see the matrix.
Main decisions
| Decision | Purpose and constraints |
|---|---|
| Tauri 2 with a Rust backend | System WebView; local backend owns file/database access. Measure package size, all-process resources and platform behavior in the actual environment. |
| SQLite, one writer and WAL | Local transactions, unique constraints, revisions and summaries without a database service. Retention/recovery never modifies source databases. |
| Svelte/TypeScript, selective ECharts imports and SVG | Filters, settings and charts. Artifacts/resources have measurements; no pure-TypeScript comparison was run, so no performance difference is claimed. |
| Local history, explicit telemetry and local session exports | Verify each source’s version, fields, lifecycle and local origin independently. Unverified IDE formats move to F1; no remote usage API. |
| Separate events, cumulative values, interval reports and quotas | Message rows, histogram sample counts and credits are not requests or tokens. Check samples per field. |
| Stable host/source identity and versioned exchange | Hostname is descriptive. Aggregate and complete normalized-detail packages import separately, retaining revisions, conflicts and archived partitions. |
| Built-in compiled adapters with version directories | Ship through application upgrades and retain historical implementations. Unknown versions try compatibility reading and keep unverified-version markers after successful validation. |
See architecture for alternatives and research for source references.
Pages and interactions
Software updates defines automatic/manual checks, optional downloads, global progress, package identity and installation/portable replacement with recovery.
M6 implements one main window and five navigation pages: Overview, Trends, Sources, Details and Settings. Overview/Trends panels can be hidden and dragged into order. Automatic collection has one global interval, initially one hour; 0 disables it. Each source can be enabled/disabled and use an interval/daily/weekly schedule. Login startup and Windows background tasks default off. System tasks check saved due rules every minute, reporting requested and actual state separately. Disabling automatic collection prevents scans. Settings explains collection with a closed window, exited process, sleeping computer or signed-out user; see scheduling.
| Page | Contents | Operations and failure feedback |
|---|---|---|
| Overview | Today summary, hourly charts, model/Agent distributions and tables; historical calls/sessions and token charts | Today, last 2 calendar days, 7/30/365 days; separate today/history drag selections update summaries/distributions. Show collection progress and failure time. |
| Trends | Calls/sessions, daily/weekly/monthly/hourly tokens, API price references, model/Agent distributions and model table; activity calendar and weekday distribution last | Click/drag selects a shared range for summaries, distributions and model table. Reset clears selection. Calendar/weekday panels show the full range. |
| Sources | Discovery/enabled state, versions, local origin, fields, last/next collection, errors and assigned user | Add roots, enable/disable, collect now, rescan, remove, reassign users and switch user; explain limitations. |
| Details | Paginated usage: time, Agent, model, category, input/cache read/output/total, duration and session; no prompts, responses or tool output | Agent/model filters, previous/next page and page jump. |
| Settings | Timezone/language/week start/interval; tiered retention, per-tier statistics and manual/all-data cleanup; startup/background; host alias; export/import | Preview schedules/configuration effects. Explicitly enable/disable Windows tasks and report actual state on failure. |
Charts include calls/sessions; total tokens, stacked cached/uncached input, output and weighted cache-input ratio; model/Agent pies with tables; today’s hourly usage; activity calendar and weekday distribution. Cost/error/latency charts depend on source capabilities. Requests and tokens do not share one numeric axis; ratios are not stacked with counts. Every chart provides a readable table, keyboard access, both themes, color-accessible palettes and explicit units. Icons/static assets are specified in asset guidance, with local previews. UI/tray/empty-state interactions have separate M6 desktop checks.
The implemented M6 layout illustrates information groups; panels remain hideable/reorderable:
[Overview] [Trends] [Sources] [Details] [Settings] [Time range] [Collect & refresh] [User]Input (including cache) | Output | Total tokens | Cached | Uncached | Cache-input ratioToday's hourly chart, model/Agent pies and tables (Overview: today)Calls/sessions, tokens, activity calendar, weekday distribution (Overview: history/Trends)Details: time / Agent / model / category / input / cache read / output / total / duration / sessionAdditional statistics
| Statistic | Purpose | Data requirements and limitations | Stage |
|---|---|---|---|
| Cache read/write and weighted hit ratio | Explain high input usage and cache reuse | Cache creation belongs to total input; unknown fields remain unknown | First version |
| Per-call token average, P50/P95 and large requests | Identify large context/high-consumption calls | Complete per-request data only; exclude session cumulative records | First version, by capability |
| Project/workspace and primary/sub-agent/auxiliary shares | Identify title generation, compaction and background usage | Project grouping defaults off; parent totals and child events must not double count | First version, by capability |
| Retries/errors/cancellation and known tokens on failed calls | Explain consumption growth and reliability changes | Request lifecycle/telemetry required; success-only logs cannot imply 100% success | M5 |
| TTFT, duration P50/P95 and output speed | Describe waiting and service changes | TTFT needs first-token time; speed uses the actual generation interval, never file times | M5 |
| Estimated/reported costs and locally attributable credits | Manage usage and compare models/cache | Local usage with versioned prices; ambiguous channels receive no price; estimates are not bills | F2; estimation/online refresh default off |
| Period comparisons, usage/cost alerts and trend anomalies | Identify sustained increases or daily spikes | Today compares equal elapsed time yesterday; daily/monthly token or single-currency estimate alerts default off; explain incomplete coverage | M6/F2; alert rules |
| Freshness, parse errors and field completeness | Explain and inspect totals | Coverage describes discovered sources/observed records, never estimated unrecorded calls | First version |
| Sessions, active days and activity calendar | Review usage patterns | DISTINCT session identities; active dates are not work duration or productivity | First version |
Do not derive code quality, saved hours, personal performance or model rankings from tokens. Carbon, GPU time and true internal cache-hit probability lack required data and have no product metrics. The UI calls cache_input_ratio “Cache hit rate”; its formula remains cached input as a fraction of input. See metric formulas.
Requirement mapping
| User requirement | Design reference | Acceptance checks |
|---|---|---|
| Refer to previous-draft | Prototype review, retaining useful groupings/chart interactions | Preserve prototype; audit historical migration separately |
| Model statistics and summaries | Data rules, separating model/provider/Agent | V01–V06, M1/M6 |
| Persist host/source history and support export/import/Merge decisions | Provenance, detail merge | V28, M1a/M6; aggregate and complete normalized-detail exchange implemented |
| API price snapshots and token estimates | Cost rules, pricing | V29/F2; check observation-time estimates/current API references separately |
| Languages and localized formats | Localization research | V31; F3 research and M6 implementation |
| Refresh today | Collection | V07–V12, M2/M6 |
| Scheduled collection | Scheduling/lifecycle | V23/V24, M1/M6 |
| Local only; attempt all Agents in stages | This page’s scope and the matrix | V17/V25; per-adapter M2–M5/M8 checks; unverified IDE formats remain F1 |
| Use native local data; defer IDEs; confirm platforms | Implementation prerequisites | Permission/stages clear; preparation passed, execution recorded separately |
| Windows first, three-platform CI and WSL builds | Platform/CI | V26/V27, M0/M7 |
| Windows/Linux Podman installation; no macOS desktop/specific hardware requirement this round | Installation, native source preparation | Lifecycle checks, container sources; host integration checked separately |
| Daily/weekly/monthly summaries | Time rules | V04–V06 |
| Other meaningful statistics | The capability-dependent table above | V01–V06, V18 |
| Retention/grouping settings | Settings rules | V13–V16, M6 |
| Charts | Page specifications above | V18, native desktop checks |
| Domestic/international Agent extensions | Tool matrix, extended coverage | Per-adapter M2–M5/M8 checks; F1 is not mandatory first-version acceptance |
| One directory per Agent, historical versions within it | Adapter layout, migration | V30, M2–M5; reused when F1 implementation begins |
| Unknown versions try the latest parser | Compatibility rules | V17/V30, M2–M5; reused for F1 |
| Small local database | Database decision | V11/V15/V20 |
| Resource-efficient desktop | Resource targets | V19–V21, actual release artifacts |
| Continue implementation after design | Plan.md, implementation scope | Current work/remaining conditions in Plan; original research retains historical scope |
Document responsibilities
Plan.md maintains unfinished work; execution.md defines outputs/order; data-contract.md defines statistics; adapters.md defines supported tools; research.md records sources; validation.md defines stable acceptance criteria. current-acceptance.md indexes current results. Detailed records preserve environment, commands, first failures and historical scope. Update the relevant design before changing behavior. Do not duplicate formulas or contradictory support statements.