Design: /usage command — usage breakdown by model¶
Issue: https://github.com/mochow13/keen-code/issues/99
Implement a /usage command that shows a table of token usage per model, with
←/→ to switch between All time / Last 7 days / Last 30 days and Esc to close.
Requirements (from the issue)¶
- Each used model has a row.
- Each row shows input and output token counts plus KV cache read and write token counts.
- Rows sorted by the most-used model in terms of input count.
←/→cycle between past 7 days, past 30 days, and all time, so usage must be stored to support all three windows.Escto go back.
Decisions (confirmed)¶
- Scope: record every provider call — main turn, manual
/compact, auto-compaction,/btw,/adversary, and subagents. - Storage: new global append-only ledger at
~/.keen/usage/usage.jsonl. Session history is not used: it is namespaced per working directory and pruned after 14 days (internal/cleanup/cleanup.go:19), which breaks "all time" and cross-project totals. - Retention / rollup: raw per-call records kept for 30 days; older records are
rolled up to one record per (UTC day, provider, model). Compaction runs
only when
/usageis opened (throttled to at most once per day) — not in/cleanup. - Cache read/write: currently merged into a single
CachedTokens; split into read and write so cache read and cache write can be shown separately. OpenAI-family providers only report cache reads, so their cache-write column renders-.
Current state (findings)¶
core.TokenUsage(internal/llm/core/message.go:134) has a singleCachedTokensfield:Input, Output, Total, Reasoning, Cached.- Providers collapse read+write into it:
- Anthropic (
internal/llm/anthropic.go:349):CacheCreationInputTokens + CacheReadInputTokens. - Bedrock (
internal/llm/bedrock.go:642):CacheReadInputTokens + CacheWriteInputTokens. - OpenAI (
internal/llm/openai.go:567), Responses (internal/llm/openai_responses.go:322), Codex (internal/llm/openai_codex.go:171): onlycached_tokens(read). - Genkit (
internal/llm/genkit.go:305): no cache fields. - One
usageevent is emitted per provider response (core.StreamEventTypeUsage), not per turn. A multi-iteration tool loop emits several, so recording per event is the most accurate for billing and survives interrupted turns. - Usage event sinks today: main turn →
llmUsageMsg→handleLLMUsage(internal/cli/repl/handlers.go:37); headlesskeen run(internal/cli/repl/headless_run.go:145); auto-compaction usage is oncore.AutoCompactionEvent.Usage(message.go:145)./btw(StreamBtw) and/adversary(StreamAdversary) ignoreStreamEventTypeUsage. Subagents stream throughcollectResult(internal/subagents/activity.go:18). - Modal widget precedent to copy: session picker
(
internal/cli/repl/widgets/session_picker.go), wiring athandlers.go:535(handleKeyMsg),repl.go:547(render),repl.go:1012(clear). - Existing helpers to reuse:
formatCompactTokens(internal/cli/repl/context_status.go:118),addCommandTable/maxColumnWidth(command_handlers.go),ModelSelection*styles andRuleStyle(internal/cli/repl/theme/styles.go).
Design¶
1. Split cache read/write in the token model¶
Add CacheReadTokens and CacheWriteTokens to core.TokenUsage; keep
CachedTokens = CacheReadTokens + CacheWriteTokens so existing context-status math
is unchanged.
| Provider | read | write |
|---|---|---|
| Anthropic | CacheReadInputTokens |
CacheCreationInputTokens |
| Bedrock | CacheReadInputTokens |
CacheWriteInputTokens |
| OpenAI / Responses / Codex | prompt_tokens_details.cached_tokens |
0 |
| Genkit | 0 | 0 |
Also update cloneHeadlessUsage (headless_run.go:316) to carry the new fields.
2. Ledger (internal/usage)¶
Append-only JSONL at ~/.keen/usage/usage.jsonl, one record per provider response:
{"ts":"2026-09-10T18:22:04Z","provider":"anthropic","model":"claude-sonnet-4-5","input":18432,"output":512,"cache_read":16000,"cache_write":1024,"reasoning":0}
Rolled-up record (one per UTC day + provider + model), date instead of ts:
{"date":"2026-08-14","provider":"anthropic","model":"claude-sonnet-4-5","input":90210,"output":4410,"cache_read":77120,"cache_write":3072,"reasoning":0,"rollup":true}
API surface:
type Record struct { TS time.Time; Provider, Model string; Input, Output, CacheRead, CacheWrite, Reasoning int; Rollup bool }withDate stringused for rollups (mutually exclusive withTS).Store{path string}:Append(rec Record) error—O_APPEND|O_CREATE|O_WRONLY, single line write,mkdir -pparent. Errors are logged, never fatal to a turn.Load() ([]Record, error)— tolerant of corrupt/truncated lines (skip and continue).Compact(now time.Time) error— see below.Summarize(records []Record, since time.Time) Summary— group byprovider/model, sum fields, sort by input tokens desc, append aTotalrow.since.IsZero()means all time. Reads raw and rolled-up records identically.type Range intwithRangeAllTime,RangeLast7Days,RangeLast30DaysandrangeSince(Range, now) time.Time.type ModelUsage struct { Provider, Model string; Input, Output, CacheRead, CacheWrite, Reasoning int }andtype Summary struct { Rows []ModelUsage; Total ModelUsage }.
Rollup compaction. Compact:
- Cheap lock via
O_CREATE|O_EXCLonusage.lock, with stale-lock recovery by mtime (e.g. > 5 min old → remove and retry). No flock dependency, so it stays portable. - Load all records. Split into
old(ts/datestrictly beforenow - 30d) andkeep. Rollups are day-granular, so a rollup day is always strictly outside the 30-day window and never contaminates the 7d/30d filters. - Group
oldby (UTC day, provider, model) and merge with any existing rollup records inkeepfor the same key (idempotent re-runs). - Rewrite the file atomically (temp file + rename) with rollups first, then
keep, preserving append-only semantics for future writes. - Remove the lock.
Throttle: a marker file ~/.keen/usage/.compacted holding the last compaction date;
MaybeCompact(now) runs at most once per calendar day and is called from the
/usage open path only.
Bounded size ≈ models × days rollup lines + 30 days of raw turns.
3. Capture points¶
Add one repl helper:
func (m *replModel) recordUsage(provider, model string, u *core.TokenUsage)
which is a no-op on nil usage and appends to the ledger. Each site becomes a one-liner:
- Main turn —
handleLLMUsage(handlers.go:37); provider/model fromm.ctx.cfg. - Headless
keen run—headless_run.go:145; same cfg. - Manual
/compact— usage events already flow throughllmUsageMsg; covered by the main-turn path. - Auto-compaction — handle
AutoCompactionEvent.UsageinllmAutoCompactionAppliedMsg. /btw— add aStreamEventTypeUsagecase inhandleBtwStreamMsg(m.ctx.cfg)./adversary— add aStreamEventTypeUsagecase inhandleAdversaryStreamMsg(adversary cfg provider/model).- Subagents — add
Usage chan<- usage.Recordtosubagents.Runnernext to the existingActivitychannel (internal/subagents/activity.go:33); emit fromcollectResultonStreamEventTypeUsageusing the profile's own provider/model fromRunner.resolvedConfig. The repl drains it likesubagentActivity.
4. UI¶
Modal overlay modeled on the session picker. Load the ledger once on open, precompute all
three windows, then ←/→ just swap an index.
────────────────────────────────────────────────────────────────────────
Usage Data
All time Last 7 days Last 30 days
────────────────────────────────────────────────────────────────────────
Model Input Output Cache read Cache write
claude-sonnet-4-5 1.24M 342k 980k 120k
gpt-5 412k 88k 210k -
gemini-2.5-pro 34k 9k - -
────────────────────────────────────────────────────────────────────────
Total 1.69M 439k 1.19M 120k
←/→ change window Esc close
────────────────────────────────────────────────────────────────────────
- Active tab bold + primary color, inactive tabs muted/faint; default All time;
←/→wrap around. - Rows sorted by input tokens desc (per the issue).
Totalrow bold, same styling approach as/context(context_status.go:188). - Numbers compacted with
formatCompactTokens;-for zero/unsupported cache columns. - Empty state:
No usage recorded yet.Footer hint reflects the active window. - Same card chrome (dim rule lines,
ModelSelection*styles) as/sessionsand/modelso it reads as one family.
5. Wiring¶
internal/cli/repl/widgets/usage_view.go+ test —UsageView{rangeIndex int, summaries [3]usage.Summary}withNextRange(),PrevRange(),CurrentSummary(),FormatUsageCard(view, width).internal/cli/repl/commands/commands.go— addUsage = "/usage"toAllandSuggestions.internal/cli/repl/command_handlers.go—case input == replcommands.Usage:callsusage.MaybeCompact(now)thenm.startUsageView().internal/cli/repl/handlers.go— gate keys inhandleKeyMsg(:535):left/rightcycle the window,esccloses; addkeyLeft/keyRightconstants.internal/cli/repl/repl.go— render inupdateViewportContent(:547); clear inhandleClearCommand(:1012).internal/cli/repl/repl_helpers.go—formatUsageCardwrapper if needed.docs/cli-usage.md— command table row,/usagesection with key table and a note on the ledger + 30-day rollup.
Note: /cleanup (internal/cleanup/cleanup.go) is intentionally not changed.
File-by-file¶
| File | Change |
|---|---|
internal/llm/core/message.go |
CacheReadTokens / CacheWriteTokens fields |
internal/llm/{anthropic,bedrock,openai,openai_responses,openai_codex,genkit}.go |
populate read/write |
internal/llm/auto_compaction.go |
pass usage through unchanged (verify) |
internal/usage/store.go |
Record, Store.Append/Load/MaybeCompact/Compact, lock |
internal/usage/summary.go |
Range, Summarize, ModelUsage, Summary |
internal/usage/*_test.go |
new tests |
internal/subagents/{activity.go,runner.go} |
usage channel + emit |
internal/cli/repl/handlers.go |
recordUsage, main/btw/adversary/auto-compaction capture, key gating |
internal/cli/repl/headless_run.go |
record usage + carry new cache fields |
internal/cli/repl/widgets/usage_view.go (+test) |
card renderer |
internal/cli/repl/commands/commands.go |
register /usage |
internal/cli/repl/command_handlers.go |
dispatch + startUsageView |
internal/cli/repl/repl.go |
render + clear |
docs/cli-usage.md |
docs |
Tests¶
internal/usage: append/load roundtrip; corrupt-line tolerance;Summarizeordering and window filtering;Compactidempotency; the 30-day boundary (raw ↔ rollup); lock contention / stale-lock recovery;MaybeCompactonce-per-day throttle.- Providers: each maps cache read/write correctly into
TokenUsage. - Widget: table render per window, tab wrapping, empty state.
- Keys:
←/→cycle the window,Esccloses, unrelated keys are swallowed. - Capture: main turn,
/btw,/adversary, auto-compaction, subagent each append a record.
Open items¶
- Zero/unsupported cache columns render
-. - Cache write is only reported by Anthropic/Bedrock; OpenAI-family rows show
-there.