Cost Management & Budget Control
Myrm provides a comprehensive multi-layer cost control system that ensures you never overspend on LLM usage — while maintaining task quality through intelligent routing.Budget Policy
Configuration (Settings → System → Budget Policy)
Set spending limits across three independent dimensions:
Each dimension triggers a 4-level progressive response:
How It Works
- WARNING — Agent receives a budget hint via cache-safe prompt injection. Eco mode activates (compression threshold lowered 20%)
- FINALIZATION — Agent is instructed to output final results immediately. Tool calls are stripped from model output
- EXCEEDED — Hard block. No further LLM calls until budget is reset or expanded
BudgetBadge and Goal status card.
Real-Time Cost Observability
Per-Message Display
Every assistant message shows detailed token breakdown:- 7 token types: prompt, completion, cached, cache_write, reasoning, citation, total
- Cost in USD with cache savings percentage
- Model tier indicator (simple / standard / reasoning)
- Cache break reason attribution when cache misses occur
Goal Panel (Active Goals)
While a goal is running, the expanded status card shows:- Burn rate — tokens/minute consumption
- Cost rate — $/minute spending
- ETA — estimated time to budget exhaustion
- Progress bar — percentage of budget consumed
Usage Statistics Dashboard (Settings → System)
Full analytics dashboard with:- Daily spending chart (7/30/365 day ranges)
- Session-level drill-down with token/cost breakdown
- Model breakdown panel (cost per model)
- Hour and weekday distribution heatmaps
- Per-agent usage cards with sparkline trends
- Routing analytics (how many requests used each tier)
Full-Chain Accounting — No Blind Spots
Every LLM call, streaming or not, lands in the same token ledger:- Unified recording — streaming and non-streaming calls are counted in one place (the adapter layer), so sub-agent spawns, dynamic workflows, background summaries, memory extraction, context compression and wiki compilation are all attributed correctly
- Honest failure paths — a failed LLM call never fabricates usage, so the ledger never inflates
- End-of-run accuracy —
message_endevents are back-filled with real usage from the tracker, so per-run totals match the sum of individual calls
Intelligent Cost Routing
ComplexityRouter (Automatic)
Myrm automatically routes requests to the most cost-effective model tier:- SIMPLE — Quick questions, greetings → cheapest model
- STANDARD — Normal tasks → balanced model
- REASONING — Complex multi-step problems → premium model
- Two-phase classification (rule scoring + LLM judge)
- Session momentum — follow-up messages inherit the conversation’s recent tier to prevent quality degradation
- Penalty feedback — records misroutes and adjusts future routing probability (24h decay)
- Min-tier floor — regenerate requests automatically escalate tier
- Zero configuration — active by default on all accounts
Privacy-Aware Routing
Sensitive content is automatically routed to privacy-appropriate models based on PII detection level.Subagent Cost Isolation
Delegation Budget
Each root task has amax_descendants = 20 guard preventing runaway subagent spawning. This ensures a single task cannot create an unbounded delegation tree that drains budget.
Per-Root Budget Scope
Budget tracking is scoped per root agent run — subagent costs accumulate against the parent’s budget dimensions, providing unified visibility.Multi-Agent Cost Pre-Flight
When the agent needs to spawn multiple sub-agents (batch delegation, council cross-review, or alternatives parallel generation), Myrm estimates the total LLM cost before execution begins:- Batch mode — estimates cost based on task count × per-task model pricing
- Council mode — accounts for
expert_count × (1 + cross_review_rounds) + 1(chair synthesis) LLM calls - Alternatives mode — estimates based on the number of parallel expert agents
- Total estimated cost in USD
- Number of LLM calls to be made
- Remaining budget
Long Task Resilience
Checkpoint & Recovery
If the server process is interrupted during a long task:- Signal-safe checkpoint — all running subagents are checkpointed to disk (JSON + fcntl file locking)
- Orphan recovery — on restart, orphan scanner detects interrupted checkpoints and publishes events to the UI
- Resume API — user can resume from checkpoint with full state restoration (messages + workspace)
Goal Continuation Guard
The continuation engine evaluates 9+ conditions before each turn:- Budget remaining check
- Cancellation / steering token
- Convergence detection
- Goal drift detection (trajectory scoring every 5 turns)
- Per-todo checkpoint (optional pause after each step)
- Protected file integrity validation
Output Compression
CLI tool output is automatically compressed using intelligent rules to minimize token consumption:- Path deduplication and summarization
- Repeated pattern collapsing
- Binary/noise filtering
- Large output truncation with summary
Stable Prefix for Cache Efficiency
System prompts and tool schemas are carefully ordered and canonicalized to maximize prompt cache hit rates:- Deterministic tool schema ordering (SSOT)
- Schema canonicalization (removing non-semantic differences)
- Cache breakpoint detection with attribution reporting