Competitor Comparison & Migration Guide
Myrm is designed as a complete AI agent workspace. Unlike terminal-only coding agents, Myrm provides a GUI-first experience with persistent sandboxes, cross-session memory, and enterprise-grade security — all while maintaining full coding capabilities.Migration confidence — You keep control: pick SaaS, self-hosted, or desktop; import Hermes skills via GUI ZIP; compare channels, memory, and security row-by-row below before you commit. For plugin-centric users (for example Grok Build), Myrm installs a full Agent package in one step (model + skills + MCP + subagents + safety profile), instead of making you assemble multiple plugins manually. Docs are available in 6 languages (zh/en/ja/ko/de/zh-TW); the app auto-detects your browser language on first visit (RFC 7231 Accept-Language negotiation) across all deployment modes — WebUI, Cloud-hosted, and Tauri desktop.
preflight vs runtime_fallback) including mobile non-hover copy, plus strict cloud fail-closed ingest (shared envelope idempotency + no silent local fallback when shared dedup backend is unavailable, with dedup_reject reason metrics). Latest focused regression run (2026-07-19): server 28 passed, control-plane 55 passed, frontend 14 passed.
For per-turn capability control, Myrm supports one-turn Skill/MCP override chips with full observability (submit/apply/noop/queue/busy/final outcome + failure reasons). Focused validation (2026-07-29): frontend 29/29 passed, server 10/10 passed; targeted coverage reached 62.47% lines on core frontend files and 72% total on server telemetry modules. This gives teams migration-safe proof for cost/reliability tuning, not guesswork.
Latest Validation Snapshot (2026-08-21 · In-place Interrupted Metacognitive Annotation & Detached Subshell Guard · Error Recovery & Robustness)
- In-place Interrupted Metacognitive Annotation & Detached Subshell Guard (vs Claude Code / OpenHands / Devin / Hermes): Myrm eliminates LLM repetition deadlocks after user cancellation and prevents runaway background subprocesses from deadlocking sandbox ports:
- In-place Metacognitive Annotation: Automatically embeds adaptive guidance into the synthetic ToolMessage upon user cancellation (
Action was cancelled by user; do not retry the exact same failing action. Please adapt your plan to user instructions.), breaking repetition deadlocks while maintaining 100% Prompt Cache prefix stability; - POSIX Streaming Lexer Guard: Detects unquoted intermediate background
&operators in compound shell commands (e.g.,npm run dev & echo started), intercepting orphaned subshells and guiding the model to use managedrun_in_background=Trueexecution; - Full-Stack Test Coverage: 93 harness unit tests + 71 module integration tests + Real LLM API streaming tests + Real Chrome E2E tests 100% green.
- In-place Metacognitive Annotation: Automatically embeds adaptive guidance into the synthetic ToolMessage upon user cancellation (
- Material SSOT:
temp-docs/materials/ERROR_RECOVERY_ROBUSTNESS_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/agent/middlewares/tooling/dangling_tool_call_middleware.py·myrm-agent-harness/src/myrm_agent_harness/agent/meta_tools/bash/_security/preflight_checks.py
Latest Validation Snapshot (2026-08-20 · Protocol-First Production Kanban vs Fragile JSONL Collaboration · Roadmap Topic 16 Done)
- Protocol-First Production Kanban Collaboration (vs MateClaw File-Based Multi-Agent): Myrm delivers an enterprise-grade asynchronous multi-agent collaboration and task tracking architecture:
- Protocol-First Kernel & Deterministic State Machine: Built directly in
myrm-agent-harnesswith atomic transactions, dependency resolution, human-in-the-loop (in_review), hallucination verification (CompletionVerifier), and timeout recovery (BlockKind), replacing brittle JSONL/markdown file polling; - Multi-Channel & Multi-Interface Unification: Full GUI Kanban in WebUI/Desktop alongside seamless
/kanbanslash command execution across Telegram, Slack, Feishu, and WeCom; - Full-Stack Test Coverage: 480 harness unit tests + 253 server/API/command integration tests + Real Chrome MCP E2E tests 100% green.
- Protocol-First Kernel & Deterministic State Machine: Built directly in
- Material SSOT:
temp-docs/materials/MULTICHANNEL_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/toolkits/kanban/·myrm-agent/myrm-agent-server/app/channels/routing/commands/kanban_command.py
Latest Validation Snapshot (2026-08-21 · OpsAggregatedSnapshotApi Full-Spectrum Observability · Roadmap Topic 14 Done)
- OpsAggregatedSnapshotApi Full-Spectrum Snapshot (vs hermesd / Vercel AI Gateway): Myrm delivers a unified single-request operational observability API:
- 8-Domain Concurrent Gathering: Seamlessly aggregates
System,Liveness,Resources(Process RSS + Memory Level),Channels,Governance(24h Cron Failures + Pending Approvals),UsageRadar,Memory, andDoctorSummaryin a single call; - Dual-Modal Light & Deep Modes:
include_doctor=falseenables sub-5ms lightweight fleet polling for Control Plane, reducing network roundtrips by 66%+;include_doctor=trueenables comprehensive diagnostics with a 2.0s hard timeout fuse; - Zero New Storage & Zero Background Polling Overhead: Pure in-memory singleton gathering and isolated session reads with zero extra DB tables and zero idle CPU waste;
- Full-Stack Test Coverage: 100% green across unit, fallback isolation, and FastAPI endpoint tests (
tests/api/ops/test_ops_snapshot.py), zero lint errors.
- 8-Domain Concurrent Gathering: Seamlessly aggregates
- Material SSOT:
temp-docs/materials/OBSERVABILITY_ANALYTICS_ADVANTAGE.md·myrm-agent/myrm-agent-server/app/schemas/ops.py·myrm-agent/myrm-agent-server/app/api/ops/
Latest Validation Snapshot (2026-08-20 · Five-Ring Closed Learning Loop Status Hub & Growth Discovery Chip · Roadmap #55 Done)
- Five-Ring Closed Learning Loop Status Hub & Discovery Chip (vs Hermes / Mem0 / Claude / Manus): Myrm delivers an industry-first visual continuous self-evolution status hub and discovery pill:
- Five-Ring Topology Status Hub (
LearningLoopFiveRingHub): Seamlessly visualizes the full continuous learning cycle across five interlocked rings: Ring 1 (Post-task Reflection & Trace Mining), Ring 2 (Skill Distillation & Proposals), Ring 3 (Runtime Advancement & Improvement Gate), Ring 4 (Periodic Memory Consolidation & Noise Pruning), and Ring 5 (Progressive User Profiling & Cross-Session Recall); - EmptyChat Growth Discovery Chip (
GrowingLoopDiscoveryChip): Subtle, high-polish pill in empty chat state informing users of current consolidated memories and active skills with one-click direct jump to/journey; - Anti-Goodhart Health Scoring: Health score strictly evaluates component readiness, absence of degradation, and gate protection efficacy rather than artificially forcing unnecessary background evolutions;
- Full-Stack Test Coverage: 2 backend schema/API integration tests passed (
test_learning_loop_integration.py) + 0 frontend lint errors.
- Five-Ring Topology Status Hub (
- Material SSOT:
temp-docs/materials/MEMORY_ADVANTAGE.md·myrm-agent/myrm-agent-server/app/api/statistics/learning_loop.py·myrm-agent/myrm-agent-frontend/src/components/features/growth/LearningLoopFiveRingHub.tsx
Latest Validation Snapshot (2026-08-20 · ComputedArtifactTrustGuard & Deliverable Write Verifier · Roadmap Topic 10 Done)
- Computed Artifact Trust & Deliverable Claim Guard (vs Prime-Agent REPL / Manus / HappyCapy): Myrm delivers an enterprise-grade computed artifact reliability and phantom delivery prevention architecture:
- ArtifactVault
vault://Pointer Protocol: Large computed artifacts and subagent outputs are automatically vaulted to disk/storage with structured pointer cards, preventing model context explosion and oral hallucination during multi-step execution; - Deliverable Write Verifier (
deliverable_write_verifier.py): Intercepts answer finalization and strictly blocks completion if the agent claims to have created, written, or modified a workspace file when no successful file write/edit tool call was recorded; - Independent Sandbox Re-Run Verification (
_rerun_verification_in_sandbox): Automatically re-runs verification commands inside the isolated sandbox when workspace files are mutated after earlier checks, ensuring 100% genuine code verification without LLM bias; - Strict Answer-Phase Gating (
answer_user_tool): Gated middleware architecture ensures full pre-answer checks (safety, mutation tracking, external evidence, and deliverable consistency) complete before final answer emission; - Full-Stack Test Coverage: 100% green across harness completion guard unit tests, edge cases, vault storage, and server-side end-to-end integration tests.
- ArtifactVault
- Material SSOT:
temp-docs/materials/AGENT_VERIFICATION_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/agent/middlewares/completion/
Latest Validation Snapshot (2026-08-20 · Declarative High-Fidelity PDF Template Registry & Print Engine · Roadmap Topic 11 #4 Done)
- Declarative PDF Template Registry & High-Fidelity Printing Engine (vs PDFx / Puppeteer Scripts / Markdown Converters): Myrm delivers an enterprise-grade structured PDF generation toolchain:
- Zero Node.js Runtime Lock: Pure Python Jinja2 interpolation coupled with direct Patchright Chromium print pipelines or standalone graceful fallback, completely eliminating external Node.js/npm container bloat;
- Manifest SSOT & Auto JSON Schema Derivation:
PdfTemplateManifestcontract strictly defines variable types, descriptions, default values, and examples, auto-generating JSON Schemas for deterministic LLM tool calling; - High-Fidelity Corporate Layouts & Paged Media CSS: Built-in standard VAT Invoices, Business Reports, and Payment Receipts featuring CJK font support and print pagination guards (
page-break-inside: avoid); - Deep Security Sanitization & SSRF Protection:
PdfRenderEngine.sanitize_htmlstrips malicious<script>tags and neutralizes dangerousfile://,javascript:,gopher://schemes; - Prompt Cache Optimization (
ToolLayer.EXTENDED): Lightweight 3-tool surface (list_pdf_templates,get_pdf_template_schema,render_pdf_template) maintains stable prefix prompt caches with 95%+ cache hit rate; - Full-Stack Test Coverage: 7 unit tests + 5 integration/concurrency stress/adversarial attack tests passed (100% green, 95% test coverage, 0 lint errors).
- Material SSOT:
temp-docs/materials/TOOL_ECOSYSTEM_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/toolkits/pdf_templates/
Latest Validation Snapshot (2026-08-20 · FastUiRestoreMode 0ms Visual Rehydration & Quota-Resilient Snapshotting · Topic 12 Done)
- Dual-Level Navigation Snapshot Bus (vs pi_agent / LobeChat / Single-Layer Storage): Myrm delivers an enterprise-grade instant visual recovery architecture across WebUI and Desktop client reloads:
- L1 (In-Memory LRU 20) + L2 (SessionStorage 3) Snapshot Bus: Instantaneous in-memory context switching combined with session-scoped persistence that rehydrates 0ms visual states upon browser refresh (F5/Cmd+R) or client restart;
- Zero-Lag Visual Rehydration: Pre-renders recent chat viewport state before asynchronous SWR data sync completes, completely eliminating cold-start blank screens and skeleton jump artifacts;
- Lightweight Data Sanitization & Quota Guard: Automatically strips oversize (>1KB) Base64 images and resets transient
loadingstates tofalse, with a 300KB hard limit per snapshot and comprehensivetry/catchfallbacks againstQuotaExceededError; - Multi-Pane & Incognito Isolation: Full integration with
resolvePaneSnapshotBasemulti-pane splits while strictly bypassing L2 persistence forincognitoModesessions to protect privacy; - Full-Stack Test Coverage: 9/9 frontend snapshot unit tests passed + 216 frontend chat store regression tests passed (100% green, zero lint errors).
- Material SSOT:
temp-docs/materials/DESKTOP_APP_ADVANTAGE.md·myrm-agent/myrm-agent-frontend/src/store/chat/chatNavigationSnapshotCache.ts·myrm-agent/myrm-agent-frontend/src/store/chat/__tests__/chatNavigationSnapshotCache.test.ts
Latest Validation Snapshot (2026-08-20 · Industrial-Grade Cross-Platform Release Engineering Pipeline · 4-Target CI Gate)
- Industrial Map-Reduce Release Engineering Pipeline (vs Hermes-Desktop / WeSight / AionUi / Community Practices): Myrm establishes a fully automated, code-enforced release gate network across macOS (Apple Silicon / Intel), Windows (x64), and Linux (AppImage):
- Map-Reduce Pipeline Architecture: Frontend built once and shared across 4 targets; macOS ARM builds publish ahead of time (40% release latency reduction);
- Pre-build Mathematical Pubkey Consistency Check (
check-updater-pubkey.sh): Validates updater public key against CI private key, hard-failing immediately on placeholder values; - Dual-Platform Authoritative Signing & Notarization: macOS 4-layer validation (
codesign+spctlGatekeeper simulation +xcrun staplernotary ticket check) + Windows Authenticode RFC 3161 timestamping; - Platform Defect Self-Healing: Automated stripping of native musl modules preventing AppImage crashes, Dummy-Swap bundling resolving Tauri sidecar packaging bug, and adaptive renaming avoiding cross-arch overwrites;
- Automated OTA Manifest Generation & Runtime Smoke Gates:
finalize-release.sh+verify-release.shpair.sigsignatures with SHA-256 cross-validation, whilesmoke-launch-runtime.shunpackages bundles on CI runners and probes live:8080/healthand WebUI ports.
- Competitor Comparison: Competitors rely on manual Markdown checklists subject to human error and measurement decay; Myrm enforces 18+ automated CI mechanical gates ensuring 100% zero-defect desktop releases.
- Verification Data:
test_desktop_finalize_fixture.py(4-target non-network signature pairing + edge cases) PASSED +test_desktop_launch_smoke_script.pyPASSED +check-fractal-docs.test.ts2 PASSED. - Material SSOT:
temp-docs/materials/DESKTOP_APP_ADVANTAGE.md(2026-08-20 section) ·myrm-agent/.github/workflows/desktop-release.yml·scripts/ci/desktop-release/_ARCH.md
Latest Validation Snapshot (2026-08-20 · Managed Skill Installation Receipt & Atomic Rollback · Roadmap #2 Done)
- Immutable Skill Install Receipts & Atomic Rollback Gate (vs Hermes / OpenClaw / Plugin Markets): Myrm delivers an enterprise-grade managed skill installation safety and audit framework:
- Atomic Snapshot Rollback (
SkillInstallTransaction): Harness-level transaction manager takes an automatic pre-installation snapshot backup of target skill directories, guaranteeing 100% clean rollback without corrupted half-installed states if any I/O, extraction, or downstream agent mounting step fails; - Lifecycle Script Blocking Gate (
LifecycleScriptGuard): In-memory inspection immediately intercepts dangerouspreinstall,postinstall,install, and suspicious shell injection vectors insidepackage.jsonmanifests before extraction; - Immutable Installation Receipt (
SkillInstallReceipt): Computes per-file SHA256 digests, total manifest hash, and security audit scores into an immutablereceipt.jsonon disk; - Verified Receipt UI & Precision Uninstallation: Frontend showcases verified receipt status, while backend uninstallations traverse immutable receipt file manifests for exact cascade cleanup with zero orphan artifacts;
- Full-Stack Test Coverage: 16 harness transaction/receipt unit tests + 4 server discovery receipt tests passed (100% green, zero lint errors).
- Atomic Snapshot Rollback (
- Material SSOT:
temp-docs/materials/TOOL_ECOSYSTEM_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/agent/skills/market/transaction.py·myrm-agent/myrm-agent-server/app/api/skills/discovery.py
Latest Validation Snapshot (2026-08-20 · Wiki Concept Explosion Governance & Sandbox Git Snapshotting · Roadmap #2 & #3 Done)
- Multi-layered Concept Explosion Defense & Isolated Git Snapshot Auditing (vs Hermes / OpenClaw / deer-flow / Obsidian): Myrm delivers an enterprise-grade wiki concept explosion governance and sandbox data integrity architecture:
- Config-Level Ingress Cap (
max_concepts_per_doc = 20): Hard ceiling enforces strict limits per source document, preventing runaway concept generation during bulk ingestion; - Catalog Context Injection (
read_index_context()): Injects existing concept taxonomy directly into concept extraction prompts, coercing the LLM to reuse existing canonical terms instead of fragmenting into synonymous sub-pages; - Batch Control & Resilience Circuit Breaker: 10-article batching combined with
canonical_registryalias deduplication andevaluate_batch_pauseautomatic pauses on quota/rate-limit anomalies; - Human-in-the-Loop (HITL) Pending Approval: All synthesized concepts require review in
WikiPendingEditsbefore permanent storage; - Native Sandbox Git Versioning: 13 mutation paths trigger automated local git snapshot commits (
vault_git.py+git_snapshot.py) with zero remote git synchronization overhead or credential exposure risks; - Full-Stack Test Coverage: 140 harness wiki unit tests + 52 server wiki integration tests + 297+ full-suite wiki regression tests passed (100% green).
- Config-Level Ingress Cap (
- Material SSOT:
temp-docs/materials/WIKI_KNOWLEDGE_BASE_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/toolkits/wiki/·myrm-agent/myrm-agent-server/app/services/wiki/
Latest Validation Snapshot (2026-08-20 · Unified Learning & Skill Evolution Timeline · Roadmap #54 Done)
- Unified Learning & Skill Evolution Timeline with Inline Governance (vs Hermes / Mem0 / Claude / Manus): Myrm delivers an industry-first unified cognitive evolution timeline and inline governance system:
- Cross-Domain Polymorphic Timeline (
LearningTimeline): Fuses factual memories, user preferences, episodic milestones, procedural rules, and skill evolution lifecycle events into a single chronological stream in/journeywith semantic color badges, timestamp filtering, and cursor pagination; - Inline Instant Governance (
is_user_locked): Enables users to directly modify memory rule content, finely calibrate importance scores (0–10 scale), delete obsolete entries, and permanently lock items (is_user_locked=True) to prevent autonomous agent sessions from overwriting verified human preferences; - Skill Lifecycle Hot-Archiving: Allows instant archiving and restoration of skills directly from the timeline, mitigating tool runtime conflicts;
- Time-Space Dual Sync: Seamlessly links timeline cards to the Memory Knowledge Graph (
MemoryKnowledgeGraph) with automatic node focusing and relationship highlighting; - Full-Stack Test Coverage: 7 backend schema/handler unit tests + 5 live HTTP API integration tests passed (
test_learning_timeline_integration.py) + 0 frontend lint errors.
- Cross-Domain Polymorphic Timeline (
- Material SSOT:
temp-docs/materials/MEMORY_ADVANTAGE.md·myrm-agent/myrm-agent-server/app/api/statistics/learning_timeline.py·myrm-agent/myrm-agent-frontend/src/components/features/growth/LearningTimeline.tsx
Latest Validation Snapshot (2026-08-20 · StorageCapabilities Explicit Contract & Fail-Closed Schema Gate · Roadmap #13 Done)
- StorageCapabilities Explicit Contract & Fail-Closed Schema Gate (vs Hermes / OpenClaw / DeerFlow / Generic Agent Stores): Myrm delivers an enterprise-grade storage capability contract and fail-closed startup validation gate across all SQLite storage layers:
- Immutable Storage Capabilities Protocol (
StorageCapabilities): Declaresschema_version,min_compatible_version,supports_atomic_batch,supports_concurrent_readers, andrequired_tablesin an immutable dataclass, embedded natively into SQLite connection hardening (harden_connection_sync/async) andconnect_async; - Fail-Closed Forward & Backward Compatibility Guard: Reading
PRAGMA user_versionon startup aborts immediately (SchemaVersionTooNewError) if an older binary encounters a newer database created by a future build, completely eliminating silent data corruption and table degradation during user downgrades, container image rollbacks, or multi-instance shared volume mounts; - Structural Shape & Minimum Version Defense: Validates essential table existence (
SchemaShapeMismatchError) and asserts minimum supported schema versions (SchemaVersionTooOldError) before any DDL or data queries run; - Stateful Migration Engine Atomic Synchronisation:
StatefulMigrationEngineatomically syncsPRAGMA user_versionupon completing all versioned migrations or baselining, with strict sub-engine isolation (sync_user_version=False) for index runners; - Full-Stack Test Coverage: 80 harness tests passed (including hardening, integrity, profile, schema gate unit tests, and live SQLite/aiosqlite lifecycle integration tests) + 1 server
init_databaselive integration test passed (100% green).
- Immutable Storage Capabilities Protocol (
- Material SSOT:
temp-docs/materials/STORAGE_CRASH_CONSISTENCY_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/utils/db/sqlite/schema_gate.py·myrm-agent/myrm-agent-server/app/database/connection.py
Latest Validation Snapshot (2026-08-20 · Maka UI Standardized Skeleton & Empty State Visual System · Roadmap #2 Done)
- 1:1 Topology-Fitted Skeletons & Resilient Empty States (vs Linear / Raycast / Manus / Generic Open-Source Agents): Myrm delivers an enterprise-grade visual loading and empty state system across all core business dashboards and journeys:
- 7 Standardized Topology Skeleton Templates (
skeleton-templates.tsx): Perfectly mirrors real layout topology acrossListSkeleton,CardGridSkeleton,TableSkeleton,FormSkeleton,ListDetailSkeleton,MetricCardsSkeleton(KPI grid), andTimelineSkeleton(chronological nodes), completely eliminating Cumulative Layout Shift (CLS reduced to 0.00); - High-Signal EmptyState Machine (
empty-state.tsx): Offers 5 clean variants (default,dashed,compact,card,error) with soft icon glows, dark/light theme awareness, complete i18n copy in 6 languages, and full A11y attributes (aria-busy/aria-live); - Full-Route Convergence & Zero Dead Ends: Replaced legacy spinning loaders (
Loader2) and unstyled empty containers acrossGrowthDashboard,LearningTimeline,ArtifactsCenter,RateLimitTab,ToolStabilitySection, andCronJobList; - In-Place Error Degradation & One-Click Retry: Gracefully renders network/sandbox cold-start failures in-place with clear guidance and an instant retry CTA, preventing blank screens or unhandled rejections;
- Full-Stack Test Coverage: 12/12 frontend primitive unit tests passed + 7/7 backend statistics API integration tests passed (100% green).
- 7 Standardized Topology Skeleton Templates (
- Material SSOT:
temp-docs/materials/USER_EXPERIENCE_ADVANTAGE.md·src/components/primitives/empty-state.tsx·src/components/primitives/skeleton-templates.tsx·src/components/features/growth/GrowthDashboard.tsx
Latest Validation Snapshot (2026-08-20 · Shared Skill Pool Cross-Agent HotSync & Dual-Track Distribution · Roadmap #3 Done)
- Universal Skill Pool Single Source of Truth & Zero-Orphan HotSync (vs QwenPaw / CoPaw / Traditional Multi-Agent Frameworks): Myrm delivers a complete global skill pool ecosystem with cross-agent real-time hot-sync and automated orphan reference cleanup:
- Dual-Track Distribution Model: Generalist agents inherit the global skill pool out-of-the-box (zero configuration overhead), while specialized agents maintain explicit whitelists to safeguard prompt token boundaries, with full support for bulk
/pool/syncsynchronization; - Automated Orphan Reference GC: When a skill is uninstalled or deactivated,
remove_skill_from_all_agentsautomatically sweeps and purges invalid skill IDs across all agent allowlists, permanently preventing stale ID prompt injection and runtime LLM execution errors; - End-to-End Real-Time HotSync Broadcast: Emits
SKILL_POOL_UPDATEDevents viaAppEventBus+ SSE and handles them with a 300ms debounce inuseGlobalEvents.tsanduseSkillStore, ensuring smooth real-time UI and agent runtime updates across WebUI, Desktop, and Cloud instances without manual restarts; - Full-Lifecycle Provenance Tracking:
Skill.installed_fromrecords the exact source channel, upstream URL, version, and installation timestamp, surfaced seamlessly inSkillDetailSheetContent; - Full-Stack Test Coverage: 171 core skill tests passed + 62 skills integration tests passed + 150 skills API tests passed (100% green).
- Dual-Track Distribution Model: Generalist agents inherit the global skill pool out-of-the-box (zero configuration overhead), while specialized agents maintain explicit whitelists to safeguard prompt token boundaries, with full support for bulk
- Material SSOT:
temp-docs/materials/TOOL_ECOSYSTEM_ADVANTAGE.md·app/core/skills/discovery/adopt.py·app/api/skills/discovery.py·components/features/skills/SkillDetailSheetContent.tsx
Latest Validation Snapshot (2026-08-20 · Universal External Agent Protocol Matrix & Zero-Code Config Ecosystem · Roadmap #1 Done)
- Standard Protocol Abstraction over Hardcoded Proprietary Cards (vs deer-flow / OpenClaw / Hermes): Myrm delivers a universal external agent integration system (
ExternalAgentsConfig.tsx+toolkits/acp/+delegate_to_agent_tool), enabling zero-barrier access and configuration for any standard CLI/ACP/SDK agent (Claude Code, Codex CLI, Gemini CLI, and custom agents):- Generic Protocol Abstraction vs Fragmented Cards: Eliminates technical debt and code bloat caused by hardcoding niche cards for experimental tools without public CLI specs, focusing on rock-solid protocol convergence;
- Unified Triple-Backend Runtime (ACP/CLI/SDK): Seamlessly supports Claude Code
--resumesession reuse, cross-turn context continuity, automatic SIGTERM/SIGKILL cleanup on timeouts, and structured error propagation; - Per-Backend Concurrency Locking: Strictly serializes same-agent executions to eliminate stdout race conditions while enabling safe parallel execution across different agents;
- Full-Stack Test Coverage: 423 ACP tests passed + 28 External Agent install/auth/delete tests passed (100% green).
- Material SSOT:
temp-docs/materials/SUBAGENT_AND_ORCHESTRATION_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/toolkits/acp/·myrm-agent-frontend/src/components/features/settings/sections/integration/ExternalAgentsConfig.tsx
Latest Validation Snapshot (2026-08-20 · WeCom Event-Driven Progress Bubble & Anti-Zombie Downgrade · Lossless Morphing)
- Seamless In-Place Morphing & Zero Packet Loss for WeCom IM (vs MateClaw / OpenClaw / Traditional Webhook Bots): Myrm delivers a dual-mode adaptive progress evolution and chunking engine tailored for Enterprise WeChat (WeCom):
- In-Place Morph-into-Answer + Sequential Proactive Delivery: Solves WeCom’s strict 2048-character single message ceiling by rendering multi-chunk responses where chunk 1 morphs the progress placeholder via
aibot_respond_msg(finish=True), and all subsequent overflow chunks are dispatched proactively viaaibot_send_msg, ensuring 100% complete long-report delivery without truncations or silent drops; - Zero-Zombie Placeholder Downgrade: Automatically detects self-built app Webhook capabilities (
capabilities.edit=False) and short-circuits placeholder dispatch to prevent uneditable “Thinking…” zombie messages from permanently polluting conversation threads; - High-Signal Stage Rolling & Thinking Scrubbing: Cleanly filters out raw
<think>tokens and tool JSON arguments into human-friendly Stage narratives with 800ms anti-jitter, maximizing clarity and trustworthiness for enterprise users; - Connection Reconnect Stream State Guard: Marks active streams as
is_force_closedupon WebSocket drops instead of dropping them, gracefully delivering final answers upon reconnection.
- In-Place Morph-into-Answer + Sequential Proactive Delivery: Solves WeCom’s strict 2048-character single message ceiling by rendering multi-chunk responses where chunk 1 morphs the progress placeholder via
- Verified: Targeted WeCom suite and Router Scrubber unit/integration tests 200 passed (0 failures, 0 zombie processes).
- Material SSOT:
temp-docs/materials/MULTICHANNEL_ADVANTAGE.md·app/channels/providers/wecom/aibot_channel.py·app/channels/routing/router_stream_scrubber.py
Latest Validation Snapshot (2026-08-20 · Profile Capability Escalation 防提权与 AST CI 门禁闭环)
- 运行时单向收敛天花板与三层立体防线 (vs mem0 / Hermes / OpenClaw / Traditional Agent Platforms): Myrm 在框架执行引擎与服务端构建了完备的防越权与提权阻断体系:
- Least Privilege Ceiling 单向收敛算法:
_merge_ruleset_with_ceiling算法保证 Agent 预设只能收紧安全策略(ALLOW ASK / DENY),严禁放宽(DENY/ASK ALLOW),杜绝 Agent 配置穿透用户安全设定的漏洞; - YOLO 模式强制双锁机制:Agent 无法单方面开启 YOLO 模式,必须由用户全局先开启方可生效;
- Marketplace 导入包安全清洗:
_sanitize_imported_security_overrides自动剥离第三方 Agent Package 中的yoloModeEnabled与通配符*/ 敏感执行提权规则; - 零依赖 AST CI 门禁:
check_profile_capability_escalation.py纯 AST 静态解析_BUILTIN_PROFILES,毫秒级拦截任何试图在 PR 中静默放宽内置策略的提权改动。
- Least Privilege Ceiling 单向收敛算法:
- Verified: Harness 单向收敛测试 57 passed + Server Marketplace 导入清洗测试 22 passed + CI AST 静态审计门禁 exit 0(100% pass rate)。
- Material SSOT:
temp-docs/materials/SECURITY_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/agent/security/channel_presets.py·myrm-agent/myrm-agent-server/app/services/agent/marketplace/import_.py·myrm-agent/myrm-agent-server/scripts/ci/check_profile_capability_escalation.py
Latest Validation Snapshot (2026-08-20 · White-box Retrieval Assertion Engine & Long-doc Deep Penetration)
- Deterministic RAG/Memory Evaluation & Long-Document Penetration (vs Citadel / Hermes / OpenClaw / Mem0): Myrm brings white-box retrieval assertion and deep span penetration testing directly into Eval Lab and Memory Command Center Doctor:
- Zero-LLM Deterministic Evaluation:
RetrievalAssertionevaluates body span recall, document ID matching, and source diversity without incurring expensive or non-deterministic LLM-as-a-judge billing; - Automated Header & Body Dissection:
split_header_and_bodycleanly strips YAML frontmatter and Markdown metadata titles, preventing false-positive hits on document headers; - Collapse Hits & Rank Inversion Protection:
collapse_retrieval_hitsdeduplicates multi-chunk hits from the same document into a single effective rank, stopping single-file flood and measuring true source diversity; - Long-Document Head vs Tail Penetration: Synthetic probes verify retrieval penetration deep into document appendices (e.g., Appendix D routing rules) rather than merely grazing document titles.
- Zero-LLM Deterministic Evaluation:
- Verified: Harness retrieval assertion tests 11 passed + Server diagnostic benchmark tests 3 passed + Eval API integration tests 10 passed (100% green).
- Material SSOT:
temp-docs/materials/MEMORY_ADVANTAGE.md(2026-08-20 header) ·myrm-agent-harness/src/myrm_agent_harness/eval/retrieval_assertions.py·myrm-agent-server/app/services/memory/diagnostics/diagnostic/diagnostic_recall_benchmark.py
Latest Validation Snapshot (2026-08-20 · Composer Inline Agent Quick-Switch · 1-Second Seamless Role Transition)
- Inline Role Hot-Swapping without Context Interruption (vs Codex / Cursor / OpenClaw / Hermes): Myrm elevates the Composer Agent Indicator (
AgentIndicator.tsx) into an interactive DropdownMenu quick-switch palette:- Zero-Interruption Multi-Role Workflows: In long, active conversations, switch instantly between Preset Agents (e.g. Code Expert, Research Analyst) and Custom Agents in 1 click without resetting session history, clearing context, or opening heavy configuration modals;
- Atomic Configuration Application: Seamlessly syncs
agentId,skill_ids,mcp_ids,system_prompt,security_preset, andauto_restore_domainsin a single atomic update (buildAgentConfig), ensuring zero partial configuration leaks; - Concurrency & Stream Guarding: Automatically locks the switch trigger when generation is active (
loading=true), preventing state collisions and prompt cache corruption mid-stream; - Progressive Disclosure: Keeps frequently used agents at the top while offering one-click entrypoints to granular tuning panels and the full Agent Management Hub.
- Verified: Frontend message-input-actions test suite 44 passed (including AgentIndicator edge cases and multi-modal avatar variants) + Server Agent CRUD & Snapshot persistence 18 passed (100% pass rate).
- Material SSOT:
temp-docs/materials/USER_EXPERIENCE_ADVANTAGE.md(§32) ·myrm-agent-frontend/src/components/features/message-input-actions/AgentIndicator.tsx
Latest Validation Snapshot (2026-08-20 · Session-Aware Multimodal Routing & Dual-Track Vision Fallback Engine)
- Zero-Loss Multimodal Direct Channel & Self-Healing Text Fallback (vs mem/pi_agent / Traditional Agent Gateways): Myrm dynamically binds vision execution to the active session model across all 5 execution surfaces (Web, IM Channels, Goal Agent, Cron, Kanban):
- Dual-Track Adaptive Execution: When the session model natively supports multimodal input (e.g., GPT-4o / Claude 3.7 / Gemini 2.0), images are forwarded directly without transcription loss; for text-only models,
VisionFallbackEnginetransforms images into structured, injection-hardened text descriptions; - Reactive Image Compression: Integrated with
ImageCompressorfor automatic payload resizing upon 413 errors; - Capacity Failover Chain: Automatically switches auxiliary vision providers upon 429, timeout, or billing errors;
- SSOT Architecture Guard: Enforced by
test_vision_fallback_execution_surfaces.pyacross all execution surfaces.
- Dual-Track Adaptive Execution: When the session model natively supports multimodal input (e.g., GPT-4o / Claude 3.7 / Gemini 2.0), images are forwarded directly without transcription loss; for text-only models,
- Verified: Vision fallback engine tests 36 passed + Surface architecture guards 6 passed + Vision toolkit integration 2 passed + Config parsers 37 passed (81 tests 100% green).
- Material SSOT:
temp-docs/materials/MULTI_MODEL_CONSENSUS_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/toolkits/llms/vision/fallback_engine.py
Latest Validation Snapshot (2026-08-20 · User-Configurable Multi-Tier Ordered Model Fallback & Streaming Graph Cascade Recovery)
- Uninterrupted Agent Execution via Multi-Tier Streaming Fallback (vs LiteLLM Router / OpenRouter / Traditional LLM Gateways): Myrm brings model retry and fallback directly into the Agent streaming lifecycle:
- In-Loop Streaming Graph Cascade: Unlike external LLM gateways that only do static request retries and break upon mid-stream chunk failures, Myrm’s Harness (
StreamRecoveryMixin) manages an ordered fallback candidate chain (_fallback_llms), cleanly rebuilding the agent execution graph and advancing failover status machines seamlessly; - Tolerant Provider Chain Extraction: Server config parser automatically skips disabled or unconfigured providers in the fallback sequence, preventing single-point chain blocks;
- Target Model Context Shield: Pre-checks target fallback model context limits (
max_context_tokens) and automatically triggers emergency compact guards if downgraded to smaller models; - Strict Safety Block Isolation:
SAFETY_BLOCKerrors are dedicatedly routed to designated safety fallback LLMs without polluting general-purpose API fallback chains.
- In-Loop Streaming Graph Cascade: Unlike external LLM gateways that only do static request retries and break upon mid-stream chunk failures, Myrm’s Harness (
- Verified: Harness stream recovery unit tests 50 passed + Server config parsers 37 passed + LLM Factory assembly 7 passed + Fallback candidate pool tests 14 passed (100% pass rate, 0 failed).
- Material SSOT:
temp-docs/materials/MULTI_MODEL_CONSENSUS_ADVANTAGE.md·myrm-agent-harness/src/myrm_agent_harness/agent/streaming/recovery/stream_recovery.py·myrm-agent/myrm-agent-server/app/core/channel_bridge/config_parsers.py
Latest Validation Snapshot (2026-08-20 · Four-Layer Tool Ordering & Maximal Prompt Cache Optimization · Zero Cache Thrashing)
- Ordered Tool Registration & Prompt Cache Preservation (vs Hermes / OpenClaw / Traditional Agent Frameworks): Myrm enforces a strict 4-layer taxonomy (
ToolRegistry: CORE COMMON EXTENDED EXTERNAL):- Deterministic Cache-Stable Tool Schemas: Tools are serialized strictly by layer priority and deterministic sorting, eliminating schema drift across turns and maximizing Prompt Cache hits across Claude, OpenAI, and DeepSeek (reducing TTFT by 60%+);
- Minimal Initial Footprint: Base agent boots with only 7 CORE meta-tools; desktop, browser, kanban, and subagent tools are dynamically mounted on-demand, preventing turn-1 token bloating;
- High-Precision Semantic Desktop Control (SDC): Leverages native OS Accessibility Trees (
@dref) with coordinate-based visual clicking ([x,y]) fallback, saving ~90% token costs compared to screenshot-heavy CV-based approaches.
- Verified: Harness tool management unit tests 35 passed + Server integration tests 38 passed + Frontend polymorphic rendering 18 passed (100% pass rate, 0 zombie processes).
- Material SSOT:
temp-docs/materials/TOOL_ECOSYSTEM_ADVANTAGE.md·temp-docs/materials/MYRM_AGENT_TOOL_INVENTORY.md·myrm-agent-harness/src/myrm_agent_harness/agent/tool_management/
Latest Validation Snapshot (2026-08-20 · Typora-Grade Split-Pane Markdown Live Preview & Unsaved Draft Guard · Live Chrome E2E)
- End-to-End WYSIWYG Knowledge Base Editing & Approval (vs Hermes / OpenClaw / Traditional Markdown Notes): Myrm delivers a professional split-pane Markdown writing and editing environment in the Knowledge Base (Wiki Concepts) and Pending Edits:
- Monaco Source + MarkdownContent Live Preview Split: Left pane provides VS Code-grade code folding, dark/light theme sync, and keybindings; right pane renders rich text, KaTeX formulas, and syntax-highlighted code blocks in real-time, eliminating the blind editing pain of plain textareas in legacy agent tools;
useDeferredValueAsync Debouncing: Keeps typing silky smooth and responsive even across multi-thousand-word architectural specifications without input frame drops;- Unsaved Draft Guard: Intercepts concept tree navigation or edit cancellation when uncommitted changes exist, presenting a “Discard unsaved changes?” dialog to protect user writing from accidental loss;
- Pro Keybindings & Responsive: Full
Cmd/Ctrl+Squick save support directly wired to narrow apply mutations; automatically collapses to an “Edit / Preview” tabbed layout on mobile devices.
- Verified: Frontend unit tests (
wikiTreeUtils.test.ts) all green + full frontend regression 598 files 4469 passed + Chrome LIVE E2E PASS (test_wiki_markdown_editor_chrome_e2e.py: seed concept open edit mode Monaco key input live preview assertion save mutation persistence loop · 21.31s). - Material SSOT:
temp-docs/materials/WIKI_KNOWLEDGE_BASE_ADVANTAGE.md(2026-08-20 header) ·myrm-agent-frontend/src/components/features/settings/sections/knowledge/wiki/WikiMarkdownEditor.tsx
Latest Validation Snapshot (2026-08-20 · Zero-Latency Device Path Interception & 7-Stage Unicode Self-Healing)
- 100% Transparent Self-Healing for Malformed Paths and Parameter Aliases (vs Hermes / OpenClaw / Claude Code): Myrm builds a multi-layered robustness defense in the file execution engine (
file_read_tool+path_security+path_hint):- 8 Parameter Aliases Normalized Silently: Handles model variances (
filePath,file_path,path,file, etc.) via Pydantic model validator with zero schema changes and zero prompt cache pollution; - 7-Stage Unicode Path Self-Healing: Transparently probes and resolves paths with curly quotes (
“data.json”), fullwidth slashes, zero-width spaces, NBSP, or macOS APFS NFD representations directly on disk, saving turns that would otherwise fail; - Pre-IO Zero-Latency Cross-Platform Device Guard: Intercepts
/dev/zero,\\.\COM1, 22 Windows reserved device names (CON,PRN,AUX,NUL,COM1-9,LPT1-9with any extension), and character/block/FIFO device files before physical I/O occurs, completely preventing sandbox OOM and hanging; - Bounded Levenshtein-2 Suggestions: Prunes search with edit distance, speeding up suggestions by 70%+ while eliminating distant hallucinated file recommendations;
- Strict Line Range Validation: Rejects
<= 0and inverted line ranges (e.g.,:50-10), preventing false “empty file” hallucinations from inverted slicing.
- 8 Parameter Aliases Normalized Silently: Handles model variances (
- Verified:
test_path_security.py+test_path_and_skill_filters.py+test_file_read_aliases_and_safety.pypassed + full regression suite 2887 passed (0 failures, 0 warnings). - Material SSOT:
temp-docs/materials/FILE_SYSTEM_INTELLIGENCE_ADVANTAGE.md(2026-08-20 header) ·myrm-agent-harness/src/myrm_agent_harness/agent/meta_tools/file_ops/file_read_tool.py·myrm-agent-harness/src/myrm_agent_harness/core/security/path_security.py
Latest Validation Snapshot (2026-08-20 · Wiki Concept Explosion Governance & Compile Output Token Budget Governance)
- 4-Layer Compile Output Budget & Vector Window Overflow Guard (vs OpenClaw / Hermes / deer-flow / LobsterAI / CoPaw / jiuwenclaw): Myrm implements strict knowledge compile output governance:
WikiCompileConfig.max_article_lengthhard budget, adaptive truncation with automatic markdown/frontmatter block closing, automatic capture ofEmbedInputTooLargeErrorwithresolve_embed_failuredowngrading to single-doc pause instead of batch crash, and end-to-end HITL Pending approvals. Verified 100% green across 130 unit tests and real LLM end-to-end tests. - High-Fidelity Compliance & Long-Document Distillation Architecture: Fourfold fidelity contract (Zero-paraphrase Prompt + Structured Claims Contract + Line-level Provenance + HITL Approval) combined with native long-context (128k-1M) and directory-level sidecars, completely eliminating the fragmentation of entities caused by naive chunk-digest pipelines.
- Material SSOT:
temp-docs/materials/WIKI_KNOWLEDGE_BASE_ADVANTAGE.md(2026-08-20 header) ·toolkits/wiki/pipeline/compiler.py·toolkits/wiki/pipeline/resilience/failure_policy.py
Latest Validation Snapshot (2026-08-20 · Native Folder Drag-and-Drop Authorization · Zero Friction Workspace)
- Drag local folders directly into chat for instant, uninterrupted Agent access (vs OpenClaw / Hermes Agent / Claude Code): Myrm integrates deeply with desktop drag-and-drop (
tauri://drag-drop) in the Tauri Desktop App. When you drag a local folder from Finder or Windows Explorer into the chat box, Myrm captures its physical path and authorizes it in the session security whitelist within milliseconds. The Agent proceeds in Turn 1 without interruption, completely removing the traditional 4-step friction of “send prompt PathPolicy block prompt suspended open OS file picker locate directory resume”. - Verified: Frontend Vitest (
useDesktopFolderDrop.test.ts) 7 passed + Backend Service test (test_session_access_service.py) 10 passed + API Integration test (test_session_access_roots_api.py) 1 passed in 3.98s (covering complete 5-step lifecycle and idempotency assertions). - Material SSOT:
temp-docs/materials/DESKTOP_APP_ADVANTAGE.md(2026-08-20 section) ·myrm-agent-frontend/src/hooks/message-input/useDesktopFolderDrop.ts
Latest Validation Snapshot (2026-08-18 · Brand Studio UI + Global User Profile · Real Chrome E2E)
- Brand identity is directly editable & automatically injected into agent context (vs Miora / WorkBuddy / Hermes / OpenClaw): Myrm is the only Agent platform that turns brand guidelines (name, tagline, primary/secondary/accent colors, typography, tone, taboos) into an editable, structured Brand Studio under Settings → Knowledge Group → Brand Style. Brand keys (
brand_*) are stored as profile memories and injected into the agent’s stable context layer (Global User Profile, priority 1) on turn 1 — generated outputs (copy, design, HTML, PPT) stay on-brand without breaking prompt cache. - Destructive reset guardrail: Clearing brand inputs requires secondary confirmation via
ConfirmDialog(destructive), and the button is disabled on empty forms. - Verified:
brandSchemaunit tests 15 passed +BrandStudioResetConfirm3 passed + Harness memory middleware 263 passed + Integration 17 passed + Chrome LIVE E2E 1 passed in 61.58s (test_brand_studio_reset_confirm_chrome_e2e.py: seed brand memory → WebUI render → clear confirm → cancel preserve / confirm delete profile memory). - Material SSOT:
temp-docs/materials/MEMORY_ADVANTAGE.md(2026-08-18 header) ·myrm-agent/myrm-agent-frontend/src/components/features/brand-studio/
Latest Validation Snapshot (2026-08-17 · Subagent HITL Approval Chain · Plan A)
- Subagent high-risk bash → parent approval card (vs Claude Code / OpenClaw / Hermes): When a child agent calls
bash_code_execute_tool, the approval interrupt propagates throughdelegate_task_toolto the parent SSE stream (tool_approval_request/action_type=subagent_approval); WebUIPolymorphicApprovalCardsupports approve/edit/reject/allow-always; the child resumes from a shared checkpointer. Real WebUI requires two Approve clicks (subagent_approvalthen bashtool_approval).Verified: Chrome LIVE E2E 5/5 (interrupt live ×4 + WebUI approval flow ×1, 168–483s,PYTEST_SAFE exit=0); TestClient interrupt_e2e ×3; delegate_task_tool 106; harness subagent 365+.Competitors: Claude Code lacks an independent child checkpointer resume chain; OpenClaw/Hermes child HITL often dead-ends in CLI.
Latest Validation Snapshot (2026-08-17 · Mem0 Import + Skill Growth UX · Real Chrome E2E)
- Mem0 users migrate without scripts (vs mem0 SDK / Hermes / OpenClaw): Upload a mem0 export JSON in Settings → Memory; Myrm auto-detects the
memorieslist, shows a dry-run review dialog with a translated Mem0 source label (not a raw i18n key) and amemoriesmapping bucket before you confirm. Leading junk entries in the export no longer misroute the batch tounknown. mem0 is a memory SDK — it has no GUI migration wizard; Hermes/OpenClaw lack a mem0 adapter. - Growth Center tells you where a proposal came from (vs MemOS / Hermes): Cards now show Manual Evolution, Memory Extraction, or Background Review — memory-extraction cases no longer mislabeled as background review, and technical growth-type badges are hidden from the product UI. Background evolution triggers dedupe in-flight tasks per chat and skip shallow turns (<3 tool steps) to avoid redundant LLM capture.
- Verified: mem0 adapter 17 passed + API dry-run 8 passed + evolution trigger guard 3 passed + Growth Center component 3 passed + Chrome E2E 2 passed (mem0 file upload review 31.9s + growth center dashboard).
- Material SSOT:
temp-docs/materials/MEMORY_ADVANTAGE.md·temp-docs/materials/SKILL_EVOLUTION_ADVANTAGE.md·myrm-agent-server/tests/e2e/test_mem0_import_review_chrome_e2e.py
Latest Validation Snapshot (2026-08-16 · Image Send Compression Pipeline · Real Chrome E2E)
- Ultra-HD photos & long screenshots always get through (vs Hermes / OpenClaw / Claude Code / ChatGPT): Myrm auto-compresses images server-side at a base64-space 4MiB threshold — aligned to per-image ceilings of OpenAI/Claude/Groq (raw bytes inflate ~4/3 when base64-embedded, so a raw-byte check would let oversized payloads through). Every source (upload / paste / drag-drop / camera frames / channel messages) funnels into the same pipeline, with 2048px fidelity downsampling keeping vision analysis sharp. Competitors only do one-shot browser-side compression with no server-side guardrails or provider-ceiling alignment.
- 80MP decompression-bomb guard (security): bounded
MAX_DECODE_PIXELS = 80_000_000— a tiny malicious file declaring a huge canvas raisesDecompressionBombErroratImage.open, 0 pixels allocated, no memory blowup; legitimate large images stay decodable (Anthropic’s 8000px/side ≈ 64MP ceiling). - Byte + pixel dual guardrails:
stream_recovery_oneshot._shrink_oversized_imageschecks both base64 byte size (SEND_COMPRESS_TRIGGER_BYTES) and pixel dimensions; PNG re-encode byte growth is not misjudged; unshrinkable images never re-send a rejected payload. - Honest scope: video/audio keep their own graded size limits (100MB/25MB); this snapshot covers the image path only.
- Verified: Harness
stream_recovery_oneshot46 passed + Servermodel_resolver_enrich/enrich_model_capabilities21 passed +image_enrichmentcompression path + Chrome E2E 1 passed (test_image_upload_stream_chrome_e2e.py: real WebUI file injection → multipart upload → thumbnail → send → LLM vision reply → API persistence check, ~3.9MiB image through real compression). - Material SSOT:
temp-docs/materials/MULTIMODAL_CAPABILITY_ADVANTAGE.md(2026-08-16 header) ·myrm-agent-harness/src/myrm_agent_harness/utils/media/image_compressor.py·myrm-agent-harness/src/myrm_agent_harness/agent/streaming/recovery/stream_recovery_oneshot.py
Latest Validation Snapshot (2026-08-16 · Session-Level Performance Composition · LLM duration by model + TTFT)
- Where the time and money of a session actually went (vs Hermes / deer-flow / CoPaw / LobsterAI / OpenClaw): Myrm’s session analytics panel breaks every session down by model — how many calls each model made and their total duration (LLM Breakdown), plus a response-speed summary with average and P95 time-to-first-token (Response Speed). Competitors offer no such user-facing aggregation: Hermes only exposes event-level
duration_seconds(no per-model session totals, no TTFT panel), deer-flow / CoPaw / LobsterAI have no LLM duration aggregation at all, and OpenClaw’s diagnostics are a security-sanitization utility, not a performance dashboard. - Honest scope: this closes the “which model is my bottleneck” question after the fact at session granularity. Live per-turn latency remains covered by the message-level Token Economics + Model Speed Test dashboard (no duplication, same panel complements it).
- Verified: server statistics suite 170 passed (incl.
_build_llm_duration_breakdown3-bucket aggregation + legacy-event filtering + API end-to-end passthrough), frontend vitest 328 passed (SessionAnalyticsDialog + DailyJournal incl. i18n known/unknown mode fallback),tsc/oxlint/verify:i18nall green, Chrome E2E 1 passed in 70.81s (real browser,/journeyfull-page render). - Material SSOT:
temp-docs/materials/OBSERVABILITY_ANALYTICS_ADVANTAGE.md(2026-08-16 header) ·myrm-agent/myrm-agent-server/app/api/statistics/session_analytics.py·myrm-agent/myrm-agent-frontend/src/components/features/settings/sections/system/SessionAnalyticsDialog.tsx
Latest Validation Snapshot (2026-08-15 · Memory E2E on Isolated Private Backends · 8/8 Chrome E2E PASS)
- Memory chain verified end-to-end on real isolated backends + real LLMs: every PRIVATE memory test boots a dedicated backend instance (own SQLite + own model config), LIVE tests run real-model flows, SHARED tests reuse the shared backend. This cycle: voice memory ACL both switches (Settings→Memory “conversation search” persisted to
/config/personalSettings, on and off directions), Memory A/B model disclosure (real WBBench sampled eval + local OpenAI-compatible embedding server, run-history discloses the actual agent model + judge placeholder), Memory A/B report matrix + history render, memory lifecycle/citations/command center/vector persistence — 8/8 Chrome E2E PASSED. - Isolated verification’s honest value (vs OpenManus/OpenClaw/Hermes/Claude Code family): because each PRIVATE test cold-starts its own backend, any server→harness wiring break surfaces immediately. This exact mechanism caught and fixed 10 server import breaks left by the harness
session/subpackaging refactor (session_activity/session_context_pins/session_continuitymoved underruntime/context/session/) —ModuleNotFoundErrorcrashed every private backend at boot. After the fix: 22 server unit tests + 8 memory Chrome E2E all green. Competitors relying on headless-browser or fully-mocked tests cannot detect this class of real-assembly failure. - Verified:
test_voice_memory_acl_chrome_e2e2/2 passed (PRIVATE, real backend + browser toggles persisted to API) ·test_memory_ab_model_disclosure_chrome_e2e1 passed (520s, real LLM eval) ·test_memory_ab_report_and_history_render1 passed · memory lifecycle/citations/command-center/vector-persistence PASSED (2026-08-15). - Material SSOT:
temp-docs/materials/MEMORY_ADVANTAGE.md(2026-08-15 header, “记忆链路 PRIVATE 隔离后端端到端全量验证”) ·ENTERPRISE_RELIABILITY_AND_TESTING_ADVANTAGE.md·myrm-agent-server/tests/e2e/test_voice_memory_acl_chrome_e2e.py
Latest Validation Snapshot (2026-08-14 · Turn-End Guaranteed Teardown · control never lingers)
- Deterministic control release on every termination path (vs KimiCU / Codex / DeepSeek): when a desktop/browser agent turn ends — normal
MESSAGE_END, user stop,ERROR,AGENT_CANCELLED, context overflow, or budget exhaustion — the inspector’s “controlling” state is released immediately and deterministically via a sharedreleaseTurnInspectorControls(chatId)helper. Competitors rely on the LLM calling aturn_endedtool (KimiCU’s documented failure mode: “screen still shows using the computer” after the agent stops), which breaks whenever a weak model forgets. Myrm is zero-dependency on model discipline — backend termination events reaching the frontend trigger the cleanup. - Multi-pane isolation + manual-panel protection:
engagedChatIdscopes cleanup to the turn’s owning chat, so parallel chat panes never close each other’s inspector;isTurnViewdistinguishes agent-driven views from user-triggered manual snapshots, so panels the user opened themselves are never force-closed. Idempotent (no-op when not engaged or chat mismatch). - User-visible win: no more “AI already stopped but the screen still says using computer” residue; parallel sessions never leak control state across panes.
- Verified: frontend Vitest 76 passed across 13 files (turn engagement ×22, teardown/clearActivePlan ×13, view-update ×6, streamConsumer ×18, scoped selectors ×8, release helper ×2, resolveStreamChatId ×5) + Chrome E2E inspector-panel lifecycle 5/5 (
test_inspector_panel_lifecycle_turn_end_releases_engaged_view, 2026-08). - Material SSOT:
temp-docs/materials/COMPUTER_USE_ADVANTAGE.md(2026-08-14 header, “回合结束保证释放” section) ·BROWSER_AUTOMATION_ADVANTAGE.md(BLCV · Multi-Chat Isolation) ·src/lib/inspector/releaseTurnInspectorControls.ts·useDesktopInspectorStore.ts·useBrowserInspectorStore.ts
Latest Validation Snapshot (2026-08-14 · Hermes Hot-Post #242/#243 · Crash Resume + Petdex GUI)
- Crash auto-continue after server restart (vs Hermes v0.19 #242 @tonysimons_): Myrm ships write-ahead
InterruptedTurnMarkerbefore each normal stream, startup scan inauto_continue_interrupted_turns()(15-minute freshness, max 2 attempts crash-loop breaker), optional Settings toggleautoContinueInterruptedTurns, and token_economics persisted on resume. Backend recovery ≥ Hermes for process crash/reboot. Honest gap: Hermes also triggers onsession.resumewhen you reopen a chat; Myrm does not yet auto-continue on chat reopen + live SSE — roadmap topic_13 #55 P1. We do not claim browser-tab reopen parity on the website until shipped. - Petdex GUI parity (vs Hermes #243 Nous Girl): PetGallery + one-click install +
/petzero-LLM palette are landed in WebUI/Tauri/cloud — ≥ Hermes CLInpx petdex installfor everyday users. Honest gap: EmptyChat discover chip for companion/petdex 0% (roadmap topic_12 #85 P1) — marketing points to Settings → Companion → Gallery, not a fake one-click chip. - Pure social posts (#241): Two-word brand praise carries zero product signal — no roadmap item; community discover gaps remain covered by existing #83 Hub / #327 Masterclass SSOT, not social reposts.
- Material SSOT:
temp-docs/materials/ERROR_RECOVERY_ROBUSTNESS_ADVANTAGE.md(2026-08-14 header) ·COMPANION_PETDEX_ADVANTAGE.md(2026-08-14 header) ·app/lifecycle/auto_continue.py·tests/lifecycle/test_auto_continue.py
Latest Validation Snapshot (2026-08-20 · Wiki Concept Explosion & Output Token Budget Governance)
- 4-Layer Compilation Budget & Vector Overflow Failsafe (vs OpenClaw / Hermes / deer-flow / LobsterAI / CoPaw / jiuwenclaw): Myrm enforces strict character/token limits during wiki article generation (
WikiCompileConfig.max_article_length), automatic syntax repair on truncation, and vector context window overflow governance (EmbedInputTooLargeErrorhandled byresolve_embed_failureinto single-doc paused state instead of cascading batch failures). Backed by HITL pending approval workflows. Verified: 130 unit/resilience tests + real MiniMax-M3 LLM end-to-end full lifecycle workflow passing 100%. - Material SSOT:
temp-docs/materials/WIKI_KNOWLEDGE_BASE_ADVANTAGE.md(2026-08-20 header) ·toolkits/wiki/pipeline/compiler.py·toolkits/wiki/pipeline/resilience/failure_policy.py
Latest Validation Snapshot (2026-08-14 · Wiki Provenance Trace Closure · message-level deep link)
- Full provenance trace chain (vs Hermes / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw): every knowledge item keeps a trace back to the exact source message — auto-archived turns write
source_chatinto raw frontmatter, the compiler preserves provenance across regenerations (re-attaching on first compile from raw sources when all sources agree), concept detail and pending-review panels deep-link “Source conversation” to the precise message (?highlight=…, chat-level fallback, natural empty state if the chat was deleted), and the raw gate rejects path traversal so provenance can never become a file-system escape vector. Verified: harnesstest_raw_gate+test_compiler27 passed ·test_skill_agent_review+test_skill_agent_wiki45 passed · servertest_wiki_api38 passed · frontend vitest 11 files 43 passed · live-API integration loop (compound→approve→publish→get_concept) + real Chrome E2E PASSED (WebUI concept-detail jump asserted end-to-end, no mocks, 2026-08-14). - User-visible win vs competitors: deer-flow offers thread-level jumps only; openclaw/jiuwenclaw have line-level data (
*.jsonl#L<line>) with zero UI; CoPaw stores a baresession_idwith no jump; hermes-agent has no provenance field. No competitor links a knowledge entry to a precise conversation message in the UI — this is a category gap Myrm alone closes. - Material SSOT:
temp-docs/materials/WIKI_KNOWLEDGE_BASE_ADVANTAGE.md(2026-08-14 header) ·toolkits/wiki/pipeline/compiler_provenance.py·raw_gate/service.py·app/api/wiki/router.py·tests/e2e/test_wiki_concept_provenance_jump_chrome_e2e.py
Latest Validation Snapshot (2026-08-14 · Hermes Hot-Post Pipeline · Cloud Egress vs IronProxy)
- Hermes IronProxy (#284 @NousResearch · 740🔖) vs Myrm Cloud sandbox credential firewall (code-level, 2026-08-14): Hermes
hermes egress setup= Docker-only TLS MITM iron-proxy + stand-in bearer tokens + CLI wizard · host LLM still uses real.envkeys · Modal/Daytona/SSH not wired. Myrm Cloud/SaaS >> Hermes iron-proxy:virtual_key.pyHMAC virtual keys (useless outside relay) +llm_relay/relay.pyin-place injection +egress_control.pyiptables blocks direct LLM egress +secrets_vaultisolation — 196 credential-security tests passed, no local MITM CA attack surface. Honest Local gap: bashcredential_env_overridesmay still inject real OAuth/API tokens into local sandbox children — Doctor surfaces this honestly (#9 P1); we reject cloning iron-proxy CLI/MITM CA. - GAIA marketing reshare (#283 @REMIX_KSA · dedupe #313): Atomic GAIA L1 37:31 headline is community marketing, not a Myrm-shipped receipt. Eval Lab breadth landed (WBBench 4 tracks + BrowseComp + Matrix + EvalManifest) — GAIA adapter not yet shipped (roadmap #72 P1). We do not paste GAIA score chips on the website without EvalLab receipts.
- Opus 5 availability (#285/#281 @Teknium): BYOK paths (OpenRouter / Anthropic Direct / local) ≥ Hermes · no Nous Portal lock-in · UsageModelBreakdown land — zero new roadmap items (pure dedupe).
- User-visible migration win: Same agent workloads, but from CLI + YAML + guess-if-keys-leaked to GUI workspace + auditable Eval Lab + cloud zero-secret sandbox + 35-channel ChatOps + one-click Agent package install — with explicit honest gaps labeled, not benchmark reposts.
- Material SSOT:
temp-docs/materials/COMPETITIVE_HERMES_ADVANTAGE.md§35 ·SECURITY_ADVANTAGE.md(2026-08-14 header) ·EVAL_LAB_BENCHMARK_ADVANTAGE.md§GAIA honest boundary
Latest Validation Snapshot (2026-08-12 · Marketplace Agent 包一键安装 + 云端安全升级)
- Marketplace Agent package install → safe force-push upgrade → rollback → dynamic roster (vs Claude Code/Cursor / Coze 3.0 / OpenClaw):
test_marketplace_import_full_chain.pyintegration 14 passed (real in-memory SQLite + real skill writes + in-process ASGI HTTP, zero mocks on critical paths) + unit regression 70 passed;import_agent_profile.py100% line coverage (2026-08-12 live run). - Verified, end-to-end: ① one-package atomic install (model + prompt + bundled skills + subagents + MCP + safety profile) with publisher→local ID remapping and full rollback on any step failure; ② force-push upgrade treats package fields serialized as
Noneas “leave untouched” — NOT NULL columns never receive empty values and local skill/subagent bindings survive the update (no more overwritten customizations); ③ every force-push writes apre-force-pushsnapshot first and emits anAGENT_CONFIG_UPDATEDevent — one-clickrollback_profilerestores the previous config; ④allow_discoveryper-agent switch controls dynamic-roster membership, toggleable live viaupdate_agent(on → roster, off → hidden, back on → rejoins); ⑤ CP transport security: missing/mismatchedX-Telemetry-Token→ 403, wrong targetagent_id→ 404 fail-closed, sandbox deployment rejects bundled skills. - User-visible win vs competitors: Claude Code/Cursor require hand-assembling subagent YAML + plugins + MCP + skill bindings in four separate steps; Coze keeps a platform-fixed team list (no local private agents, no per-agent discovery switch); OpenClaw ships fixed config files. Myrm is the only workspace where an entire Agent arrives in one package, upgrades never clobber local customizations, every upgrade is one click away from undo, and private agents can opt in/out of team discovery at will.
- Honest scope note: force-push/rollback currently covers the agent-profile layer (skills, subagents, MCP bindings, tools, memory policy, workspace policy, cron verify); full skill-file content diffs per push are not yet rendered in a dedicated upgrade GUI — snapshot restore is API-driven today.
- Material SSOT:
temp-docs/materials/SKILL_MARKET_AND_ADOPTION_ADVANTAGE.md§市场 Agent 包一键安装 ·app/api/internal/import_agent_profile.py·app/services/agent/marketplace/import_.py·tests/integration/test_marketplace_import_full_chain.py
Latest Validation Snapshot (2026-08-11 · LLM 结构化 JSON 输出解析鲁棒性 · 六竞品代码级实证)
- One self-repairing extraction layer, 17+ modules (vs OpenClaw / Hermes / deer-flow / LobsterAI / CoPaw / jiuwenclaw): harness robust parsers
parse_llm_json_object/parse_llm_json_list(utils/chat_utils.py) now back 17+ business modules — skill evolution, wiki compilation, contradiction detection, memory extraction (implicit feedback / cognitive deriver / staleness review / proactive extraction), sidecar summaries, browser session structured extraction, semantic judge, and delegation orchestration. Unit 4899 passed; integration with real LLM 13 passed (tests/integration/llm_extraction/+ staleness reviewer + wiki reindex, no mock); harness architecture gates green. - Verified, code-level per competitor (2026-08): OpenClaw production code calls
JSON.parsedirectly (agent-run-control-shared.ts:267,state/openclaw-state-ownership.ts:79) — a fence or one prose sentence breaks it. Hermesextract_json_candidate(tools/delegation_output_schema.py:80) only strips fences + first/last bracket slicing then barejson.loads; itsjiter_preload.pyis just the OpenAI SDK’s Rust accelerator, it repairs nothing. deer-flow has no LLM-text JSON extractor at all (alljson.loadsare on its own file/message data). LobsterAIparseJsonObject(sessionDiagnostics/archive.ts:63) is bareJSON.parse+ catch→null. CoPawextract_json_payload(utils/structured_output.py:59) does fence + balanced-bracket scan but takes the first JSON and never escapes control chars or strips trailing commas;json_repairis only used for config files. jiuwenclawextract_json_from_response(server/hooks/executor.py:270) tries direct→regex-fence→{}slice then silently returns{}. - Four capabilities no competitor has: ① pick the last (not first) valid JSON block — survives reasoning-model draft-then-final outputs; ② escape stray control characters inside JSON string literals (
_escape_control_chars_in_strings); ③ strip trailing commas (_strip_trailing_commas); ④require_keyschema-aware filtering so an array vs object ambiguity resolves to the object carrying the target key. This is the M1-M17 hardening that replaced fragile hand-rolled parsing across the harness. - User-visible win: skill/wiki/memory/judge flows stop silently failing when a model wraps JSON in prose, adds a fence, prints a draft block first, or emits
\n/\tbare inside strings — no manual retries, no “retry once” loops. - Material SSOT:
temp-docs/materials/REASONING_MODEL_COMPATIBILITY_ADVANTAGE.md§LLM 结构化 JSON 输出解析鲁棒性 ·myrm-agent-harness/src/myrm_agent_harness/utils/chat_utils.py
Latest Validation Snapshot (2026-08-09 · Kanban IN_REVIEW Human Approval Gate · 9-State Lifecycle)
- Task-level human approval gate end-to-end (vs Hermes / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw): harness
test_in_review.py11/11 + servertest_in_review_api.py14/14 = 25 tests, 0 failures (2026-08-09 live run,scripts/dev/run-pytest-safe.sh). - Verified: Tasks created with
require_approval=trueland in IN_REVIEW (never COMPLETED) after dispatcher verification;/approve→ COMPLETED with dependent promotion;/reject→ READY for rework with reason written back toerror; manual move out of IN_REVIEW returns 409 (gate cannot be bypassed); approve/reject are idempotent outside IN_REVIEW; concurrent approve succeeds exactly once; rejected reason + review history surface in the worker context;REVIEW_REQUESTEDnotification routes to the source chat viaBACKGROUND_TASK_DONEwithsuppress_web_push;kanban_add_tasksupportsrequire_approval(orchestrator set). - User-visible win vs competitors: Kanban lifecycle is now 9 states + 23 events (added
IN_REVIEWstate andREVIEW_REQUESTED/APPROVED/REJECTEDevents). No competitor has a task-level human approval gate: Hermes completes tasks directly (no approval state), OpenClaw’s Workboard is display-only (zero LLM tools), and deer-flow/LobsterAI/CoPaw/jiuwenclaw have no native Kanban at all. Myrm is the only agent workspace where a verified long-running task can be parked in review and must be explicitly approved or rejected by an operator — approval is operator-driven via REST/GUI, not a bypassable LLM tool. - Honest scope note: Approval is synchronous REST/GUI (no inline IM approve button yet); review-history context surfaces in worker context via tool output (no dedicated GUI panel yet); the gate applies per-task on opt-in (
require_approval), not globally. - Material SSOT:
temp-docs/materials/KANBAN_TASK_SCHEDULING_ADVANTAGE.md§IN_REVIEW ·toolkits/kanban/dispatcher.py·services/kanban/review_ops.py·api/kanban/routes/tasks.py
Latest Validation Snapshot (2026-08-03 · Wiki Compile Protection Chain · Wiki Roadmap #6)
- 7-layer compile protection chain integrity (vs Hermes / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw): Harness compile_resilience 12/12 + queue 13/13 + compiler core+extended 52/52 + survey 10/10 + Server wiki_queue_resilience 2/2 + wiki_api+ingest 47/47 + import_raw_gate 4/4 + Frontend wikiQueuePoll 7/7 + WikiSection.evidence 1/1 = 148 tests, 0 failures (~2min, 8 batches, harness + server + frontend).
- Verified: Incremental SHA256 hash filtering (
_filter_changed_files) compiles only changed files;batch_size=10limits per-batch LLM cost;evaluate_batch_pausecircuit breaker (AUTH/BILLING → immediate pause, RATE_LIMIT×3 → pause, ≥5 failures → forced pause);is_compile_pausedskips processing when paused;reset_stale_processingrecovers stuck items after 5min;WikiPendingEditsapprove/reject/edit HITL with stale detection;WikiQueuePanelreal-time SSE monitoring with pause/resume/cancel/retry controls +WikiCompilePhaseBar3-phase progress visualization. No competitor implements any compile protection mechanism. - Why scope confirm gate is not needed: Pre-compile file selection UX is fundamentally poor — users cannot meaningfully “select” files to compile without reading each one’s content. Post-compile Pending Edits provide the effective HITL: users see extracted concept names and content, enabling meaningful approve/reject/edit decisions. The 7-layer protection chain (incremental + batch + circuit breaker + HITL) already controls LLM costs comprehensively.
Latest Validation Snapshot (2026-08-03 · SaaS Billing Transparency Panel · Roadmap #1)
- SaaS Billing Transparency Panel (vs TRAE / Cursor / Claude Code / Codex / OpenClaw / Maka): CP billing catalog 12/12 pytest (tier_multipliers + topup_available + model_tier + burn_table edge cases) + FE TypeScript compilation verified + Chrome E2E Account page render = 12 tests, 0 failures.
- Verified: CP
burn_table.pyexposesTIER_MULTIPLIER(LITE 1×/STANDARD 3×/FRONTIER 10×) +TIER_EXAMPLESviacatalog.pytier_multipliersAPI field; FEQuotaDisplay.tsxrenders 5-dimension SaaS billing panel (subscription WU progress bar + topup WU compact balance + daily refresh allowance + free models list + tier multiplier rate card) + competitive pricing explainer (6-locale i18n: Claude Code per-session / Cursor monthly / Codex per-completion / Myrm BYOK per-token) + reliability trust bar;cp-billing.tspatchedbilling_provider+topup_available+yearly_checkout_available; NaN-safe arithmetic (limit > 0guard instats.chat.per/stats.search.per+StatCardpercentage). - User-visible competitive advantage: Myrm is the only AI agent product to publicly display model tier multipliers, daily refresh quotas, free model lists, and a cross-platform pricing comparison — all in a single GUI panel. TRAE hides WU conversion rates; Cursor bundles monthly flat-rate without per-model transparency; Claude Code charges per-session without model weight disclosure; Codex exposes only a single
rollout_budgetdimension; OpenClaw shows OAuth provider-level percentage windows only (no WU breakdown); Maka provides post-hoc usage statistics tables only (no real-time quota, no multiplier). - Honest scope note: Competitor Pricing Model Explainer is static i18n text (not live API); reliability trust bar links to approval/PDV docs (not live SLA metrics); topup WU display is conditional (renders only when
topupWu > 0). - Material SSOT:
temp-docs/materials/USER_EXPERIENCE_ADVANTAGE.md§30 ·QuotaDisplay.tsx·QuotaWidgets.tsx·burn_table.py·catalog.py
Latest Validation Snapshot (2026-08-03 · Humanize Dual-Layer SSOT · Roadmap #6+#8)
- Humanize + scope + save-skill preview dialect (vs OpenWorker / Hermes / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw):
humanize.test.ts+utils.humanize.test.ts12/12 +PolymorphicApprovalCard+ToolCallApproval+SaveSkillApprovalPreview18/18 +SingleApprovalCard.saveSkill+shellCommandDisplay+buildToolApprovalRequest19/19 = 49 vitest, 0 failures (~30s, 3 batches, no Chrome MCP — MCP navigate timeout;curl :3000200 OK). - Verified: FE
lib/humanize/SSOT wires ProgressSteps + three approval outlets (Single / Polymorphic / ToolCall);resolveScopeNoteexternal only viaEXTERNAL_TOOLS(fixed false external onabout:blank/ colon-like browser targets);SaveSkillApprovalPreview+skill_managenormalize;PtcHintBadgessix-locale PTC hints; six-localehumanize.*+toolApproval.ptc.*i18n verify passed in pretest. - User-visible win vs OpenWorker: six locales vs EN-only humanize; three-outlet scope parity (OW: ApprovalCard + InboxItemCard — Myrm has no inbox product); save-skill preview + PTC badges are Myrm-only polish. vs six claw repos: no FE tool-line humanize SSOT at all.
- Honest scope note: no HumanLine bold emphasis (OW has TitleText/LineText — ROI ~2/10); no Transcript declined
humanizeAskrow (Progress cancelled→ask tense covers); Chrome live approval Drawer E2E not signed this round; vitest coverage API fails under Bun — do not cite coverage %. - Material SSOT:
temp-docs/materials/USER_EXPERIENCE_ADVANTAGE.md(Humanize Dual-Layer subsection) ·myrm-agent-frontend/src/lib/humanize/_ARCH.md·components/approval/_ARCH.md
Latest Validation Snapshot (2026-07-31 · Assessment Import Pipeline · Project Milestones)
- Assessment Artifact Smart Import (vs OpenClaw / Hermes Agent / Deer Flow / LobsterAI / CoPaw / JiuwenClaw): backend artifact API 8/8 + milestone API 15/15 + assessment service 6/6 + statistics API 4/4 = 33 pytest, 0 failures; frontend import candidates 7/7 + metrics 5/5 + error mapping 9/9 = 21 vitest, 0 failures. Total 54 tests, 0 failures (< 12s, 4 batches, no Chrome MCP).
- Verified:
GET /files/artifacts?project_id=project-scoped candidate filtering + semantic candidate probe (reusesparse_assessment_markdown);POST /projects/{id}/milestones/import-assessmentidempotent ledger + failure rollback + structuredimport_reason;POST /statistics/assessment-import/eventsfunnel metrics (attempt/success/fail/dropped) +GET summaryaggregation; frontendProjectMilestonePanelcandidate list (importable-first ordering) + one-click import + manual ID fallback + structured error i18n mapping (6 reasons) + telemetry buffer with drop recovery. - User-visible competitive advantage: All 6 competitor repos lack “assessment artifact → project milestone/task” import pipeline; no semantic candidate probe (verify importability before display); no idempotent ledger (duplicate import rejection); no failure auto-rollback; no structured error feedback; no funnel observability. Myrm exclusively delivers the full “assessment output → candidate probe → one-click import → idempotent guarantee → failure rollback → structured feedback → funnel observability” closed loop.
- Honest boundaries: Candidate ordering currently splits by “importable yes/no” + recency; no content-relevance semantic ranking yet (next optimization);
dropped_reportuses global context (project-level troubleshooting needs future bucket); frontend summary max 90 days vs backend 365 days (low-frequency scenario, acceptable). - Material SSOT:
temp-docs/materials/PROJECT_MILESTONE_ROADMAP_ADVANTAGE.md·app/api/files/artifact_api.py·app/api/statistics/assessment_import.py·ProjectMilestonePanel.tsx·assessmentImportMetrics.ts
Latest Validation Snapshot (2026-07-31 · MCP Elicitation HITL Bridge · Roadmap #6)
- MCP Elicitation → GUI HITL bridge (vs Hermes / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw): harness
TestBuildElicitationCallback13/13 + servertest_mcp_elicitation_handler16/16 = 29 tests, 0 failures (~4s, 2 batches, no Chrome MCP). - Verified: Harness
MCPSessionActor._build_elicitation_callbackbridges MCP SDKElicitRequestto business-layer handler with 5-path safety (accept→accept, decline→decline, timeout→cancel, exception→decline, unexpected→decline); URL-mode elicitation explicitly declined; serverbuild_mcp_elicitation_handlercreatesApprovalRecordviaApprovalRegistry+asyncio.Eventsuspension + SSE push to frontend;_normalize_decisionmaps 8 decision variants to 3 MCP actions (accept/decline/cancel); frontendPolymorphicApprovalCardrendersmcp_elicitationwith amber alert header + server name +ExpiryCountdown. - User-visible win vs competitors: Myrm provides GUI approval card with countdown + persistent audit trail — Hermes uses CLI
readlineblocking (no timeout, no audit, no GUI); OpenClaw’s elicitation bridge is Codex-plugin-specific (auto-fill heuristics, not general-purpose user confirmation); deer-flow / LobsterAI / CoPaw / jiuwenclaw have no elicitation support. - Honest scope note:
ElicitResult.content(form data collection) not implemented — no competitor supports general-purpose user form input;edited_payloadinfrastructure supports future integration if needed. URL-mode elicitation declined (browser redirect not applicable in server-hosted architecture). - Material SSOT:
temp-docs/materials/TOOL_ECOSYSTEM_ADVANTAGE.md·toolkits/mcp/session_actor.py·services/agent/backends/mcp_elicitation_handler.py·PolymorphicApprovalCard.tsx
Latest Validation Snapshot (2026-08-03 · Artifact + Wiki Ingest Pipeline · Wiki Roadmap #5)
- Artifact system + wiki ingest pipeline integrity (vs Hermes / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw): Harness artifacts core 160/160 + Server artifact_api + wiki_ingest 22/22 + Server artifacts core + deliverable/scanner 23/23 + Harness artifacts_ready + artifact_observer 14/14 + Harness raw_gate + vault_archive 8/8 = 227 pytest, 0 failures (~1min, 5 batches, harness + server).
- Verified:
ArtifactProcessortemplate method pattern (event parse → filter → persist → metadata) + XSS protection (active content forced download) +short_file_idSSE support;wiki_ingest_tool3-way dispatch (URL/file/text) + binary document parsing + large file auto-splitting + conflict policy FAIL + security scan;deliverable/scanner40+ extension support + 5MB size limit + code block filtering;ArtifactVaultput_file + sha256 hash storage + DB Artifact + ArtifactVersion versioning. Architecture correctly separates Artifacts (work output storage/viewing) from Wiki (knowledge base ingest/compile/query) — connected on-demand viawiki_ingest_tool(Agent-initiated precise write). No competitor implements automatic deliverable→Wiki ingestion.
Latest Validation Snapshot (2026-08-03 · Obsidian Import + File Watch Coverage · Wiki Roadmap #4)
- Obsidian vault import + file watch + Second Brain cron idempotency (vs Hermes / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw): Harness wiki structure/vault_archive/raw_gate 33/33 + Server obsidian adapter/export/reveal 49/49 + file_watch_service/browse_watch 6/6 + wiki_api + vault_portability 34/34 + second_brain_onboarding 10/10 + local_actions/project_api 68/68 + workspace_sync/migration 12/12 + Harness file_ops 41/41 = 253 pytest, 0 failures (~2min, 5 batches, harness + server).
- Verified:
_process_obsidian_vault[router.py:1990-2044] full pipeline (scan + frontmatter parse + image migration + conflict detection + security governance + audit trail); 4 import channels (obsidian/obsidian-zip/folder/zip);WorkspaceFileWatchServiceref-counting + debounce 0.8s + temp file filtering +is_dangerous_pathsafety + max 32 watch paths + SSE event push; Open in Obsidian cross-platform (macOS/Windows/Linux); Project mount sync + SSE auto-refresh. Fixed 1 pre-existing bug:test_apply_second_brain_preset_idempotent_reuses_cronmockSimpleNamespacemissingcommand/job_typeattributes + third cron job (wiki_maintain) not mocked → Pydantic validation failure.
Latest Validation Snapshot (2026-08-03 · OKF Extended Provenance Metadata + Session Orientation · Wiki Roadmap #1+#2)
- Wiki provenance metadata full-stack + session orientation parity (vs Hermes llm-wiki / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw): Harness cognitive_map/sidecar/index_routing 41/41 + retrieval/query 63/63 + compiler/frontmatter 87/87 + Server wiki 283/283 = 474 pytest, 0 failures (~4min, 4 batches, harness + server).
- Verified:
WikiProvenance(StrEnum)6-member enum infrontmatter_contract.py;pending_editsDBprovenancecolumn with idempotentALTER TABLE; YAML_coerce_enum_valuesserialization fix; compiler + contradiction_synthesis magic strings replaced; ServerConceptResponse.provenanceparsed from frontmatter; Frontend provenance Badge inWikiPendingEdits.tsx+WikiConceptDetailPanel.tsx(dark mode + 5-locale i18n). Session orientation:read_hot_context()+read_log_context()auto-prepend in everywiki_query;build_hot_snapshot()zero-LLM vault status;parse_index_entries()+match_index_entries()token-overlap routing;render_schema_markdown()SSOT auto-generation; L0/L1 sidecar hierarchical navigation. - User-visible win vs competitors: Hermes session orientation requires Agent to manually execute 3
read_filesteps per session (SCHEMA + index + log) — Myrm automates all three in the query engine with zero manual operations and zero LLM cost. Provenance badges let users see exactly how each wiki page was created (compiled, chat save, import, contradiction synthesis, etc.) — Hermes has no equivalent provenance tracking UI. - Material SSOT:
temp-docs/materials/WIKI_KNOWLEDGE_BASE_ADVANTAGE.md·toolkits/wiki/core/frontmatter_contract.py·pipeline/cognitive_map/·retrieval/query.py
Latest Validation Snapshot (2026-08-03 · Hermes Wiki Vault Migration Parity · Wiki Roadmap #3)
- Migration + Wiki import parity (vs Hermes llm-wiki / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw): Migration services 225/225 + Wiki Import/API 79/79 + Memory Import Adapters 49/49 + Architecture Closure 3/3 = 356 pytest, 0 failures (~2min, 6 batches, server).
- Verified: Hermes codebase has zero wiki keywords (full scan of hermes-agent/ + hermes-workspace/); llm-wiki is an optional SKILL producing plain Markdown folders; Migration Wizard 5-lane coverage complete (SOUL/MEMORY/USER/skills/cron/MCP);
POST /wiki/import/{folder,zip,obsidian,obsidian-zip}4 channels cover any Markdown folder import with conflict detection + security governance + audit trail. - Pre-existing bugs fixed: (1)
agent_has_moa_overlay_refsdelegated tois_moa_preset_configuredwhich requiredenabled: true— migration skip logic now checks reference_model_selections directly (regardless of enabled state); (2)MemoryCommandMigrationImportSourceLiteral missinggbrain+pi→ manifest validation crash on 7-source payload. - User-visible win vs competitors: No competitor has cross-product wiki vault migration; Myrm Wizard auto-discovers 5 sources + 4 wiki import channels + security governance. Hermes has no migration tooling.
- Material SSOT:
temp-docs/materials/WIKI_KNOWLEDGE_BASE_ADVANTAGE.md·services/migration/·api/wiki/router.py
Latest Validation Snapshot (2026-07-30 · Wiki Maintain Cron + SCHEMA + Compile Index + Query Log · Roadmap #31–#33+R1+W36)
- Wiki maintain cron + compile index seed + SCHEMA + query log orientation (vs Hermes llm-wiki / OpenClaw / deer-flow / LobsterAI / CoPaw / jiuwenclaw): maintain server 6/6 + schema writer 5/5 + compile index R1 2/2 +
test_query29/29 + second_brain/cron router 4/4 = 46 pytest, 0 failures (~45s, harness + server lanes, no Chrome MCP). - Verified: Cron blueprint
wiki_maintain(structuralfast lint vsfulldrift/backlink) via router__wiki_maintain__:{mode}→maintain_runnerSSOT; Settings maintain mode Select + compile-busy 409 skip; Second Brain preset adds weekend 03:00mode: full(triple cron with read-it-later + morning delta); compile_extract_concepts_from_docinjectsread_index_context(); autowiki/SCHEMA.md+ preset vault seed; zero-LLM query prefix addsRecent activity logfromlog.md(≤1.5k) alongside hot context. - User-visible win vs competitors: Hermes llm-wiki relies on manual SKILL lint/orientation + manual log scanning — no GUI cron maintain, no compile index seed, no auto SCHEMA contract file. Six reference claw repos lack Settings maintain mode + scheduled vault hygiene + query log orientation parity.
- Honest scope note: latest hot event may appear in both hot and log sections (accepted low-cost tradeoff); Hermes-style full tag taxonomy in SCHEMA remains WORTH:NO for GUI-first users (R3/W37 rejected).
- Material SSOT:
temp-docs/materials/WIKI_KNOWLEDGE_BASE_ADVANTAGE.md·services/wiki/maintain_runner.py·pipeline/cognitive_map/schema_writer.py·retrieval/query.py
Latest Validation Snapshot (2026-07-30 · Chat Wiki Knowledge Quick Lane · Roadmap #30)
- Chat Wiki knowledge quick lane (vs Hermes llm-wiki / OpenClaw memory-wiki / deer-flow / LobsterAI / CoPaw / jiuwenclaw): intent + knowledge_query_service + wiki_knowledge_lane + API query = 16 pytest, 0 failures (~6s, server lane, Chrome MCP E2E pending sign-off).
- Verified:
should_use_wiki_knowledge_lanegate (agent mode +enable_wiki+ vault ready + question intent);execute_wiki_knowledge_querySSOT (SettingsPOST /wiki/queryshared with lane);wiki_knowledge_laneemits STATUS → SOURCES → MESSAGE →execution_lane=wiki_knowledge; query failure yields explicit failed STATUS (no bare raise, no silent GeneralAgent fallback); FE ProgressStepswiki_knowledge_lane/_clear+message_endlane field. - User-visible win vs competitors: Settings already supports zero-LLM wiki query — short knowledge questions in Chat no longer burn multi-turn GeneralAgent tool calls. GUI Chat zero-LLM wiki lane is unique; Hermes uses SKILL/CLI multi-step; OpenClaw uses memory-wiki tool chain; six reference repos lack Settings+Chat retrieval parity with ProgressSteps + SOURCES Drawer.
- Honest scope note: no lite synthesis; non-gate queries still use GeneralAgent; Chrome MCP Chat E2E + FE vitest
completionEvents.wikiKnowledgeLanepending local bun sign-off. - Material SSOT:
temp-docs/materials/WIKI_KNOWLEDGE_BASE_ADVANTAGE.md·USER_EXPERIENCE_ADVANTAGE.md§22 #30 subsection ·services/wiki/_ARCH.md·lanes/_ARCH.md
Latest Validation Snapshot (2026-07-30 · ChatExecutionPrewarm Turn1 · Roadmap #29)
- Turn1 cold-start prewarm (vs Hermes CLI parallel init / OpenClaw TUI pre-send): coordinator 4/4 + prewarm API 3/3 + CDP log helper 2/2 = 9 pytest, 0 failures (~10s, server lane). Chrome E2E lane adds
Turn prewarm requestedlog count ≥1 before existing 2msg1build execution-cache proof (test_execution_cache_chrome_e2e.py). - Verified: EmptyChat mount + MessageInput focus + AgentConfigPanel switch →
POST /agents/chats/{id}/prewarm; send pathjoin_for_turn(0.3s)+coalesced_acquire; SSE ProgressStepsturn_prewarm_agent/turn_prewarm_memory+ unconditionalturn_prewarm_memory_clearonbrief_pending; FE module-level inflight dedupe; autoOnMount does not DELETE on unmount (first-send EmptyChat→ChatWindow safe). - User-visible win vs competitors: Hermes v0.19 parallel init is CLI-only; OpenClaw prewarm is TUI before first send — neither exposes GUI EmptyChat/focus/agent-switch triggers with ProgressSteps waiting UX. Turn2+ still uses per-chat
execution_cache_reuse(Chrome E2E 2msg1build). - Honest scope note: join window 0.3s — very slow memory may still miss brief on Turn1 (same class as legacy 250ms serial); prewarm init failure has no GUI toast (send cold fallback). Frontend vitest 4 cases not run in agent env (no bun/node) — run locally:
bun run test src/hooks/chat/__tests__/useChatTurnPrewarm.test.ts. - Material SSOT:
temp-docs/materials/COMPETITIVE_HERMES_ADVANTAGE.md§31 ·USER_EXPERIENCE_ADVANTAGE.md§22 Turn1 subsection ·prewarm/_ARCH.md
Latest Validation Snapshot (2026-07-30 · Hermes 8-Article Series + OpenClaw Comparison · Article 13)
- Series takeaway (vendor blog finale, no comments): OpenClaw = stable tool · Hermes = growing partner. Myrm = GUI-first growing partner plus OpenClaw-grade Channel/Skill/Tool skeleton, same on Local / Tauri / Cloud CP.
- What migrators actually get (honest):
- Myrm ≫: GUI Migration Wizard (5 sources, dry-run/confirm/rollback) vs
hermes claw migrateCLI; 4-step onboarding vs terminal-only; 8-type memory + Wiki vs Hermes 2,200-char MEMORY.md; HITL write approval vs Hermes direct writes. - Myrm ≥: Cron + proactive nudges (Cron GUI + push + situation_report + HeartbeatEvaluator); FTS5 cross-session search (
conversation_search); cloud sleep/wake (CPsleep_sandbox); subagents (SubagentDashboard); domestic models via Provider GUI;/learnGUI workflow (Settings Learn wizard + 5 scenario chips + draft review +[use skill]invoke guide + server SSOT rewrite — Jul 2026: 35 FE vitest + server learn pytest). - Delivered (2026-07-31): Vision Fallback epic #1+#2+ABC ✅ — all ExecutionSurfaces + Settings chain health check + one-click recommend + chat gap CTA; 40 directed tests passed (19 harness + 5 health + 6 config_parsers + 10 FE vitest). Hermes sessions import → memory #3 still pending.
- Myrm ≫: GUI Migration Wizard (5 sources, dry-run/confirm/rollback) vs
- WORTH:NO from series: Python RPC subagent pattern (sandbox code-as-action ≥); Nous Portal bundled API; RL batch-trajectory research pipeline (not GUI-primary).
- Material SSOT:
temp-docs/materials/COMPETITIVE_HERMES_ADVANTAGE.md§30 ·temp-docs/roadmap/hermes_openclaw_series_summary_cross_reference_roadmap.md
Latest Validation Snapshot (2026-08-08 · Vision Agent Toolkit #2 sealed)
- Agent-active vision + video dual-slot parity (vs OpenClaw media runner / Hermes auxiliary / DeerFlow view_image gate): Harness vision 63/63 + converter→GeneralAgent video integration 2/2 + real-LLM video agent-stream 3/3 + frontend
visionCapability11/11 + Chrome Settings READ smokes 2/2 = 81+ tests, 0 failures. - Verified: 统一收敛至
file_read_tool多模态直接读取与自适应降级(VisionFallbackEngine) · Settings separatevideoFallbackModelslot +POST /vision-health·pick_video_fallback_model_cfgsend-to-end (converter → runtime context →file_read→video_reader) · Chat SSEvision_backendbadge (vlm / native_video / frame). - User-visible win vs competitors: OpenClaw 仅有被动媒体转译;Hermes 依赖 CLI 配置与散落辅助工具;Claude Code / DeepSeek Harness 缺乏纯文本主模型降级兜底 — Myrm 既保持
file_read_tool统一看图零 Prompt 污染,又通过VisionFallbackEngine保证纯文本模型无缝解析沙箱图像,兼备极致性能与普适性。 - Honest scope: No pre-upload H.264 transcode health probe; no vtracer SVG trace path; no Hermes-style zero-config provider auto-scan (by design — billing/compliance).
Latest Validation Snapshot (2026-08-12 · Model Failover Progress UI)
- Chat failover visibility (vs Hermes / OpenClaw / DeerFlow): Frontend Vitest 7/7 + Chrome E2E
test_model_failover_chrome_e2e.pyPASSED (~93s, Aug 2026). - Verified:
MODEL_FAILOVER+STATUS model_failoverSSE auto-create assistant placeholder before first MESSAGE chunk; ProgressSteps show backup-model switch (model_failover_*); Harness rebindsStreamContext.agentafter failover so reply continues on fallback LLM. - User-visible win: Competitors often silent-switch or toast-only; Myrm shows in-thread progress while the conversation keeps going.
Latest Validation Snapshot (2026-07-31 · Vision Fallback GUI closure)
- Vision auxiliary capacity (vs Hermes auxiliary_client / OpenClaw media runner):
./myrm test19+5+6 passed +bun run testvisionCapability/visionConfigGap 10 passed = 40 tests, 0 failures (~40s total, minimal directed set). - Verified: Ordered provider failover (402/rate-limit/overload; AUTH/MODEL_NOT_FOUND no switch) · all ExecutionSurfaces share
resolve_vision_fallback_chain_for_agent· Settings Test vision chain +resolved_model· Use this model recommend card · chat attach toast Go to Settings →/settings/models?sub=default. - Honest scope: No Hermes-style zero-config Provider auto-scan (by design — billing/compliance); no per-agent visionFallback slot (separate epic). (Chat model failover ProgressSteps — see 2026-08-12 snapshot above.)
Latest Validation Snapshot (2026-07-30 · Wiki Evidence Closure v2 #4+#5)
- Evidence citation GUI (vs OpenClaw / Hermes / deer-flow / LobsterAI / CoPaw / jiuwenclaw):
test_query_closure_v24/4 +test_best_first5/5 +test_source_citations6/6 +test_wiki_recall_benchmark_gate1/1 = 16 tests, 0 failures (~6s, 2 batches, harness-only, no Chrome MCP — Node unavailable in agent env; vitest not run this round). - Verified: Chat/Settings citation Drawer shows claim status + compile confidence + raw excerpt; fresh supported claims rank ahead of stale contested; Settings Query retrieval path trace (index/seeds/concepts); citation score derived (not hardcoded 1.0).
- User-visible win vs competitors: OpenClaw has CLI
raw-claim/source-evidence+ claim confidence rerank (query.test.ts:602) but no GUI Drawer; Hermes llm-wiki skill / deer-flow / LobsterAI / CoPaw / jiuwenclaw have no claim-level citation GUI; none expose Settings retrieval trace cards. - Honest scope note: No
route-questionmode yet (OpenClaw has it; frontmatter schema pending); Chat omits retrieval trace (Settings debug surface); no Trust card / SSE lineage (by design).
Latest Validation Snapshot (2026-07-30 · Wiki ↔ Memory boundary #27)
- Wiki vs Memory write boundary (vs Hermes / OpenClaw / Mem0 concept):
test_wiki_memory_boundary7/7 + save guard 3/3 + extract prompt 2/2 + agent wiring 1/1 = 25 tests, 0 failures (~5s, 3 batches, harness-only, no Chrome MCP — backend guard has no browser interaction path). - Verified:
memory_save_toolhard-rejects document-like knowledge/event when wiki is enabled (>800 chars or ≥3 markdown headings →wiki_ingest_tool);persist_extracted_memorieshard-filters the same heuristics on auto-extract semantic/episodic; extraction promptwiki_boundary_enabledwhen agent vault is active; Settings en/zh copy explains Wiki vs Memory roles. - User-visible win vs competitors: Hermes caps flat MEMORY.md at 2,200 characters (
hermes-agent/website/docs/user-guide/features/memory.md) — no compiled wiki vault or dual-path code enforcement; OpenClaw uses MEMORY.md files without vector auto-extract filter; Mem0 articulates Wiki≠MemCon separation — Myrm adds write-path enforcement plus compile pipeline + unifiedcorpus=allrecall. - Honest scope note: Splitting one long article into multiple short memory entries is not specially blocked; no standalone Boundary Card (ultimate review WORTH:NO).
Latest Validation Snapshot (2026-08-20 · Multi-root Topology & Drag-Drop Mount)
- Multi-root Workspace Zero-Friction Drag & Drop and Topology Awareness (vs Cursor / Windsurf / Claude Code / Cline):
- Test Verification: Backend
test_session_access_roots_prompt.py+test_session_access_roots_api.py+test_session_access_service.py13/13 passed; Harness security & middlewaretest_path_security.py+test_session_access.py+test_session_access_middleware.py47/47 passed; Frontend component & hookSessionAccessRootsBar.test.tsx+useDesktopFolderDrop.test.ts10/10 passed; full chain 70 automated tests, 0 failures. - Key Capabilities:
- Zero-Friction Drag & Drop Whitelisting: Native drag-and-drop on Desktop/WebUI instantly authorizes local directories with session persistence and 0 permission interruptions.
- Deterministic Topology Prompt Injection: System prompts format
[Mounted Workspace Directories]with deterministic alphabetical sorting, preserving 100% Prompt KV Cache hit rate while eliminating LLM path blindness. - Interactive Frontend Chip Suite: Hover to reveal full absolute paths, one-click copy to clipboard with toast notification, and instant access revocation.
- User Migration Payoff:
- vs Cursor / Windsurf: Eliminates heavy background Language Servers and File Watchers that cause CPU saturation and battery drain; Myrm uses lightweight topology awareness and direct tool invocation with 1/100th the resource footprint.
- vs Claude Code / Cline: Eliminates continuous CLI confirmation prompts and interruption of user focus across multi-repo microservice workflows.
- Test Verification: Backend
Latest Validation Snapshot (2026-07-31 · Vision Fallback GUI closure)
- Wiki external source sync GUI + Google Drive (vs OpenClaw / Hermes / deer-flow / LobsterAI / CoPaw / jiuwenclaw / OpenWiki): gdrive 3/3 + gmail/rss/status/oauth 12/12 + state/config/defaults/blueprint/hygiene/second_brain 15/15 + gmail_html 2/2 = 32 tests, 0 failures (~45s, 3 batches, server lane, no Chrome MCP).
- Verified: Settings External Sources panel (Gmail label + Google Drive folder + RSS + integration mirror → wiki
raw/viapublish_raw); zero-LLM pull; OAuth auto-enables GmailReadLater;google_drive_authorizedreconnect hint when Drive read scope missing; persisted sync state +syncIssueerror surfacing; Cron blueprintread_it_laterrouter job (__wiki_source_sync__, empty tools). - User-visible win vs competitors: Full GUI ingest → compile → search loop for saved mail, Drive docs, and feeds — OpenClaw routes Gmail through Pub/Sub chat hooks (not a wiki vault pipeline); Hermes/OpenClaw google-workspace CLI can read Drive but not into a compiled wiki raw pipeline + Settings GUI + cron; deer-flow / LobsterAI / CoPaw / jiuwenclaw have no equivalent; OpenWiki Gmail is env/config-driven with no Drive connector.
- Honest scope note: Gmail label and Drive folder ID typed manually — no label dropdown / Drive Picker (ultimate review WORTH:NO); no OneDrive (broken stub removed); Integrations API
drive_read_enabledsymmetry WORTH:NO (raw OAuth scope already exposed). Chrome MCP E2E not run; behavior proven via mocked Gmail/Drive API unit tests.
Latest Validation Snapshot (2026-07-27)
- Clickable deliverable paths in chat (vs OpenWorker Cowork / Hermes) (2026-08-03): Harness
short_file_idSSE 3/3 + Server processor/seed fixture 3/3 + live API seed 200; chat`workspace/...`/@file_NNN→ ArtifactPortal (DeliverableReferenceLink+openWorkspaceFileInPortalSSOT). OpenWorker needs[Title](artifact:path)+ RightRail; Hermes@file:is agent-internal only, not clickable in GUI. Chrome MCP E2E pending mux recovery (test_deliverable_link_chrome_e2e.py). - Office Document Pipeline (vs WPS Lingxi / OpenClaw / Hermes / DeerFlow / LobsterAI / CoPaw / jiuwenclaw): file_parsers 213/213 + document_reader 26/26 + docx read-chain E2E 16/16 + deliverable bundle 23/23 = 278 tests, 0 failures (2026-07-28 batched run). Verified: LegacyFormatParser OLE2 + soffice auto-conversion, docx.py cell_map structure metadata, full-chain Goal deliverable ZIP bundle. Write fidelity (in-place): Harness
OfficeBashAuditpost-bash OPC/formula diff + baseline-missing honest warn + corrupt Office package warn + optional LibreOffice recalc error scan + layout QA (18 pytest, 2026-07-28) — no OSS competitor has an equivalent harness-level post-bash audit. WPS Lingxi still leads on fixed government-template fidelity via proprietary kernel; Myrm leads on legacy format handling, structured parsing, delivery bundle, and 3-deployment independence. - SpreadsheetEditor Full-Fidelity GUI Editing (vs WorkBuddy editor_sdk): Formula roundtrip + style preservation + merge preservation + numFmt preservation + degradation warning banner + IndexedDB draft auto-persist & recovery — 41/41 Vitest tests, 0 failures (2026-08-03). Fixed P0 silent data corruption: previous
sheet_to_json/aoa_to_sheetpipeline silently dropped formulas, styles, merged cells, and number formats. Now iterates SheetJS cell objects directly withcellStyles:true. WorkBuddy has comparable fidelity viaeditor_sdk; Myrm now matches on roundtrip correctness and adds degradation warnings + crash-recovery drafts. - AI-Native Selection→Agent Architecture (2026-08-03 code-level verified): 3 of 5 artifact types have selection→AI interaction (Code via
SelectionToolbar, Document viaDocumentSelectionToolbar, HTML viaElementPickerToolbar), all sharinguseSelectionActionhook (dirtyArtifacts injection + Agent busy queueing + AgentBusyError fallback). 29/29 selection tests passed (useSelectionAction 8 + SelectionToolbar 17 + DocumentSelectionToolbar 4). All 6 competitors (OpenClaw, Hermes, DeerFlow, LobsterAI, CoPaw, jiuwenclaw) have zero GUI spreadsheet editors and zero selection→AI interaction — verified by grep across all competitor codebases with 0 matches. - Frontend Agent State Management Architecture (vs CopilotKit AG-UI useAgent/useAgentEvents Hooks): multiplexChunkBridge 16/16 + SSE handlers (gap+clarification+completion) 24/24 + messageStreamHandler core 27/27 + Zustand store (snapshot+navigation+subagent+model config) 30/30 + integration & event dispatch 39/39 + UI data model merge 6/6 + state recovery (budget+memory+reasoning) 40/40 + plan lifecycle 13/13 = 195 tests, 0 failures. Fixed 1 pre-existing bug (Japanese locale assertion out-of-sync).
- Why CopilotKit SDK Hooks are obsolete here: CopilotKit’s
useAgent()/useAgentEvents()target SDK embedding (Agent inside a third-party app). Myrm, as an independent AI product, uses Zustand store + 13 dedicated Handler modules +streamConsumer(with disconnect retry / Last-Event-ID stream resumption / chunk buffering / capability_gap deferred re-send / missed HITL recovery / historical state hydration) — 66 SSE event types vs CopilotKit’s ~10 abstract events, 6.6x granularity, precise re-render with zero Hook overhead. - Migration benefit: Developers migrating from CopilotKit AG-UI get real-time streaming, auto-reconnect, and full UI state machine out-of-the-box — no SDK hooks, EventSource management, or reconnection logic required.
- Declarative Agent UI Framework (vs CopilotKit Generative UI useRenderTool/A2UI): Frontend interactive-ui components 65/65 + Harness render_ui_tool + update_ui_data_tool 30/30 = 95 tests, 0 failures. Myrm ships 23 whitelisted components (10 form + 5 layout + 6 display + 3 basic) + 8 built-in validation rules + conditional rendering (visible binding) + UIComponentErrorBoundary per-component isolation + validate_ui_adjacency structure validation + update_ui_data_tool incremental deep merge — Agent outputs JSON schema, zero frontend code needed.
- Why CopilotKit Generative UI is obsolete here: CopilotKit requires frontend devs to hand-write a render callback for every tool. No built-in validation, error boundaries, or incremental updates. Myrm’s declarative registry model is superior in security, consistency, and development efficiency.
Latest Validation Snapshot (2026-07-26)
- Task Navigation & AI Auto-Routing (vs mattpocock/skills
ask-matt): DiscoverCapability 57/57 + Clarification/ask_question 15/15 + capability_gap SSE 37/37 + StreamDispatcher 17/17 + Kanban API 361/361 + Kanban Service 102/102 + Kanban Integration 33/33 + Kanban Channels 89/89 = 711 tests, 0 failures. - Verification Seam Gate (vs mattpocock/skills
/to-spec): Goal Engine (VerificationGatekeeper + ShellCriterion + SemanticCriterion + CompletionGuard + 熔断保护) 209/209 + PlanConfirmMiddleware 13/13 + Kanban Verifier + Criteria Integration 47/47 = 269 tests, 0 failures. - Pre-existing bug fixed:
kanban_command_handler.py:326—by_agentdict withNonekey causedTypeErrorin sorted(); now handles gracefully. - Why CLI navigation is obsolete here: Matt’s ask-matt requires users to manually choose a workflow path from 17 isolated CLI skills. Myrm’s 6-layer auto-routing (ActionMode + DiscoverCapability + capability_gap + ask_question + task-planning + KanbanPipeline) makes the user simply describe what they want — the Agent automatically plans and executes.
- Why manual
/to-specis obsolete here: Matt’s to-spec requires users to manually trigger seam definition. Myrm has 5-layer automated verification: GoalMode acceptance_criteria (user-defined) + PlanConfirm HITL (plan review) + Kanban completion_criteria (structured shell+semantic) + VerificationGatekeeper (auto-execution after task completion) + CompletionGuard (anti-hallucination). - User migration payoff: developers using Matt’s skills lose zero capability when switching to Myrm; they gain AI-automatic task routing and programmatic verification instead of memorizing slash commands.
- Real-time Task Monitoring (vs mattpocock/skills GTE Workbench concept): GoalStatusCard+GoalControlPlane+GoalPlanStepsList 27/27 + SubagentDashboard 2/2 + Kanban DnD+Markdown 48/48 + Goal Engine 164/164 + Subagent Engine 179/179 + Kanban API 250/250 + Kanban Integration 100/100 = 770 tests, 0 failures. Myrm provides 9-layer real-time monitoring (GoalStatusCard fixed overlay + GoalPlanStepsList + GoalControlPlane + SubagentDashboard + ArtifactPortal + BrowserLiveView + DesktopLiveView + KanbanBoardView + MobileStatusBoard) — all accessible at 0-1 click distance. Portal supports three layout modes: overlay (quick preview), side-by-side (wide screens auto-switch at ≥1280px, user-toggleable), and fullscreen (immersive editing) — with user preference persisted via localStorage.
- Intelligent Runtime HITL vs Task-level AFK/HITL Declarations (vs mattpocock/skills
/to-tickets): Kanban task_runner 25/25 + unattended_mode 2/2 + approval_flow+yolo 111/111 + unattended_mode_guard 2/2 + approval_edge_cases+batch+ptc 38/38 + rate_limiter+denial+subagent_safety 22/22 + interception+correction+scheduler 104/104 + security_config+execution_policy+permission 169/169 + server approvals 67/67 + personality_yolo+payload 19/19 + HITL_resume+kanban_binding 11/11 + subagent_approval_integration 4/4 + batch_decisions+session 72/72 + workspace_boundary 12/12 = 658 tests, 0 failures. Myrm’s design: Kanban=always AFK (yolo+unattended), Chat=natural HITL, Cron=natural AFK, GoalMode=intelligent switching. Runtime ToolApproval triggers per-operation (not per-task), with 4-level Allow-Always learning. Matt’s task-level AFK/HITL tag is a CLI limitation patch — it creates semantic conflicts (AFK task + dangerous op = ?) that Myrm’s operation-level approach solves elegantly. - Context Management & Smart Handoff (vs mattpocock/skills ask-matt Context Hygiene +
/handoff): ContextBudgetGuard 24/24 + ConversationForkManager 24/24 + ContextUsageIndicator 25/25 + Context pipeline 111/111 + Filter+cache_ttl_prune+resume 68/68 + Handoff API 24/24 + SummarizeProcessor 49/49 + Split Turn Prefix Summary 46/46 = 371 tests, 0 failures. Myrm provides 7-layer context management: (1) ContextBudgetGuard 4-layer overflow protection (2) 50+ file auto-compress pipeline (3) Split Turn active turn prefix summary (4) compactChat manual compress API (5) checkpoint-based fork at ≥75% usage (6) ContextUsageIndicator ring+CTA (7) HandoffDialog cross-channel migration. Matt’s/handoffis a manual CLI command producing text summaries that lose context; Myrm’s Fork preserves full LangGraph checkpoint state. Pre-existing test bug fixed: summarize_processor_NoOpMetricAttributeError.
Latest Validation Snapshot (2026-07-23)
- Shell pattern Allow-Always full chain: Harness pattern tests 20 passed; server SHPOIB + allowlist API 11 passed; frontend derive-pattern parity 13 passed; Chrome LIVE_AGENT E2E 1 passed (bash HITL → Allow always this pattern → Settings list/delete → next run auto-approved).
- User-visible win vs competitors: 4-level Allow-Always (permission / tool / exact / command pattern) with Settings CRUD — Claude Code & OpenClaw stop at tool-name or CLI signing; compound shell (
&&/pipe) never persisted as pattern. - Dev-gate reliability under real parallel load:
./myrm ready --chromepreflight +./myrm testfocused suites all passed (54 ready/install regressions + 32 dev-gate contract tests + 3 Chrome E2E flows: READ spill / READ expired / LIVE marketplace chat). In temporary shared-stack degradation windows, attach preflight recovered via built-in retry without requiring users to kill other running pytest jobs. - Migration payoff: teams moving from CLI-only pipelines get deterministic “ready → test → real Chrome” signoff in one workflow, with explicit health evidence (
clientHot, runtime IDs, lane-aware queueing) instead of ad-hoc shell coordination.
Latest Validation Snapshot (2026-07-24)
- TTFT startup path proof (minimal batch, low load): harness
test_backend_detector.py21 passed with focused module coverage 84% (backend_detector); server TTFT chain tests (test_stream_loop_ttft.py,test_stream_collector_coverage.py,test_usage_aggregation_coverage.py) 18 passed with focused coverage 48%-72% onstream_loop/stream_collector/usage_aggregation. - Resource envelope (measured, not estimated):
/usr/bin/time -lpeak RSS across these batches stayed around 90MB / 337MB / 144MB / 355MB; no monotonic memory climb observed during reruns. - User-facing takeaway: first-reply latency instrumentation now stays verifiable end-to-end while preserving startup stability of external-agent detection under realistic local constraints.
- Honest scope note: this 2026-07-24 batch intentionally avoided heavy browser E2E to keep CPU/memory impact low; latest real Chrome MCP evidence remains documented in the 2026-07-23 and 2026-07-20 snapshots above.
Latest Validation Snapshot (2026-07-20)
- Browser automation baseline (targeted, low-load batches): 109 harness tests passed (
snapshot, dispatcher/event routing, takeover). - Contract parity checks: server SSE event parity 1 passed; frontend event schema 6 passed.
- Quick interaction stability (focused regression): frontend
deep-link-listener+flow-pad-inline-mode+intent pagesuites passed 36/36; server gate/finish/integration suites passed 43/43. - Coverage evidence (focused files): frontend (
deep-link-listener.tsx,flow-pad-modal.tsx) 78.28% overall in focused run; server modulesdesktop_control/gate.py89%,background_job_finish_handler.py94%. - Real browser proof (not paper analysis): real Chrome MCP
take_snapshot+evaluate_scripton live WebUI pages; verified/intent/ask?text=...auto-returns home and opens FlowPad with prefilled text. - Chrome UI E2E:
test_background_tasks_panel_chrome_e2e.py3 passed. - Delegation control Chrome LIVE E2E:
test_subagent_dashboard_chrome_e2e.py3/3 (340s); session-scoped delegation pause + Dashboard cancel/token/model signed off on real Chrome. - Pre-existing bug fixed during validation:
mode=runningseed fixture returned 500 due toUnboundLocalError; fixed in API fixture route and covered by integration regression. - Resource envelope during verification: memory free percentage stayed in the 36%-40% range in this run (no high-pressure >80% used state observed).
- Known environment caveat (honest): captcha coordinator tests still depend on local
patchrightavailability; this does not block the browser snapshot/telemetry baseline verification above.
vs 六大竞品 — LLM 结构化 JSON 输出解析鲁棒性
一句话:模型把 JSON 夹在散文里、包进代码围栏、先吐一版草稿再给最终答案、字符串里塞裸换行——这些场景下,Myrm 的结构化流水线照样一次跑通,多数竞品要么抛错要么静默吐空。
逐竞品证据(2026-08 代码级):
- OpenClaw:生产代码直接
JSON.parse(src/talk/agent-run-control-shared.ts:267、src/state/openclaw-state-ownership.ts:79)——LLM 输出夹带一句话或一个 fence 即抛错。 - Hermes:
extract_json_candidate(tools/delegation_output_schema.py:80)剥离 fence + 首尾括号切片后裸json.loads,失败报错转重试;agent/jiter_preload.py只是 OpenAI SDK 的 Rust 加速解析器,不修复语法。 - deer-flow:backend 全仓
json.loads仅解析自有文件/消息数据(app/channels/*.py),没有面向 LLM 自由文本的 JSON 提取器。 - LobsterAI:
parseJsonObject(src/main/sessionDiagnostics/archive.ts:63)裸JSON.parse+ catch 返回 null,输出稍脏即丢。 - CoPaw:
extract_json_payload(plugins/apps/qwenpaw-creator/backend/utils/structured_output.py:59)fence + 平衡扫描但取第一个 JSON;json_repair仅用于配置文件(src/qwenpaw/config/utils.py:18)。 - jiuwenclaw:
extract_json_from_response(jiuwenswarm/server/hooks/executor.py:270)3 级回退(直接→正则 fence→{}切片)全失败返回{}静默吞错;json_repair仅 JSON5/特定场景。
require_key(只认含目标键的对象)天然抗草稿污染,skill 进化 / wiki 编译 / 记忆抽取 / judge 判卷全部受益。
Myrm 实现:myrm-agent-harness/src/myrm_agent_harness/utils/chat_utils.py — parse_llm_json_object / parse_llm_json_list → _iter_json_blocks 平衡块迭代 → _escape_control_chars_in_strings → _strip_trailing_commas。判卷另走 max_tokens=600 防思考截断。
At a Glance
vs PilotDeck — Agent Operating System
PilotDeck (open-sourced by Tsinghua THUNLP) positions itself as an “Agent OS”, featuring WorkSpace isolation, white-box memory (Dream Mode), smart model routing for cost reduction, and Always-on background execution.Where Myrm Goes Further
Migration from PilotDeck
Migrating from PilotDeck upgrades you from a “CLI-centric operating system” to a “GUI-first modern agent workstation”. You retain and surpass all its advanced features like multi-project isolation, cost-effective routing, and background execution—but everything becomes visual, draggable, rollback-ready, and seamlessly integrated with the vast MCP plugin ecosystem.vs 360 Security OpenClaw (Enterprise Wrapper)
360 Security OpenClaw is a closed-source enterprise wrapper built around the OpenClaw core. It aims to lower the barrier to entry with a “Shrimp Coach” (guided setup) and manual “Token Cost Modes” (Lightweight, Economy, Full-power).Where Myrm Goes Further
Result: Myrm delivers a more automated, native GUI experience. Instead of a “coach” asking questions, Myrm provides ready-to-use templates. Dual-phase routing is more token-efficient and accurate than competitors’ pure LLM-judge approach, with 10x more cost transparency.
vs OpenClacky — Token-Optimized Local Agent
OpenClacky markets itself as a cost-efficient local AI agent, claiming 1/6 the token cost of Hermes. Its core strategies: 16 minimal tools, idle compression, Insert-then-Compress, dual cache marking, and BYOK multi-model routing.Where Myrm Goes Further
Key Architectural Differences
Compression cost: OpenClacky calls LLM to generate compression summaries — every compression incurs API cost. Myrm uses a deterministic rule engine (Dedup → Truncate → Remove) with zero LLM overhead.vs Tokenomics Lite — Per-Turn Invalid Token Cleanup (2026-08-11 · Article 53)
Tokenomics (closed-source blog series, no GitHub) describes a Lite engine as layer 3 of a 12-engine stack: per-turn cleanup of whitespace in JSON/Markdown, data-URI removal, and external image URL →[image] placeholders — claimed 8–15% savings per turn, <1ms, on by default.
Where Myrm Already Goes Further
Honest boundary: Lite’s sub-ms JSON minify has marginal ROI once Myrm’s dedup + semantic compression + archive pipeline is active. No separate “Lite mode” productization needed — the harness 11-step context pipeline covers and exceeds it on Local, Tauri, and Cloud.
Migration pitch: If you liked Tokenomics Lite’s “clean invalid tokens every turn” idea, Myrm ships a deeper multi-layer equivalent by default — plus Turn1 token CI gates, cache observability, and zero-LLM rule compression that competitors charge LLM tokens for.
vs WorkBuddy / Codex / QoderWork — Interactive Study Apps (2026-08-11 · Article 52)
Community benchmark: same bare prompt (formula sheet → junior-high physics review web app). Ranking WorkBuddy Plan > Codex Goal > QoderWork Chat (63 comments, closed source).Where Myrm Is Stronger (Engine)
Where Myrm Is Weaker Today (Product UX — honest)
Migration pitch (honest):
- From Codex Goal: You keep Goal + acceptance + preview; gain BYOK, context pipeline, and budget transparency. Tutoring page polish catches up via #168 preset — not a blocker for power users today.
- From WorkBuddy Plan: Engine ≥; one-click “Plan → starred difficulty review app” UX is the gap (#89 + #168). Use Goal + acceptance_criteria + HtmlPreview until preset ships.
- From QoderWork: Strict upgrade on preview, verification, memory, channels, and sandbox.
CacheKeepAliveManager sends lightweight probes every 4 minutes during idle to prevent the Anthropic/Qwen 5-minute cache TTL from expiring, ensuring consistent TTFT when users resume (0.5-1s vs 2-5s cold restart).
Migration from OpenClacky to Myrm
Result: 8 upgrades, 0 equivalent, 0 downgrades. OpenClacky users gain zero-cost compression, intelligent cache protection, persistent cross-session memory, and a complete GUI workspace while maintaining all token efficiency benefits.
Unified Tool Gateway & Flexible BYOK (vs Hermes / OpenClaw)
While competitors often force users to manage dozens of API keys or rely on rigid, error-prone gateways, Myrm introduces a 4-in-1 Unified Tool Gateway with an Elastic BYOK Fallback mechanism.Where Myrm Goes Further
- 4-in-1 Gateway: One subscription unlocks LLM, Web Search, Image Gen, and TTS capabilities. Zero configuration required out of the box.
- Elastic Try-Catch Fallback: If the official gateway is unavailable or your Work Units (WU) run out, Myrm automatically and seamlessly degrades to your locally configured API keys. Your business never stops.
- Quota Rollover: Unlike traditional SaaS where unused quotas expire, Myrm allows unused Work Units to roll over to the next month.
- Visual Gateway Status: A dedicated GUI dashboard shows the status of all tools (Gateway Managed, Custom Key, or Unconfigured), with smart prompts for Pro users to enable gateway features.
- Financial-Grade Security: PAT tokens are transmitted via secure POST bodies and validated with strict regex rules, preventing SSRF attacks and guaranteeing tokens never leak into URL logs.
- Extreme Network Resilience: Built with imperative on-demand polling and an 8-second hard timeout using
AbortController, eliminating the UI freezes and API quota drain common in competitors’ auto-polling dashboards.
Competitor Pain Points
- Hermes: Uses a “black and white” hard switch for its gateway. If the gateway fails or runs out of credits, the agent simply crashes and stops working.
- OpenClaw: Relies entirely on user-provided keys via extensions, offering no unified gateway or billing.
- Cloud Browser Dependency: Hermes relies on third-party cloud browsers (like Browser Use) for web tasks, introducing high latency and privacy risks. Myrm insists on local/self-hosted sandboxes for browser automation, ensuring zero latency and total data privacy.
Migration Wins
- Zero Key Management: Stop juggling API keys for OpenAI, Firecrawl, FAL, etc.
- Peace of Mind: The elastic fallback means you get the convenience of a managed gateway with the reliability of your own backup keys.
Tool-Call Lineage DAG & Step-Level Security Verdicts (vs Claude Code / Hermes / OpenClaw / DeerFlow)
Competitors treat tool calls as a black box: when the agent edits a file, runs a command, or hits the network, you cannot tell which instruction caused it, whether the safety system reviewed it, or how it fits into the overall execution. Myrm ships a full tool-call instruction-lineage DAG with step-level security verdicts attached.Where Myrm Goes Further
- Precise instruction → tool-call lineage: every
tool_start/tool_endis paired by a uniquetool_call_id— concurrent same-name tools (e.g. two parallelbashcalls) never cross-wire, and legacy events without an id fall back safely. Real LLM streamingtasks_stepsevents merge into the same lineage by id, so nothing is lost between the lifecycle events and the live progress stream. - Step-level security verdicts: every
security_auditdecision (allow / deny / risk-flagged, with reason and timestamp) is anchored to the exact tool call and surfaced on the execution timeline. ADENYor tainted step is painted rose in session replay; approved steps carry their reason too. No decision is ever silently dropped — an audit without a matching tool call logs a warning instead of disappearing. - Accurate success semantics: a step that completed normally is
success=true— competitors often paint “completed” steps as failures or conflate terminal status with errors. Failed steps keep their error detail and are painted red with the reason in the Inspector. - Session replay with verdicts: scrub any past session and see every sensitive action — which instruction triggered it, what it did, and the security verdict with its reason. In-product, no external SaaS, no extra fees.
- No internal leakage: internal pairing metadata (
tool_call_id,message_id) never leaks into the displayed tool input — the Inspector shows the real arguments, not framework plumbing. Verified end-to-end with a real-LLM Chrome E2E (2026-08-15: harness 187 passed, server integration 16 passed, frontend 10 passed, Chrome E2E 1 passed in 283.22s).
Competitor Pain Points
- Claude Code: conversational resume-level lineage only — no per-step tool-call lineage, no security verdicts attached to individual steps.
- Hermes / OpenClaw / DeerFlow: no tool-call lineage DAG, no step-level security label propagation, no verdict visualization in any UI.
- All competitors: no way to answer “which instruction caused this sensitive operation and did the safety system review it?” from a replay timeline.
Migration Wins
- Auditability you can see: replace “trust the logs” with a replay timeline that shows every sensitive action, its cause, and its verdict.
- Faster incident review: when something goes wrong, pinpoint the exact instruction → tool call → decision chain instead of grepping logs.
- Enterprise-ready: step-level verdicts attached to the trace make compliance review straightforward — no separate audit tool to maintain.
Organization-Wide Agent Audit Timeline (vs Claude / Hermes / OpenClaw / any single-sandbox audit)
When agents run across many team sandboxes, competitors give organizations no answer: no cross-sandbox ledger, no way to attribute a sensitive action to the right member or sandbox, and failed scans are silently skipped — so “audit coverage” is whatever happened to get collected, with gaps you cannot see. Myrm ships an org-wide audit aggregation where every member sandbox reports into one unified timeline.Where Myrm Goes Further
- Fan-out aggregation across all member sandboxes: one query reaches every sandbox’s local event ledger and merges tool calls, approvals, and security decisions into a single reverse-chronological timeline — attribution to member and sandbox included. Two data layers (
security_auditevents +tool_start/tool_endlifecycle) feed the same KPI set, so counts stay consistent end-to-end. - Exact security-deny accounting, not a vibe:
security_deny_totalcounts only true deny decisions (BLOCK/DENY/REDACT/LEAK) — the measurement can’t be inflated byALLOWdecisions or by events that were never reviewed. The same deny predicate is enforced on the server aggregation layer and re-validated in the frontend rendering, so a red “security blocks” KPI always reflects a real denial. - Scan failures are surfaced, never silently skipped: an
ACTIVEsandbox without a container is explicitly reported as failed instead of vanishing from the audit — coverage gaps are visible so you never mistake “not collected” for “nothing happened”. The same honesty applies to event grading: an audit event that can’t be matched to a tool call logs a warning rather than disappearing. - 1-N sandbox coverage regardless of status: audit coverage is reported per sandbox, with unscannable or failed sandboxes counted separately — the audit view always tells you how complete the picture is.
Competitor Pain Points
- Claude / Hermes / OpenClaw: per-sandbox logs only — an admin has no org view, no cross-sandbox aggregation, and no way to see which member’s sandbox produced a sensitive action.
- All competitors: deny counting is a per-session afterthought — no org-wide
BLOCK/DENY/REDACT/LEAKtotal, and failed sandbox scans drop out of the picture with no warning. - SaaS bolt-ons (external SIEM/audit product): extra cost, data leaves your sandbox boundary, and per-session granularity is lost in aggregation.
Migration Wins
- One dashboard instead of N log bundles: admins see the whole org’s agent activity on a single timeline with member + sandbox attribution.
- Prove coverage, not just activity: failed and unscannable sandboxes are explicit — you know when the audit picture is incomplete.
- Compliance-ready deny accounting: exact org-wide deny counts with reasons are exportable and attributable per member and sandbox.
WebUI Security & Local/Remote Separation
Most competitor web interfaces are either dangerously exposed when bridged to the public internet (LAN/Tunnel naked exposure) or overly burdensome for single-user local development (forcing DBs and logins forlocalhost).
Myrm solves the “last mile” of local-first deployment with an Ockham’s Razor approach to WebUI security.
Where Myrm Leads
- Zero-Friction Local vs. Ironclad Remote: Myrm’s WebUI auto-detects
localhostloopbacks to bypass login seamlessly. The moment you expose it via a tunnel or LAN, it enforces strict password protection. - Zero-Trust WebSocket Gateway: Competitors often secure their HTTP routes but leave WebSockets (used for realtime logs or voice) unauthenticated. Myrm’s
WsAuthMiddlewarephysically drops the WebSocket handshake (HTTP 403) if the session cookie is missing or invalid. - 3-Layer CORS Defense-in-Depth: Scheme+host dual input validation (wildcard
*rejected outright) + WebSocket Origin Guard (unified security boundary reusing CORS config, 4003 rejection for unauthorized origins) + HostAllowlist DNS rebinding protection (dynamic ingress/tunnel hostname allowlist). Native Tauritauri://localhostscheme support. Cloud deployment default-deny. 29 CORS security tests verified. - Instant Session Eviction: When a password is changed or protection is toggled, Myrm automatically rotates the global HMAC Session Signing Key. This instantly invalidates all existing sessions globally (like a kill switch for stolen devices), whereas competitors often leave old JWTs valid.
- Single-Tenant “Vault” vs. Heavy DBs: Instead of a bloated Postgres/MySQL user table with RBAC rules designed for SaaS, the local mode uses a single highly-encrypted
admin.jsonvault file. Secure, portable, and zero-dependency.
Data Loss Prevention (DLP) & Privacy-Aware Routing
Agent conversations may contain sensitive personal information. Competitors (e.g., ClawVault) only perform simple “detect → placeholder → restore” at the chat message level, missing tool parameters, tool results, and streaming output as leakage channels.How Myrm Leads
2,033 offline/privacy/security tests passed (verified 2026-07-14), covering PrivacyRouting (44) + Hardware API & deploy modes (48) + LLM Fallback degradation (153) + full security suite (1,766) + OfflineGuardian notifications (22). vs QoderWork’s “offline mode” which only binds Qwen local edition — no privacy routing, no hardware recommendations, no circuit breaker degradation. Myrm leads across all 9 offline/sovereignty dimensions.
Multi-Agent Orchestration: Deterministic Scheduling (vs Hermes / OpenClaw / JiuwenSwarm)
While competitors rely on single-mode orchestration (e.g., Hermes’ Kanban Swarm, OpenClaw’s basic linear spawn, or JiuwenSwarm’s static DAG workflow scripts), Myrm introduces an 8-Mode Deterministic Orchestration Engine backed by a 5-Layer Tool Security Fence and 4-Dimension Budget Control. Unlike JiuwenSwarm’s CLI-onlyworkflow.py approach (which requires users to write Python scripts to define execution order), Myrm’s Leader Agent dynamically schedules sub-agents through natural language — with runtime steer/cancel capabilities that static DAGs cannot provide.
Where Myrm Goes Further
Result: Myrm replaces the “prompt and pray” multi-agent paradigm with deterministic software engineering patterns. Your multi-agent pipelines execute reliably, securely, and within budget every single time. 1,680 orchestration tests passed, 0 failures (Jul 2026).
:::tip vs JiuwenSwarm — Natural Language Over workflow.py
JiuwenSwarm’s Plan/Performance/Swarm modes and Swarmflow often rely on pre-written
workflow.py static DAGs, with CLI/TUI as the primary experience (Web on :5173 is secondary). Myrm replaces script orchestration with Leader natural language + an 8-mode deterministic engine (including Swarm Fission / DAG / Council), plus runtime steer/cancel, batch/council cost preflight, and AgentWorkMap GUI across Local/Tauri/cloud. Jiuwen Skill Hub / team evolution maps to TemplateMarket + 8-source Skill Market + 4D evolution GUI (338+1008 tests). Migration blockers: 0 (Jul 2026 README 22-point seal).
:::
:::tip vs “Universal Agent” Narrative — Dan Shipper Context Rot
Dan Shipper (Every CEO) argues context rot makes a single do-everything agent degrade in quality; specialization and hierarchy remain essential. Myrm matches this philosophy: single-agent default (zero orchestration noise for simple tasks) + on-demand specialist teams (TemplateMarket + 8-mode engine) + fork-isolated context (11-step pipeline against rot). You stay editor/manager/producer — Goal/Kanban/IFM for final decisions. 0 new Roadmap items (Jul 2026 narrative seal).
:::
:::tip vs Doubao Search — Agent-Ready Structured Hits + Multi-Engine Failover
Volcengine Doubao Search returns authority tier, publish time, query-focused summary, and Markdown excerpts for agents—not just URLs. Myrm ships a Volcengine Provider MVP (Settings-selectable · NeedSummary long summaries) and is stronger overall: 11-engine priority chain (auto hop on quota exhaustion) + RSG sufficiency guard + Completion Guard freshness gate + all-model grounding rules + domain diversity ranking. Doubao migration has 6 pending items (field schema / explicit params / onboarding / cross-verify / Sources drawer / skip fetch)—see Roadmap v5. No OpenClaw CLI skill install required.
:::
:::tip vs Hermes “10x” — Memory·Skill·Cron Compounding Loop
One’s usage guide: Memory for facts · Skills for procedures · Cron for schedule · Self-verify against bad automation. Myrm is ≥ on all engines: structured memory (not flat MEMORY.md) · Settings /learn wizard · Cron acceptance_criteria + VerificationGatekeeper · SkillCurator+Pin · /journey GrowthDashboard · Wiki ≥ Obsidian. MSC Gate Closure (2026-07-31): Settings “Compounding Loop” four-row checklist · Cron 6-field create audit (real pause→confirm→resume gate · run-history acceptance editor · /settings/cron?job= deep link) · in-chat audit card. Hermes migration blockers: 0. No VPS+CLI — cloud sandbox + IM on three deployments.
:::
:::tip vs AWS Codex Agent Team — Code-Enforced Orchestration
AWS’s sample-codex-agent-team uses 6 Prompt-only SKILL.md templates (brainstorm → spec → coordination → review → documentation → project management) to guide multi-agent collaboration — entirely relying on LLM compliance. Myrm replaces every step with code-enforced primitives: run_council for multi-expert brainstorming with cross-review, execute_dag_plan for deterministic task DAGs, run_with_verification for adversarial verification with automatic retries, and 31-component Kanban GUI for real-time project management. None of these can be bypassed by the LLM.
Their 12-field “Handoff Contract” (role, objective, file scope, acceptance criteria, verification command, etc.) is also purely Prompt-based. Myrm enforces all of these through DelegateTaskInput Pydantic schema + 8 code-level checks (payload-hash dedup, depth/capacity limits, type allowlist, delegation pause, verifier mode) — plus 4 mechanisms competitors lack entirely. 2,488 tests passed, 0 failures.
:::
Smart Concurrency Router — Eliminating Read-Write Races (vs Hermes / OpenClaw)
In a highly concurrent Agent sandbox, an LLM often hallucinates operations that can corrupt data, such as trying to “read file X and write to file X simultaneously” in a single parallel tool call batch. In competitor architectures, this causes dirty reads and catastrophic write overwrites. Myrm introduces the Smart Concurrency Router, which actively intercepts and re-routes conflicting operations at the middleware layer.Where Myrm Goes Further
Result: Myrm completely eliminates file-level race conditions caused by LLM hallucination. Conflicting batches are gracefully unrolled into a sequential queue with zero performance penalty on unrelated parallel tasks.
Headless Agent: Zero-Deadlock Background Tasks (vs Hermes / OpenClaw)
In headless environments like SaaS scheduled jobs (Cron), batch processing, or background automation, an Agent runs without human supervision. If the LLM hallucinates and decides to call a human-in-the-loop (HITL) tool (like asking a question or rendering a UI form), competitors’ architectures will hang indefinitely waiting for user input that will never arrive. This causes catastrophic deadlocks and burns massive compute resources. Myrm introduces a robust Tag-Based Environment Degradation architecture at the lowest framework layer.Where Myrm Goes Further
Result: Myrm ensures bullet-proof reliability for unattended workflows. By physically stripping interactive tools from the LLM’s context during background tasks, it eliminates the root cause of Cron job deadlocks before the LLM can even attempt to make a mistake.
Additionally, Myrm’s Stuck Task Watchdog detects agent tasks that hang at the infrastructure level (e.g., LLM API half-open connections, tool deadlocks) and automatically cancels them after a configurable timeout (default 600s). The watchdog runs inside the existing 60-second janitor cycle with zero new components, sends a localized timeout notification to the user (5 languages), and releases the session so subsequent messages can be processed. Competitors like OpenClaw require container-level
killContainer (harsh), while Hermes relies on manual /stop — Myrm handles it transparently.
Extreme Scenario Anti-Explosion (vs Hermes / OpenClaw)
In autonomous multimodal scenarios (e.g., Computer Use) or very long sessions, agents inevitably hit the limits of context windows, causing frequent OOM crashes, dropped API connections, or total memory loss in competitors. Myrm implements a 4-layer Extreme Anti-Explosion Moat, ensuring that your agent never crashes and never loses the conversation context.Where Myrm Goes Further
Result: Myrm ensures a silky-smooth experience even under extreme loads. You save massive amounts of tokens by stripping historical images, and you never have to worry about the agent suddenly dying and wiping your hard work.
Web Search + Web Fetch — Dual Engine (vs Hermes / OpenClaw / Claude Code)
Myrm ships web_search and web_fetch as first-class built-in tools. Competitors either lack them, pass through raw API results, or charge per fetch via cloud APIs.What You Get
Competitor Pain Points
:::note Honest boundary
Hermes
web_extract appears “zero-config” because it uses LLM summarization instead of local embedding — but that costs tokens per page. Myrm’s web_fetch works locally with DOM pruning out of the box; fetch_and_extract adds vector+Reranker when configured. SSRF protections are comparable — Myrm’s edge is the filter pipeline, not basic security.
:::
Migration Wins
See Web Search & Fetch Guide for setup and configuration.
Citation Tracing & Source Display
Every citation in AI responses is traceable to its original source. Supports 4 source types (web search, MCP tool calls, knowledge base documents, conversation history) with hover preview showing full snippet, domain, and favicon. Touch devices automatically switch to tap mode.Document Parsing Engine
Myrm includes an 11-module professional document parsing engine with deep parsing from PDF to Office to Jupyter Notebook formats (184 tests verified):vs Hermes Agent (v0.15 Velocity) — Multi-Agent Platform
Hermes v0.15 (Velocity Release, 2026-05-28) refactored its core loop from 16k to 3.8k lines across 14 modules, added Kanban Swarm orchestration (104 PRs), rewrote session_search to be LLM-free (~20ms, from ~30s), and introduced Promptware defense (3 chokepoints with ~15 Brainworm/C2 patterns). A significant “pay down tech debt” release.Where Myrm Leads
- Event-driven Kanban vs Hermes’ polling — instant task dispatch with heartbeat + zombie detection + 3-layer anti-conflict (atomic claim + ownership enforcement + assignment audit trail, 817 tests)
- 6-rule diagnostic engine with severity auto-escalation (warning→error→critical) — detects stranded tasks, repeated failures, stuck blocked, dead dependencies, triage stalls, and block→unblock cycling (O(1) per-card evaluation vs Hermes’ O(N) event scan)
- 8 orchestration modes (Spawn/Chain/Batch/DAG/Verified/Swarm Fission/Alternatives Race/Council Debate) vs v0.15’s single Swarm pattern
- CompletionGuard with physical evidence verification — Hermes has no completion verifier
- Pipeline template wizard with discovery questions + role auto-matching — Hermes’
hermes kanban swarmis a fixed CLI command - 35+ messaging channels (25 with output hints) vs ~23 channels (19 with platform hints) — Myrm covers DingTalk, Teams, GoogleChat, LINE, IRC, iMessage, Voice, Webhook, Zalo that Hermes lacks
- 8 memory types with knowledge graph vs 2,200-character flat memory (MEMORY.md + USER.md with § delimiter)
- Web Search + Web Fetch dual engine — 8 search engines with BM25/Reranker filtering + 3-tier local fetch with DOM pruning ($0/month) vs Hermes’ cloud API passthrough + Firecrawl/LLM web_extract
- FTS5 + Qdrant hybrid search with scope/lineage filtering — Hermes v0.15 rewrote to local-only text search (~20ms, but no semantic retrieval)
- 108-pattern security scanning vs Hermes threat_patterns.py (~20 patterns in 3 scopes)
- 22+ middleware pipeline vs 8-layer fixed stack — Myrm’s middleware architecture is independently configurable per component
- 4-layer model discipline (CORE→ENFORCEMENT→FAMILY→ESCALATION) vs Hermes’ model-gated text blocks — Myrm includes ESCALATION_CONTRACT for automatic model self-upgrade
- GUI-first with Tauri desktop vs CLI-only (v0.15 added TUI multi-session, still terminal-bound)
- 4D budget control (Token + USD + time + max descendants) — Hermes v0.15 added per-task timeout only (1 dimension)
- 3-level intelligent model routing — Hermes v0.15 allows manual per-task model selection (no auto-routing)
- 6-layer Prompt Cache vs Hermes’
system_and_3strategy (4 breakpoints, Anthropic-only) — Myrm supports 5+ providers with cache break detection and anti-thrashing - Cache-friendly memory injection: Myrm splits memory into Stable (SystemMessage, cache-safe) + Learned (HumanMessage with UNTRUSTED_DATA isolation, never pollutes System Prompt). Hermes claims “User Message injection” in their article but their code (
system_prompt.py:424-486) merges memory into System Prompt volatile layer — breaking cache on every memory update. 429 cache-specific tests verified - IP persona protection without extra cost: User-configured
system_promptis injected viaUserInstructionsMiddlewareatpriority="highest"with[ABSOLUTE OBEDIENCE OVERRIDE], ensuring the LLM strictly follows IP identity facts. Hermes relies on plain-textSOUL.mdwith no priority enforcement. AI Builder auto-generates persona definitions including anti-fabrication instructions — 765 middleware tests verified - 4-layer workspace context injection:
workspace_rules_middlewareauto-discovers project rules from 15 file formats (.myrm.md,AGENTS.md,CLAUDE.md,.cursorrules,.clinerules,.windsurfrules,SOUL.md,MEMORY.md,.cursor/rules/*.mdc,.claude/CLAUDE.md,.github/copilot-instructions.mdand more) and injects them as<workspace_context>at layer [2] of the 4-layer stable prefix architecture. Includes prompt injection detection (scan_inputscore≥0.8 blocks), invisible Unicode stripping, YAML frontmatter cleanup, inode deduplication, and head/tail truncation within a 20K char budget. Each layer is per-scope (cross-user/per-user/per-workspace) for maximum Prompt Cache efficiency. Hermes/WorkBuddy require manualMEMORY.md+SOUL.mdfile creation with zero priority control — 17 integration tests verified
Skill Evolution — True Self-Improving Agent
Myrm’s Skill Evolution System (42 modules, native built-in) implements all engineerable concepts from Self-Improving Agent research. Hermes’ “self-evolution” is an external CLI wrapper around a third-party AGPL-3.0 tool (darwinian_evolver), not a native capability.
vs ECC (Everything Claude Code)
continuous-learning-v2: ECC distills atomic “instinct” YAML habits in the background — powerful for power users, but their v2.1 project scoping relies on a simple 2-level system (git remote hash → project vs global), with auto-promotion when an instinct appears in 2+ projects (false-positive risk). Myrm ships the same idea (idle distillation → skill proposals) with architecturally superior isolation: 5-level scope hierarchy (GLOBAL > AGENT > CHANNEL > CONVERSATION > TASK) + namespace system + per-agent CoW evolution. The Insights Inbox in Agent settings ensures proposals stay drafts until you approve; dismiss persists a negative exemplar so the agent stops re-proposing; built-in agents are read-only in the inbox. Verified with 3,519 tests covering the full continuous learning pipeline including skill injection, model discipline, evolution, search, and security. For MCP, Myrm now ships GUI pre-enable scan + verify + runtime fail-closed (see MCP Security Gate below) — ahead of Hermes log-only flows; we still do not ship ECC’s /aside fork chats, /context-budget tree-map, or 102-hook full static packs (roadmap #6).
Skill Module Architecture — 8-Dimension Advantage
Beyond evolution, Myrm’s skill system architecture is fundamentally more sophisticated across all dimensions:Skill Ecosystem Discovery & Import — 5-Source Aggregation vs Single Source
While competitors scan a single local directory (.claude/skills/, .agents/skills/), Myrm aggregates skills from 7 parallel discovery sources (including ModelScope 80K+ and Aliyun AgentExplorer) with a full GUI import pipeline:
Result: 9 upgrades, 0 equivalent, 0 downgrades. Myrm treats skill discovery as a first-class GUI experience with enterprise-grade safety and deterministic error contracts across preview/confirm, while competitors remain limited to local directory scanning.
Evolution Validation Pipeline — 5-Layer vs Zero Validation
Competitor proposals suggest users write manualvalidator/*.md test cases for skills — a fundamentally flawed concept for SOP prompt skills that adds high friction. Myrm’s 5-layer automatic validation pipeline provides comprehensive protection with zero user effort:
Key insight: Most skills are Markdown SOP prompts (e.g., “When reviewing code, follow these steps…”). You cannot write deterministic test cases for prompt guidance — our 5-layer pipeline solves this elegantly through complementary validation types. 3,519 tests verified.
Evolution Transparency — 6 GUI Panels vs Flat File
Competitor proposals suggest syncing DB records to aHISTORY.md file in each skill directory for “transparency.” Myrm provides 6 dedicated GUI panels with rich interactivity that makes flat text files obsolete:
Negative memory: 3 DB tables (
evolution_constraints + evolution_rejections + execution_analyses) automatically learn from failures. The engine injects past constraints into variant generation prompts via get_evolution_constraints(), preventing repeat mistakes — a closed-loop system that requires zero user maintenance. 332 additional tests verified.
Bounded Edit Control — 6-Layer Soft Constraints vs Hard Circuit-Break
Competitor proposals enforce a rigid “max 1 change point” rule with AST/Diff circuit-breaking. This is fundamentally flawed: a single bug fix often legitimately touches 2-3 related sections, and Markdown has no reliable AST for counting “change points.” Myrm uses 6 layers of soft constraints that maintain fix quality without artificial limits:
Combined with hard threshold rejection (score < 0.6 or accuracy < 0.7 = auto-reject), this system achieves bounded editing without sacrificing fix completeness. Monaco DiffEditor provides line-level red/green change highlighting — more precise than the competitor’s “highlight a single section.” 1783 tests verified.
Framework-Driven Engine — Python Framework + 5-Layer Customization vs Pure-Prompt Files
Some competitors drive skill evolution entirely through editable Markdown prompt files (e.g.,reflect.skill.md, edit.skill.md). While this appears flexible, it’s fundamentally fragile — one wrong edit crashes the entire pipeline, it doesn’t work in SaaS deployments, and creates dual-source-of-truth conflicts.
Myrm uses a Python framework engine with 7 modular prompt components hardcoded for safety, combined with a 5-layer structured customization system:
This gives users structured, safe customization without the risks of raw prompt file editing. Works identically across Local, Tauri, and SaaS — no filesystem dependency. 1923 tests verified.
Evolution Visualization — 9-Panel Full Lifecycle vs Command-Line Only
Some competitors rely entirely on Cron jobs and command-line output for skill evolution monitoring. Myrm provides 9 dedicated visualization surfaces covering the full evolution lifecycle:
No command-line output to parse, no log files to tail. Everything is visual, interactive, and actionable. 2900 tests verified.
Cross-Skill Global Rules — 3-Layer Organic Learning vs Auto-Writing AGENTS.md
Some competitors propose auto-injecting rules intoAGENTS.md or system prompts when common failure patterns appear across skills. This approach is architecturally dangerous: it modifies user-managed workspace files, creates dual source-of-truth conflicts, and breaks Prompt Cache efficiency.
Myrm uses a 3-layer organic learning system that achieves the same goal safely:
Key advantages over auto-write approaches:
- Safe: Never modifies user files or system prompts — learned rules flow through MemoryContextMiddleware
- Cache-friendly: Learned rules go into HumanMessage position, not System Prompt, preserving KV Cache
- Scoped: Rules can be global or tool-specific (
tool_namefield), not just blanket global injection - Secure: All learned data passes through
sanitize()+wrap_untrusted()boundary markers - Three-platform: Works identically on Local, Tauri, and SaaS — no filesystem dependency
Multi-Agent Skill Binding — DB-Level 3-Layer Isolation vs Config File Management
Some competitors use CLI config files to bind skills to agents. Myrm’s approach is fundamentally stronger with a 3-layer binding architecture:
Additional advantages:
- 8 orchestration patterns (Spawn / Chain / Batch / DAG / Verified / Swarm Fission / Alternatives Race / Council Debate) vs competitor’s single pattern
- GUI-first management: AgentEditPanel with 12 config tabs vs CLI-only
- ProfileTimeMachine: Config snapshot versioning with GUI rollback vs migration scripts
- Multi-competitor import: One-click migration from OpenClaw, Hermes, Cursor, Codex, Claude Code, ChatGPT, gbrain, and Pi (Windsurf/Trae on the roadmap)
- Channel skill triggers: Slash command bindings via
ChannelSkillCommandHandler
Migration from Hermes (Skill System)
Result: 7 upgrades, 1 tradeoff, 0 downgrades. Users gain intelligent security, secure config management, and organized skill lifecycle with zero capability loss.
Smart Context Archive References — Content-Addressed Storage vs Simple Reference IDs
Some competitors (like AgenticX) propose assigning simple “Reference IDs” to large tool outputs, immediately replacing content with an ID. This has a fundamental flaw: the LLM must process the content at least once, so immediate replacement forces repeated full restores. Myrm’s ContextArchiveReference system takes a smarter approach:
298 tests verified across archive reference, cache TTL prune, compactor offload, compress pipeline, cache metrics, and context budget.
Streaming Resilience & LLM Infrastructure — Enterprise-Grade vs Basic Retry
Some competitors propose simple “retry on truncation” or basic multi-key rotation. Myrm’s infrastructure goes far deeper: Streaming Truncation Recovery —StreamTruncationRecoveryMixin handles output interruptions transparently:
End-to-End Event Propagation — 48+ structured SSE event types vs competitors’ 2–3 basic signals:
Multi-Key Hot Failover —
KeyPoolLLM + CredentialPool:
OpenAPI Bridge — direct tool generation vs MCP intermediary:
978 tests verified: Auth middleware 54 + Streaming recovery 481 + Key pool 21 + OpenAPI bridge 84 + Core events 25 + LLM infra (fallback/routing/consensus) 313.
vs MiniMax Mavis — Multi-Agent Team Platform
MiniMax Mavis is a closed-source SaaS multi-agent system with Leader-Worker-Verifier architecture, available exclusively through Lark (Feishu) integration.What Mavis Does Well
- Leader-Worker-Verifier pattern — clear separation of planning, execution, and verification roles
- IM-native experience — multi-agent collaboration directly within Lark chat
- “Instant reply” with background execution — acknowledges the user immediately while tasks run asynchronously
- Context isolation between workers — each worker operates independently
Where Myrm Goes Further
Migration from Mavis to Myrm
Result: 7 upgrades, 1 equivalent, 0 downgrades. Users gain open-source data sovereignty, model freedom, and multi-platform access with zero capability loss.
vs Claude Code — Fork Subagent & Prompt Cache
Claude Code uses a “Fork Subagent” design with byte-level prompt prefix alignment to reuse KV Cache, reducing sub-agent costs by up to 90%.Where Myrm Goes Further
Migration from Claude Code to Myrm
Result: 6 upgrades, 0 equivalent, 0 downgrades. Myrm delivers superior cache economics with 3 unique capabilities (Break Detection, Anti-Thrashing, Resume-Aware) that Claude Code completely lacks.
SubagentExecutor Reliability (Jul 2026)
Beyond cache economics, Myrm’s dedicated SubagentExecutor engine handles what users feel when delegating work:
Regression coverage: 10,838 orchestration full-chain tests passed with 0 failures (Jul 11 2026): delegate core (97) + tournament & verification (5) + spawn tools & security registry (21) + sub_agents full suite incl. manager/executor/event/checkpoint/orphan recovery (603) + parallel & resume compact (5) + Kanban engine full suite (429) + streaming & stream_recovery (637) + security full suite incl. context_budget/isolation/permissions (1,761) + LLM toolkit incl. credential_pool/routing/fallback (1,344) + context_budget dedicated (60) + credential_pool dedicated (36) + error_classifier dedicated (175) + Item #13~15 verification (2,431): sub_agents suite (603) + security suite (1,761) + multiplexer & thread_sharing (27) + frontend Vitest SubagentDashboard+SubagentStore (17) + EventForwarder (12) + TeammateMailbox (11) + Item #17 verification (105): consensus engine (64) + council orchestration (25) + council integration (16) + consensus types (8) + Item #18~19 verification (978): memory scope binding (9) + ACP full suite (386) + A2A resolver+SSRF (44) + MCP toolkit full suite (539) + Item #20+TokenBudget verification (847): budget_guard (59) + manager_core (186) + context_budget (71) + executor+delegation (165) + config+heartbeat (70) + stream+lifecycle (121) + token_control+memory+loop_guard (175) + Item #32+#36 verification (230): dynamic_workflow engine (64) + DW e2e (9) + pipeline+templates (47) + orchestrator+swarm (99) + swarm_fission+marketplace (11).
Core Preset Tool Availability — SCIP Phase 0+1 (Jul 2026)
Fixes the silent dead delegate class of bugs where a browser/analysis subagent preset referenced stale tool names,filter_tools() produced zero tools, and the parent agent believed delegation succeeded.
Honest competitor note: OpenClaw and Hermes already support browser delegation paths; Myrm adds preset SSOT fail-fast, cache-safe peripheral skill binding, and L3 parent toolkit session merge so production delegates do not silently no-op or fail after enabling Browser.
SCIP bundle: 65 passed, 1 skipped (harness preset/guardrail tests + server binding/spawn smoke + frontend hint vitest, Jul 2026) · API Live E2E
example_domain_seen ~6.5s (mimo-v2.5-pro, Jul 8 2026).
Claude Code 2.1.154~2.1.157 Harness Upgrade
With Opus 4.8, Claude Code introduced/effort (6-level reasoning budget), Dynamic Workflows, lean system prompt, and enhanced --resume. Here’s how Myrm compares:
Result: 10 upgrades, 0 equivalent, 0 downgrades.
Dynamic Workflows — Real-World Pain Points
Based on user feedback from Claude Code’s Dynamic Workflows feature (2026-05 data):
Key architectural difference: Claude Code DW generates JavaScript orchestration scripts at runtime with keyword triggers (non-deterministic, can fail or diverge). Myrm offers a dual-path strategy: declarative DAG plans (deterministic, verifiable before execution) for production workloads, plus an optional Dynamic Workflow mode (manual toggle, Python PTC sandbox, deterministic
workflow_id, SQLite event sourcing for idempotent sub-agent replay) for ad-hoc parallel scripting — all with a first-class Human-in-the-Loop (HITL) GUI.
Confidence-graded results: Myrm’s DW summarization automatically classifies each finding with a 4-tier confidence badge — ✅ Verified (backed by execution evidence), ⚠️ Unverified (LLM reasoning only), ❌ Refuted (contradicted by evidence), 💥 Failed (task errored). Users instantly distinguish reliable conclusions from LLM speculation. No competitor offers this capability.
vs Scrapling & BrowserUse — Fully Autonomous Hybrid Browser Engine
While typical AI agents use standard Playwright/Selenium wrappers (like BrowserUse) that crash constantly on dynamic pages, and traditional scraper frameworks (like Scrapling) require developers to manually write code to bypass anti-bot measures, Myrm Agent introduces a Fully Autonomous Hybrid Browser Engine.What Scrapling & BrowserUse Do Well
- BrowserUse: Wraps Playwright for LLMs to click elements, but relies on slow LLM reasoning every time a locator fails.
- Scrapling: Provides powerful stealth tools (
camoufox/curl_cffiHTTP fetching) but requires developers to manually wire them into a crawler script.
Where Myrm Goes Further
Migration Wins
Users migrating from raw Playwright automation or other Agent frameworks gain:- Zero-Crash UI Navigation: End the frustration of “Element not found” errors ruining a 20-minute agent task.
- Blazing Fast Data Retrieval: Don’t burn memory on full Chromium instances for simple text extraction; Myrm degrades to HTTP seamlessly.
- Production Grade Reliability: Auto-detects 10 CAPTCHA providers and solves them via CapSolver API with FallbackSolver chain (manual HITL takeover when needed). Passes Cloudflare invisibly.
Credential Vault — Passwords & 2FA Never Enter the LLM
When agents automate login flows, the default pattern in most frameworks is catastrophic: the model generates the password string and passes it throughtype or fill tool arguments. That value persists in chat history, MCP logs, and retry buffers.
Myrm’s Form Credential Vault separates knowing a credential from using it:
Where Myrm Goes Further
Honest Comparison with FSB
FSB pioneered the vault-boundary pattern for browser automation (label reference → extension decrypts → DOM fill). Myrm adopts the same security principle and extends it to desktop Computer Use, native TOTP, and unified product GUI — without requiring a separate Chrome extension. FSB still leads on payment-card-specific APIs (use_payment_method); Myrm covers password fields and 2FA today.
Migration Wins
- From Hermes / OpenClaw: Stop pasting passwords into prompts; configure once in Settings.
- From FSB: Same mental model (labels), plus desktop apps and TOTP in one workspace.
- For enterprise: Combine vault with 12-dimension permissions, structured audit trail (37 decision types + Prometheus metrics), and incognito sessions for sensitive runs.
MCP Security Gate — Know Risk Before You Enable
Third-party MCP servers are a growing attack surface: poisoned tool descriptions, sensitive path access, and runtime tool injection can compromise an agent mid-conversation. Myrm gates MCP before it reaches your workspace:
Honest scope: Full 102-hook static rule packs are roadmap — this gate covers the real user path (Settings → enable → chat). Regression: harness + server API tests (21+ cases).
Migration wins
- From Hermes: Stop discovering bad MCP only in logs — block or confirm in Settings first.
- From OpenClaw: Env filtering is not enough; get pre-enable scan + runtime disconnect.
MCP Protocol Architecture — Persistent, Event-Driven, Cache-Friendly
Myrm’s MCP implementation uses persistent warm connections with event-driven tool discovery — not the cold-start reconnect pattern used by competitors.
Why this matters for cost: Every time a competitor reconnects to an MCP server, the tool list changes token positions in the prompt, invalidating the LLM’s prompt cache. With 10+ MCP servers enabled, this can cost 2.00 extra per hour in wasted cache misses. Myrm’s frozen proxy approach keeps the tool section byte-stable across turns.
Migration wins
- From Hermes: No more “tool not found” errors after idle timeouts — persistent connections stay warm.
- From OpenClaw: Stop paying for prompt cache misses caused by tool list churn on every reconnect.
Shell Command Visual Approval — See Every Pipe Before You Allow
When an agent runscurl … | bash, a single monospace line hides which segment downloads code and which executes it.
Myrm’s Shell Command Display splits pipelines into spans with per-segment risk coloring — the same mental model OpenClaw users expect, enabled by default and wired into our 6-layer security stack.
Honest limits: Very long commands are truncated with a clear UI notice. Local use requires both frontend and backend running; if you open only the WebUI, you get explicit startup guidance—not opaque parse errors.
Migration wins
- From OpenClaw: Familiar segmented shell view — plus 12-dimension permissions, leak detection, and structured audit trail.
- From Hermes / Claude Code: Stop approving blind one-liners — see pipe segments and risk before one click.
vs MemPalace — AI Memory System (14.9K+ Stars)
MemPalace is a standalone AI memory system using architectural metaphors (Wing/Room/Closet/Drawer) with a “store everything verbatim” philosophy, achieving 96.6% R@5 on LongMemEval. It operates as an MCP tool that external AI assistants can call.What MemPalace Does Well
- Verbatim storage — stores raw conversations without lossy summarization
- Architectural organization — Wing/Room/Closet/Drawer hierarchy gives AI a “navigation map”
- 4-layer memory stack — L0 identity (~50 tokens) through L3 deep search, keeping wake-up cost under 900 tokens
- Multi-format ingestion — normalizes Claude, ChatGPT, Codex, Gemini, Slack exports into a unified format
- Local-first — runs entirely on your machine with ChromaDB, zero cloud API calls
Where Myrm Goes Further
Migration from MemPalace to Myrm
Result: 10 upgrades, 0 equivalent, 0 downgrades. MemPalace users gain a complete AI agent platform where memory is natively integrated — not an external add-on — with 12 capabilities MemPalace doesn’t offer (intelligent forgetting, GUI management, safety scanning, preference tracking, health diagnostics, and more).
Dual-Path Data Visualization — No-Code Charts & Interactive Tables
Competitors like Codex and Claude Code output plain Markdown tables when working with data. Myrm provides a dual-path architecture that covers everything from simple data display to complex interactive dashboards:
Path 1 — Generative UI (zero-code): The agent renders 24 built-in component types (UIChart, UITable, UIProgress, UIBadge, UICard, UIGrid, UITabs, and more) via SSE events — secured by front-end + back-end dual whitelist, with 8 validation rules, i18n support, and conditional rendering. Users see rich interactive data visualizations without writing a single line of code.
Path 2 — React Artifact (full-feature): For complex visualizations, the agent generates custom React components using Recharts, Echarts, D3, or Chart.js. These render in a Sandpack-powered live preview with code editing, console output, and Tailwind CSS support.
Verified by 697 tests (2026-07-03) → render_ui full pipeline 87 tests (2026-07-04): Harness A2UI spec 24 ✅ | Interactive UI 59 ✅ | Artifact renderers 47 ✅ | ArtifactPortalStore + SSE 15 ✅ | Chat export + schema 39 ✅ | Backend Artifact system 22 ✅ | Backend Share API (CSP + HMAC + traversal) 41 ✅ | Backend Deploy API 15 ✅ | Backend Artifact file_id chain 3 ✅ | Deploy integration 4 ✅ | Harness Core Artifacts 27 ✅ | Harness Agent Artifacts 128 ✅ | Inline Artifact Events 9 ✅ | Artifact Judge 9 ✅ | Inline Artifact Push 7 ✅ | Mermaid Theme 18 ✅ | Render UI Tool 9 ✅ | Artifact Services (share_token + share_bundle) 13 ✅ | UI Artifact Stream 7 ✅ | DocumentSelectionToolbar + useSelectionAction 12 ✅ | VersionHistory + VersionHistoryBanner 20 ✅ | reactCodeProcessor 35 ✅ | reactPreviewConstants 21 ✅ | Artifact navigation full-stack (PortalTabs + ArtifactsCenter + SpreadsheetPreview + DataGrid) 105 ✅ | Image annotation editor 54 ✅ | render_ui SSE wiring 14 ✅ | run_bind + fail-closed integration 13 ✅ | real LLM E2E + live Chrome DOM 1 ✅
vs Doubao Pro — “AI→Office Delivery” Full-Cycle Closed Loop
Doubao Pro (ByteDance) bundles Feishu’s Office suite (Docs/Sheets/Slides) for a full “AI generate → format → output” workflow. Myrm achieves the same — and better — with a lighter Artifact architecture:
Verified by 52 tests (2026.6.30): DataGrid 11/11 ✅ + CsvParser 19/19 ✅ + SelectionToolbar 17/17 ✅ + LocaleKeys 5/5 ✅
Verdict: Doubao’s advantage is Feishu Office ecosystem moat (thousands of engineer-years) — not replicable via technology. Myrm’s Artifact system (Monaco + DataGrid + HTML Preview + SelectionToolbar + one-click deploy) delivers a complete “AI→delivery” closed loop that’s lighter, faster, and easier to maintain for AI Agent scenarios.
vs Trae/WorkBuddy/deer-flow — Report Formatting & PDF Export
Users judge AI office output by one criterion: “Can this go straight to the boss?” Trae is praised for clean formatting, WorkBuddy’s stock-advisor generates magazine-style PDF reports. Myrm leads across the entire report generation pipeline:
Myrm exclusive advantages: Sandbox pre-installs PDF toolchain (pdfkit + reportlab + pdfplumber + pdf2image + img2pdf + Pillow) — Agent generates professional PDFs with zero setup; 8 report Skills cover all scenarios from daily briefings to deep research to competitive analysis to data visualization; daily-briefing + cron delivers automated briefings to 26+ IM channels.
4,612 report+artifact full-stack tests passing (2026-07-14): Artifact system 159 + Skill system 398 + File/PDF/Briefing/Cron 492 + Code execution 1,272 + Frontend 2,291.
vs Doubao Pro — Anti-Freeze & Granular Progress
Doubao Pro’s real user pain points: (1) PDF processing freezes at 87% on a 50-page file; (2) page unresponsive for ~30 seconds during long document generation. Myrm solves both at the architecture level:
Verified by 169 tests (2026.7.11): ProgressSteps 25/25 ✅ (incl. semantic tool labels + workflowStage) + LiveTerminal 27/27 ✅ + PetDispatch 8/8 ✅ + ArchiveRestore 2/2 ✅ + MessageStream handlers 29/29 ✅ (toolLifecycle/gapEvents/statusStream/tasksSteps/petDispatch/clarification/completion) + structured_clarify E2E (SSE unwrap + option.id + live API interrupt/resume 1/1 ✅) + Harness step_builder 38/38 ✅ + scrubbing 5/5 ✅ + event_handlers 47/47 ✅ (reason extraction + sensitive info scrub)
Verdict: Doubao’s freezing issues stem from “poorly implemented” (main-thread blocking + no granular feedback) rather than “not implemented”. Myrm’s architecture (SSE + WebWorker + PTC notify + tree-shaped progress + Live Terminal) builds a complete anti-freeze & progress feedback system that far exceeds a simple progress bar.
vs Marvis (Tencent) — Split-view Live Streaming Workstation
Marvis (Tencent) offers a split-view interface: “left side shows your Mac mini’s real-time desktop, right side is the chat — feels like sitting in front of the computer.” Myrm already has this — and more:
Verified (2026.7.14): 454 VNC+ArtifactPortal full-stack tests passing — Frontend: ArtifactPortal Store 4 + Portal interaction 49 + VisualApproval 4 + Entitlements 3 + ChatWindow 5; Backend: VNC Routes 7 + Desktop Snapshot 5 + Takeover Integration 7 + Permissions 5 + Deploy 11 + Artifact System 22 + Stream+DeepLinks 29 + Share 54; Harness: VNC 63 + Artifacts 168; Control Plane: VNC Proxy 18.
Verdict: Myrm comprehensively surpasses Marvis. Real-time video streaming is on par (noVNC vs WebRTC), but Myrm additionally provides DOM element-level inspectors (BrowserLiveView + DesktopLiveView), Browser Takeover human-agent collaboration (Agent pauses for user control, auto-resumes with learning feedback), cloud sandbox + local dual-mode, and macOS permission auto-guidance — all unique capabilities Marvis lacks.
vs Claude Office Visualizer — CLI Status Dashboard
Claude Office Visualizer renders Claude Code CLI status as a pixel-art office animation, showing agent state, context usage, and task progress through a standalone Next.js + PixiJS application. It addresses a real pain point: CLI users can’t see what the agent is doing without staring at terminal output.What Claude Office Does Well
- Visual agent metaphor — pixel characters that walk, think, and interact based on real CLI state
- 12 whiteboard modes — multiple visualization layouts for different data views
- Tour overlay — 7-step interactive guide for new users
- Attention system — screen flash notifications when the agent needs user input
- Docker deployment — easy self-hosted setup
Why Myrm Doesn’t Need This
Myrm is a GUI-first application. The problems Claude Office Visualizer solves — “I can’t see agent state” and “CLI output is boring” — don’t exist in Myrm’s architecture.Migration from Claude Office Visualizer
Result: 5 upgrades, 1 equivalent, 0 downgrades. Users move from a separate observation window to a fully integrated GUI with precise data, interactive controls, and a complete agent platform.
vs Coze / workflow platforms — batch LLM fan-out vs SubAgent batch (Jul 2026)
Workflow tools (including Coze) often ship a batch LLM or fan-out node: same prompt template × N inputs, cheap Lite calls, workflow-level billing. Myrm intentionally does not expose an in-agentllm_map primitive. Homogeneous or heterogeneous bulk work goes through delegate_task_tool (mode=batch|parallel) — each item becomes a full SubAgent with its own sandbox, tool policy, budget guard, GUI cost approval (≥$0.50 batches), subagent_control_tool (cancel/steer/list), and audit trail.
Migration win: Teams leaving Coze batch nodes for real operational batch work get traceable, approvable, recoverable jobs instead of a single fan-out bill with no row-level forensics.
vs Coze 3.0 / Lobster / Vercel v0 — Artifact Deploy & Read-Only Link (Sites 2.0)
Coze keeps projects inside its ecosystem; Lobster excels at public static links but not multi-platform GUI publish from an agent workspace; v0 targets developers building with AI. Myrm targets GUI users who want a shareable result from chat artifacts, not a separate site builder.
What you get: Landing page, report, or mini-app from chat → ~30s Deploy URL or instant read-only Link for reviewers.
Dual-path publish — GUI zero-token + Agent conversational: Configure Vercel, Cloudflare Pages, Netlify, or a custom webhook in Settings → Hosting Targets, then either click the Globe on any HTML artifact card → Publish (zero LLM tokens, server-side only) or let the Agent publish mid-conversation via the
artifact_publish tool (say “give me a live link” and the Agent does it). The Agent tool is conditionally loaded — zero prompt-token overhead when hosting is unconfigured. Preflight blocks bad artifacts before egress; each target keeps its own publication history and stale-version redeploy banner. Clawith still routes deploy through 5 Agent tools (permanent token cost); Myrm loads the tool only when needed (104+5 hosting pytest cases, 3 conditional-mount integration scenarios).
Test coverage (2026-07-29): 109 hosting pytest cases (104 module tests + 5 new artifact_publish tests including 3 conditional-mount integration scenarios), covering multi-target CRUD, webhook full chain, SSRF, redeploy upsert, WebSocket status, credential edge cases, and Agent tool factory — real DB/vault/preflight/orchestrator; external provider HTTP mocked only at egress boundary.
Beyond deploy — the full artifact editing lifecycle: Myrm’s Artifact Portal isn’t just a viewer. Code artifacts open in a Monaco Editor (VS Code core) where you can edit directly. The SelectionToolbar lets you highlight code and invoke Modify / Explain / Optimize / Comment actions — surgical AI edits without leaving the preview. Your edits are silently tracked and auto-injected into your next message so the Agent continues from your version. React artifacts render live via Sandpack with hot-reload. Version Diff View (exclusive) — click the Diff button to see exactly what AI changed via VS Code-grade Monaco DiffEditor, with Inline/Side-by-side toggle, mobile-responsive auto-inline, version labels (V2→V3), and real-time diffing during generation. SHA-256 integrity checks and snapshot rollback round out the lifecycle. Competing tools like WorkBuddy rely on external document integrations (e.g. Tencent Docs) rather than a self-built editing system.
Honest limits: Long-term public sites should use Deploy + your own domain; pure code artifacts must be exported as html first (preflight + UI gates align). Share revocation is fully implemented (one-click revoke → instant 404 Link Revoked, revoke-proof token fingerprints, revoked password links never show the password gate, revocation audit log) — chat and artifact shares both covered.
vs Codex — MCP Ecosystem, Agent Templates & Multi-Channel Automation
Codex recently showcased Linear MCP integration and automated PM workflows. Myrm’s architecture already provides vastly superior capabilities in every dimension — verified by 543 tests.
Key takeaway: Codex’s “Linear MCP” is simply connecting an external MCP server — something Myrm already supports with a full GUI, security scanning, and 35 prebuilt service integrations. The “PM Agent” workflow requires no new platform features; Myrm’s existing template system + channel monitoring + Kanban pipeline already enables this with one-click setup.
Microsoft To Do connector — zero-config where competitors require app registration
The built-in Integration Catalog ships a pre-configured Microsoft To Do entry (npx -y microsoft-todo-mcp), so users connect through a device-code sign-in that reuses the public Microsoft client ID — no Azure App registration, no Client ID/Secret, no API key. The connect dialog inlines the guidance plus a “Learn more” link. Chinese users find the entry by its official Chinese name 微软待办 or the tags 待办/微软 — a discoverability layer no competitor (OpenClaw / CoPaw / LobsterAI / jiuwenclaw) implements. Verified end-to-end: backend catalog + API tests, frontend Vitest, and live Chrome E2E.
vs Codex (OpenAI) — Appshots & /goal GA
Codex recently shipped two flagship features: Appshots (window capture + text extraction via ⌘⌘) and /goal mode GA (long-running autonomous task execution). Here’s how Myrm compares:Appshots (Window Capture + Text Extraction)
What you get: When the agent wants to click your desktop, you see where — not just “Approve?” in text. On Tauri, a system-level red box stays visible even if the chat window is covered. Both macOS and Windows users get full Appshot capabilities — no competitor offers Windows window capture + text extraction.
Honest limit: OS overlay smoke requires the Tauri desktop app; browser-only WebUI gets inline + AttentionBar but not the OS frame.
/goal Mode
Lock Screen & Remote
User Pain Points (from comments) Myrm Already Solves
Result: Codex’s Appshots and /goal are simplified subsets of Myrm’s existing capabilities. Myrm covers 3 platforms (vs Mac-only), offers DAG execution (vs linear), and provides enterprise-grade budget control and completion verification that Codex lacks entirely.
Vibe Canvas — “Point-and-Fix” Artifact Interaction
What you get: Click any element in an Artifact preview, Myrm captures the exact HTML structure and sends it to the Agent with your instruction. No screenshots, no coordinate guessing — precise DOM-level “point and fix.”
Browser Extension Bridge — Authenticated Session Takeover
What you get: Let your Agent use your real Chrome session (Gmail, SaaS apps, internal tools) with fine-grained domain control. Every action requires your approval. Real-time status shows exactly what the extension is doing. Myrm does not bulk-read your OS Chrome/Edge
History.db or bookmark database — logged-in access is Extension Bridge + encrypted Session Vault only.
Why Myrm doesn’t need Chrome Cookie Import: OpenClaw imports Chrome system profile cookies because it lacks persistent session storage — every browser launch starts fresh. Myrm’s SessionVault (AES-256-GCM) + auto_restore_domains solves this at the architecture level: authenticate once via takeover → session encrypted and saved → automatically restored on every future visit. This covers all auth methods (password, 2FA, QR, OAuth), saves cookies + localStorage (not just cookies), works across all platforms and deployment modes (including cloud-hosted), and requires zero maintenance when Chrome updates its cookie format. 196 related tests passed.
vs 360 LobsterAI — Consumer Agent Platform
360 LobsterAI is a consumer-oriented agent product (backed by Zhou Hongyi) featuring manual model-tier selection, 100+ preset “expert lobsters,” and a “coaching” onboarding flow.Cost Intelligence
Agent Templates
Multi-Channel Access
Migration from 360 LobsterAI to Myrm
Users migrating from 360 LobsterAI gain:- Automatic cost optimization — no need to manually select “lite” vs “full” mode
- Smarter routing — system learns preferences and improves over time
- 35+ channels vs 3-4, with per-channel agent binding
- Enterprise security — 6-layer defense vs consumer-grade
- Unlimited extensibility — Skill Marketplace + custom agents vs fixed presets
- Full data ownership — self-hosted option, no vendor lock-in
Smart desktop download & release (vs Hermes / OpenClaw / Cursor)
When shipping desktop agents, competitors often send users to GitHub Releases or force manual CPU architecture picks; CI races can corrupt update manifests. Myrm (verified v0.1.39) offers consumer-grade download + production OTA + an honest release contract:
Result: Day-one install goes from “understand GitHub” to “click Download” — the gate for mainstream users.
Local Migration Wizard (v1.4)
Myrm provides a GUI-first competitor migration wizard on Local and Tauri deployments. It supports 14 competitor adapters (12 ready) with auto-discovery, MCP config auto-conversion, model auto-mapping, and full dry-run/confirm/rollback lifecycle.Five discovery sources (Local / Tauri)
Four preview lanes
- Persona → Agent — SOUL-style instructions attach to a target agent profile.
- Facts → Global memory — structured memory rows with batch rollback.
- Skills → Review queue — imported skills require approval before activation; optional Agent binding.
- API keys → Opt-in — never silently copied; user confirms each secret.
Workflow
Scan → Preview (dry-run) → Confirm → Result in Settings → Migration. Server-side resolve_competitor_import_source() forces the correct adapter (fixes OpenClaw source=auto mis-routing to Hermes).
GUI-First Seamless Skill Migration & Atomic Persistence
When users import third-party ecosystem skill bundles (like Hermes ZIP exports), competitor architectures often rely on basic file overrides (os.replace) and sequential database writes. Under SaaS high-concurrency conditions, this leads to catastrophic “split-brain” scenarios (database records inserted but files missing) or API freezes.
Myrm introduces a Zero-Cost Hot Migration & Atomic Persistence Engine:
Result: Competitor servers will crash or leak threads when 50 users upload skill ZIPs simultaneously. Myrm processes massive concurrent imports with 0ms thread-blocking and zero dirty-write risks.
vs Hermes claw migrate
Hermes cron import (Aug 2026):
test_hermes_cron_converter.py 7/7 passed (~5s) — schedule mapping, omit-model, skip no_agent/script, coverage row, plan build. Not verified: Chrome MCP E2E for cron import UI.
Honest score: 9.2/10 for local eight-source GUI migration — not “perfect for every competitor artifact.” Not available on SaaS cloud sandboxes (no access to the user’s host filesystem); use Local WebUI or Tauri desktop only.
Verified (2026-06-08)
Automated regression on a developer machine (minimal batch, ~7s total):
Total migration-specific tests: 276 all passed (Jul 25, 2026 — covering 14 competitor adapters / dry-run sessions / import adapters / MCP config converter / model migration / architecture guards / multi-competitor discovery / E2E / API / skill batch import). Competitors checked (OpenClaw, Hermes, DeerFlow, CodePilot, LobsterAI, CoPaw, jiuwenclaw, PilotDeck, Grok Build): no GUI 14-source migration wizard with MCP auto-conversion; Hermes offers CLI
claw migrate only; Grok Build supports 2 sources (Claude + Cursor) via TUI modal.
What you get after migrating (user-facing)
- Keep your persona and habits — SOUL-style instructions land on an Agent profile you pick, not a one-size-fits-all default.
- Keep structured memory — facts and episodic sessions import with batch rollback if you change your mind.
- Keep skills under control — imported skills sit in a review queue; you approve and bind them to the right Agent.
- Keep skill usage history — call counts, last-used dates, pinned status, and lifecycle state are automatically preserved from Hermes
.usage.json. Your most-used skills stay prioritized; pinned skills remain pinned; no cold-start re-learning period. - See your savings before you commit — the dry-run preview shows a Token Economics comparison card: your current skill-loading cost (competitor full injection) vs Myrm’s on-demand loading. Typical savings: 94% (e.g., 15 skills: 7,500 → 450 tokens per turn).
- No silent key theft — API keys import only when you opt in lane-by-lane.
- Honest follow-up — MCP servers and messaging channels are guided in Settings (we do not claim one-click channel parity with Hermes CLI hints).
- Hermes cron under control — scheduled jobs import paused; preview shows skips; Result links to Settings → Cron; batch rollback available. Kanban is not auto-migrated.
What You Gain by Choosing Myrm
Never Lose Context
42+ module memory system — the most complete AMO implementation in the industry. 8 memory types, 7-signal retrieval fusion, 8-layer staleness defense (including LLM-driven per-fact expiry review), and cross-session consolidation with one-click rollback. What Anthropic Dreaming and Mem0 partially cover, Myrm delivers end-to-end.
Work From Anywhere
Approve tasks, steer agents, and monitor progress from any device via 35+ messaging channels.
Smart Tool Usage
3-tier tool layering + on-demand loading + 3D health monitoring + 14-type error diagnostics with expert fix suggestions. Tools are selected precisely, used effectively, and self-heal on failure.
Enterprise Security
6-layer defense-in-depth with budget control, PII protection, taint tracking, and audit trails.
Zero Vendor Lock-in
100+ models, self-hosted or cloud. Your data is always yours.
vs OpenClaw 2026.6.1 — 跨越”能用”到”好用”的终极形态
OpenClaw 2026.6.1 是其重磅更新,主打 Windows 原生节点、技能工坊、工作看板以及稳定性修复。然而,从其官方发布和大量用户社区反馈来看,其依然面临着”更新即崩”(如飞书/QQ断连)、任务缺乏直观可视化干预、以及”缝合感”较强的痛点。 Myrm 在架构和体验设计上实现了降维打击:Where Myrm Goes Further
一句话总结:从 OpenClaw 迁移到 Myrm,你将告别”修一天跑一小时”的极客折腾期,获得一个拥有全天候健壮队列、极致 UI/UX 细节以及跨端原生体验的企业级 AI 工作站。
实时语音对讲(Realtime Voice & Multimodal Talk)
OpenClaw 拥有成熟的语音架构(多传输层、插件化提供商、电话集成),但 Myrm 在用户实际体验和成本维度实现了超越:
关键数据:1,176 项语音相关测试全部通过(语音核心 632 + Discord Voice 全栈 485 + Channel Security 59),覆盖 Realtime API / TTS 5 提供商 / STT 5 提供商 / WebSocket 全双工 / Barge-in / Agent Bridge / Vision 融合 / Discord Wake Word / VoiceReceiver DAVE 加密 / VoiceFollowManager 自动跟随。
vs DeerFlow — 彻底告别大模型崩溃死循环
DeerFlow 作为一个优秀的开源智能体框架,在健壮性处理上遇到了典型的业界难题:大模型极易因为长日志和异常输出引发死循环崩溃。Myrm 将 DeerFlow 暴露出的核心痛点转化为原生中间件级别的降维打击解决方案。工具调用安全网(Tool Calling Safety Net)
无论是使用 LangChain 还是原生 API,执行长线任务时极易遭遇致命的崩溃死循环。Myrm 实现了免人工干预的 100% 自动容错:
用户感知收益:当别的 Agent 跑到一半因为“日志太长”或“JSON 残缺”而崩溃挂起时,Myrm 会自动截取超长日志存入本地文件,并用一段礼貌的提示语告诉大模型:“输出过长已落盘,请使用读文件工具查看某某路径”。它永远能优雅地处理各种边界崩溃。
智能记忆检索(Context-Aware TF-IDF Memory Retrieval)
在注入长期记忆时,普通框架会盲目按时间或事实置信度排序,导致上下文被大量“高置信度但与当前任务完全无关”的噪音记忆占满,反而忘记了核心指令。 Myrm 引入了 TF-IDF 与置信度双重加权检索:在注水前,先利用tiktoken 提取当前对话最近几轮作为 current_context,计算所有记忆与该上下文的余弦相似度,并结合事实的置信度进行双端加权评分(similarity*0.6 + confidence*0.4)。
用户感知收益:如果你积累了数百条记忆,迁移到 Myrm 后绝不会体验到因为“记得太多而变笨”的痛点。我们在极低的 Token 预算内提供“最精准”的记忆召回。
智能防泄漏与零卡顿的标题生成 (O(1) Anti-Blocking Title Generation)
当用户粘贴超大日志或代码块时,后台生成标题的正则清洗往往会卡死整个服务器(Event Loop Blocking);此外,廉价模型生成的标题经常带有Title: 等废话前缀,且容易暴露 API Key 等敏感信息。
Myrm 在底层实现了 O(1) 早期截断与军工级脱敏管线:
用户感知收益:哪怕你粘贴了 1MB 的错误日志,Myrm 也能瞬间为你生成一个安全、精准、无废话前缀的标题,且整个过程对服务器性能零冲击。细节体验 100% 碾压竞品。
深度研究与长程任务挂机(Deep Research & Offline Guardian)
DeerFlow 1.x 作为专业 Deep Research 框架广受好评,但 2.0 已完全重写为通用 Agent Harness(README 明确说”shares no code with v1”),原深度研究能力仅在 1.x 分支维护。Myrm 提供了完整的深度研究引擎与军工级长任务挂机保障:
用户感知收益:你可以发起一个数小时的深度研究任务后安心关闭浏览器甚至合上笔记本——Myrm 会自动阻止系统休眠、将任务状态持久化到数据库。即使服务器意外重启,Myrm 也会自动恢复中断的任务并在完成后推送系统通知。研究开始前,你的 Wiki 知识库会被自动检索,已有知识不再重复搜索,直接节省 token 费用;旧报告会自动入库并在后续研究中复用和验证,确保知识不断迭代更新。研究成果通过三层机制确保不丢失:同会话内直接可见,跨会话通过 Wiki 自动归档并被后续研究召回,普通对话中 AI 也能通过记忆检索工具访问历史研究。对比类查询(如”React vs Vue”)自动按统一维度规划研究步骤并在报告中生成汇总对比表。研究过程中,每轮研究的实时成本和预算使用情况都在 UI 上清晰可见。这种级别的深度研究能力与长任务可靠性,在所有竞品中独一无二。(9,487 项相关测试全部通过)
Long Report Reading (Auto TOC)
When an Agent reply contains multiple headings, Myrm automatically builds a table of contents beside the message — click any chapter to jump, and the highlight follows as you scroll. ChatGPT, Hermes, and OpenClaw do not offer equivalent in-chat navigation for multi-section reports; users typically copy content into Notion or Word to get an outline.Multi-Agent Orchestration & Zero-Trust Verification (vs Hermes Agent / Claude Code)
In the realm of Multi-Agent architectures, most frameworks rely on blind trust and non-deterministic LLM improvisation. Myrm breaks the serial bottleneck and completely eliminates the pain points of “hallucinated success” and “budget burnout” found in competitors like Hermes and Claude Code.Where Myrm Goes Further
Migration Wins
If you migrate from Hermes or Claude Code’s Dynamic Workflows, you gain:- Absolute Cost Predictability: Never wake up to a massive API bill because an agent got stuck in a loop.
- Genuine Execution Evidence: When Myrm says a task is complete, it’s backed by OS-level execution logs, not an LLM’s imagination.
- Flawless UI Experience: Background agents remain strictly in the background, and human intervention is rich and corrective rather than binary.
Visual Approval, Artifact Lifecycle & Intelligent Template Reuse (vs Codex)
Codex offers basic screenshot approval and no artifact management. Myrm delivers a complete visual approval system across Web, Desktop, and Mobile with a full artifact lifecycle from generation to deployment — verified by 347 artifact delivery full-stack tests (harness registry 137 + server processing 105 + frontend Portal 105). Zero-click delivery pipeline: Agent generates file → ArtifactRegistry auto-registers with dedup → SSE event → Portal auto-opens with preview (competitors only call OS Open, no versioning, no publish, desktop-only).Where Myrm Goes Further
Migration Wins
If you switch from Codex, you gain:- Multi-platform visual approval: From basic screenshots to Web inline highlights, Tauri OS overlay, and mobile swipe gestures
- Artifacts become permanent assets: Complete version management with SHA256 tamper-proof verification and multi-target hosting publish from Settings
- Agent learns your style: No manual template saving needed — Agent automatically learns your preferred report formats and coding patterns
- Precision collaboration: Select any code in the Artifact Portal and trigger modify/explain/optimize/comment directly with the Agent
AI Companion & Desktop Pet Status Visualization (vs Codex)
While Codex offers a basic floating pet showing ~4 states, Myrm provides a full AI companion system with 65 sprite vitest cases (7 files, incl. R06 PSUA bridge/bubble/away-completion; run locally with bun to verify) plus 2 companion pytest passed (install/serve; Aug 2026 agent run):
Migration win: From a simple 4-state icon to a Hermes-aligned 7-state Petdex companion with correct waiting on approvals, Tauri pop-out desktop workflow, Agent identity sync, evolution, and cross-platform consistency.
Voice Input & Full-Duplex Voice Sessions (vs Codex / Claude / ChatGPT)
Myrm provides the most comprehensive voice integration among AI agents — 4,200+ lines of full-stack voice code across 3 session modes:
Migration win: From basic mic input to full-duplex voice conversations with 3 modes, adaptive barge-in, Agent Bridge for server-side execution, and zero-config browser fallback.
Workspace Organization & Project Management (vs Codex / Claude / ChatGPT)
Migration win: From single-dimension tree folders to 3D filtering (project + date + source), milestone roadmaps, shared memory, auto context injection, and multi-agent concurrency safety — all verified by 39 project-specific + 115 workspace tests.
Thinking Process Visualization (vs Codex / ChatGPT / Claude)
Migration win: From limited thinking display to full 6-level intensity control with per-model persistence, 6-tag parsing for all LLMs (think/thinking/thought/antthinking/reasoning/REASONING_SCRATCHPAD), 3-layer Thinking Signature management (proactive cleaning → error recovery → stream scrubbing), and 3-format export — verified by 152 automated tests.
Top Agent Architecture Teardown (“Agent Harness Parsing”)
In a recent industry deep dive, the “Agent Harness” (the complete software infrastructure wrapping the LLM) was confirmed as the decisive factor in Agent performance. Compared with theoretical ideals and competitor implementations, Myrm’s architecture demonstrates crushing advantages in these key areas:How Myrm Leads
Migration Wins
Moving away from competitors relying on simple ReAct loops and basic memory, you gain:- Never getting lost in long documents: Smart previews and on-demand loading ensure the model always focuses on high-signal Tokens.
- Stop paying for empty loops: If the AI gets stuck, it immediately suspends execution and prompts you for instructions/parameters, rather than spinning in a costly loop.
- Engineering-grade reliable delivery: Code that hasn’t passed linting cannot “claim to be done.” Say goodbye to hallucinatory success.
Automation & Hooks Configuration (vs Codex / Claude Code)
Migration win: From CLI-only hooks that users report as “not easy to use” to full GUI automation panels — security policies, risk rules, event triggers, and per-agent behavior all configurable through professional visual editors. 405 automated tests (140 Hook core + 121 Planner + 49 Middleware + 95 CompletionGuard) verify the entire lifecycle pipeline.
Execution Safety & Error Prevention (vs Codex)
Migration win: From “the model should check git status before editing” (prompt-level advice) to automatic 5-layer security evaluation + auto-snapshot on every file mutation + 7-algorithm loop detection — all transparent, zero user intervention. 203 automated tests verify the entire safety pipeline.
Error Prevention UX (vs Codex)
Migration win: Every “Error Prevention UX” enhancement proposed for Codex is already implemented in Myrm with deeper, architecture-level protection. Shell risk visualization, visual approval overlays, and approval card editing are exclusive to Myrm. 742 automated tests (163 frontend + 579 Harness) verify the entire error prevention pipeline.
Long-Task Automated Orchestration (vs Codex Manual Workflows)
When handling long-running, multi-file refactoring tasks spanning days, many developers rely on “manual constraint workflows” popularized by tools like Codex—manually splitting subtasks, creating.codex-work/tasks/*.md folders, and even writing handoff documents like handoff.md by hand.
While this semi-manual mode prevents LLM context pollution, it should never be a burden for the end user. Myrm upgrades this “geek’s workshop” into automated infrastructure:
Myrm’s Lead
Conclusion: Migrating to Myrm frees your hands completely. Say goodbye to the artificial workflows you imposed on yourself just to “keep the model from acting dumb.” Let the machine do the orchestration it was meant to do.
Session-Level Performance Observability (vs Hermes / DeerFlow / CoPaw / LobsterAI / OpenClaw)
Where Myrm Goes Further
When a long task feels slow, competitors give you nothing to act on — no breakdown of where the time went and no number for how long you actually waited for the first token. Myrm’s session analytics panel turns “the agent felt slow” into concrete, actionable data:- LLM duration breakdown by model: every session’s LLM calls are aggregated per model (
call_count+total_duration_ms), so you can see at a glance which model is your cost bottleneck and which one is your time bottleneck. - Response speed summary: time-to-first-token aggregated at the session level (average + P95 + sample count), so a single slow reply stands out against the session baseline instead of being lost in one big duration number.
- Zero configuration: the panel is fed entirely by existing event data — no new instrumentation, no extra LLM calls, no prompt-cache impact.
Competitor Pain Points
None of the five competitors aggregates LLM call durations by model or exposes a time-to-first-token summary in a user-facing panel.
Migration Wins
- Bottleneck identification in one glance — open any session’s analytics and immediately see which model consumed the most wall-clock time.
- Slow replies are no longer invisible — the P95 TTFT number and per-reply comparison tell you exactly which response was the slow one.
- No extra cost — the feature is derived from data Myrm already records, so it adds zero token spend and zero configuration.
Ready to Start?
Quick Start
Get up and running in under 2 minutes.
Local Deployment
Self-host Myrm on your own infrastructure.
记忆系统:告别臃肿与幻觉,真正“懂你”的 AI
许多竞品(如 Mem0、Hermes 或 Supermemory)在记忆系统上存在显著痛点:要么提取过多导致记忆库臃肿、检索准确率下降;要么只存摘要导致精确信息丢失;要么必须依赖云端服务。Myrm 的记忆系统在架构上实现了全面超越:- 严格精度门控 (No-Op Default):我们拒绝像竞品那样“宁滥勿缺”地提取日常闲聊。Myrm 默认静默,仅在检测到高杠杆业务价值(如代码规范、架构偏好)时才提取入库,彻底解决记忆臃肿与 AI 幻觉问题。
- 原文无损保留 (Verbatim Storage):采用双轨存储,不仅保留用于语义检索的摘要,更 100% 无损保留原始对话和代码片段。当你需要精确数字或代码时,Myrm 绝不含糊。
- 异步画像投影 (Cognitive Deriver):无需你刻意调教,Agent 会在后台静默分析你的沟通风格、决策逻辑,并实时投影到下一次对话中,实现“零延迟”的默契。
- 物理级无痕模式 (Incognito Mode):独家支持阅后即焚,底层记忆引擎彻底卸载,确保敏感项目数据绝对零泄漏。
- 全景可视化与零配置:提供豪华的记忆指挥中心 GUI,打破黑盒;且内置 SQLite+Qdrant,无需配置 Docker 或 API Key,断网 100% 可用。
- 记忆冲突自动裁决 (Conflict Arbitration):当 AI 学到的新知识与已有记忆矛盾时,自动检测并路由至用户裁决(保留旧值 / 采用新值 / 自由编辑合并 / 双弃),低风险冲突 24h 自动保留旧值(KEEP_OLD,json_set 合并保留字段);高风险(重要度 ≥0.9)强制人工裁决;重复冲突去重。竞品 Cognee/Mem0 静默覆盖导致用户无感知、Hermes 只能手动编辑文件——56 项验证测试确认
-
7 维热门记忆浮顶 (Hot Memory Floating):通过
frequency_factor对数衰减追踪每条记忆的访问频率,结合 semantic/recency/importance/preference/rating/confidence 共 7 维加权几何均值评分,常用的偏好和规则自动在检索中排名更高。零配置、零延迟。竞品 OpenSquilla 的 Memory Dream 仅是单一 24h 定时任务且无频率评分——6 项验证测试确认 - 6 层规则防护体系 (6-Layer Rule Protection):is_user_locked 防覆盖锁 + PendingRecord 强制审批 + ConfidenceApprovalFlow 多信号风控 + Evolution Lock 防进化 + ProfileSnapshot Agent级快照回滚 + Skill Growth Audit 审计,从根源防止规则损坏,比“版本历史+回滚”的事后补救更优雅。竞品 Hermes 无版本历史、无回滚、直接覆盖文件——103 项验证测试确认
-
多智能体记忆隔离 (Multi-Agent Memory Isolation):
MemoryScope提供 agent_id+channel_id+conversation_id+task_id 四维隔离,代码助手学到的编码规范不会污染写作助手的文风偏好。SQLite 4 张核心表均含 agent_id 列,向量层支持 namespace-aware 多层检索。ProceduralMemory 拥有 9 个结构化字段(trigger/action/trigger_keywords/tool_name/tool_rule_priority/source/language/is_user_locked/scope)。竞品 Hermes 为单智能体+Markdown 文件+4 个 frontmatter 字段——177 项验证测试确认 -
GEPA 人机协同审批闭环 (Human-in-the-loop Approval):Agent 自动归纳的规则必须经过
PendingRecord审批流(pending→approved/rejected),用户在 GUI 中一键批准或拒绝,支持审批时编辑内容和批量操作。用户修正过的规则自动上锁(is_user_locked),Agent 后台整理和维护永远跳过锁定规则。独创ConfidenceApprovalFlow多信号风控(置信度+diff比例+历史有效率),高质量规则静默通过、低质量降级人工审核。竞品 Hermes 的 Agent 直接写入规则库且自我评估“总觉得自己做得好”——261 项验证测试确认(harness 175 + 前端 84 + server 2) - Goal Learnings 避坑经验闭环 (Pitfall Extraction & Injection):Goal 任务完成时自动提取 3 类前瞻性经验(Patterns 有效模式 / Gotchas 坑点警告 / Context 项目事实),每条经验需通过 confidence≥0.7+importance≥0.6 双门槛质量过滤后才存储。新 Goal 启动时自动检索历史经验并注入 metadata,让后续同类任务直接受益。竞品 agentmemory 仅有单一 Lesson 类型+纯关键词 text.includes() 检索+无质量过滤——96 项验证测试确认
- 检索证据三层可视化 (Retrieval Evidence UI):每条 AI 回复显示”引用了 X 条记忆”药丸标签 + 记忆预算百分比;点击展开详情面板可查看每条引用的类型、相关度分数、内容摘要、命名空间和溯源聊天跳转;用户还可直接评分反馈,评分回流到 7 维检索系统。双通道后端架构(LLM 主动 citation + 系统级 tool lifecycle 追踪)确保引用数据完整。ChatGPT 仅显示一行”✨Using memory”无详情,Hermes/Mem0 完全没有 GUI——51 项验证测试确认
-
情感化 AI 伙伴系统 (Companion System):Web + Tauri 同构 GUI 宠物 — Hermes 有 Desktop Petdex + CLI,但 无 Web GUI 移除与
/pet0-token。Myrm:PetGallery 社区装宠、chip 切换、GUI 移除、manifest fail-open、DELETE↔config SSOT;R05 Hermes 7 态 + 审批/澄清 waiting 行(五源 GUI 阻塞);独有 RPG 进化/Observer — 202 vitest + 2 pytest(2026-07-31 实测)。详见COMPANION_PETDEX_ADVANTAGE.md
vs. DeerFlow 2.0: From Framework to Industrial Harness
DeerFlow 2.0 introduced a 14-layer middleware chain and Docker sandboxing, but its architecture is heavily coupled with cloud-native microservices (requiring a standalone LangGraph Server). Why migrate to MyrmAgent?- 17+ Layer Middleware Kernel: We provide a more granular, strictly ordered middleware chain (17+ layers) with static DI Graph validation, completely eliminating implicit dependency bugs.
- True Agent-in-Sandbox: Unlike DeerFlow which creates sandboxes from within the agent, our architecture relies on the external deployment layer (Server/Tauri) to provide polymorphic sandboxing (Docker for cloud, lightweight process isolation for desktop). This ensures zero overhead on local machines.
- Zero-Copy ArtifactVault: For concurrent sub-agents, we use a
vault://pointer protocol for large files, achieving zero-copy sharing and completely preventing LLM context explosion. - Security-in-Depth Skills: While both support Markdown-based skills, we add a 3-layer defense system (Trust Decay + Security Scanning + Content Escaping) to ensure third-party skills never compromise your host.
vs Hermes / DeerFlow / OpenClaw / Cursor — Single-Track Task Progress (2026-07-02)
Most agents ship a single todo/plan tool that is always bound to Turn 1 — inflating prompts and offering no guard when the model tries to finish early. Myrm converged progress to Planning (todo_write) + Kanban, all off by default (~zero Turn1 token tax):
User benefit: Turn on Planning only for 3+ step jobs — you get a live todo tree in chat, Goal long tasks auto-enable it, files and guard stay in sync across reconnects, and you are not paying token tax for tools you did not ask for.
Verified (2026-07-03): harness progress + todo conditional 38/38 + server planning integration 24/24 (1 skipped) + live LLM E2E 3/3 (including resume with existing todos) + frontend
tasksStepsMerge 5/5; fixed two production bugs (task_workspace_root passthrough + workspace_todos_exist false negatives). Product WebUI dropped @playwright/test headless CI — UI integration uses real Chrome MCP + API prepare scripts.
2026 Latest Core Advantages Summary (vs Hermes, OpenClaw, DeerFlow, etc.)
Among the numerous open-source AI Agent frameworks and products, Myrm Agent stands out with its robust underlying architecture, ultimate user experience, and enterprise-grade security isolation:1. Crash-Proof Underlying Self-Healing Mechanism
When using APIs compatible with the OpenAI format, network jitter or frequent user cancellations often interrupt tool calls, leaving invalidtool_call_ids in the history and causing strictly validating models to crash.
Our Advantage: Myrm Agent runs a transparent self-healing middleware stack before every LLM request. tool_history_hygiene deduplicates tool results and rewrites duplicate tool-call IDs; dangling_tool_call repair fixes orphaned calls after cancellations. On provider errors, stream recovery retries once and shows progress in the Web UI. Chat history stays valid under network jitter—no crashes from malformed tool history.
2. Exclusive 4-Level Progressive Skill Loading
Equipping an Agent with a massive number of skills usually causes the context window to explode, increasing token consumption and degrading instruction-following capability. Our Advantage: Our exclusive 4-level progressive loading strategy loads the skill body into the context only when truly needed, achieving ultra-low token consumption. Supports/slash commands for precise activation, perfectly combining the Agent’s autonomy with your precise control.
3. Serverless-Grade Sandbox Hibernation & Wake-on-Demand
Traditional cloud deployments incur exorbitant costs if a resident sandbox instance is maintained for every user. Our Advantage: The control plane natively supports idle hibernation and wake-on-demand for sandbox environments. When a user is offline, the state of their 100% physically isolated exclusive sandbox is persisted and put to sleep. Upon receiving a new message or a cron trigger, it wakes up instantly, drastically reducing cloud idle costs.4. Diagrams Without Product Bloat
Most competitors ship heavy whiteboards or force manual layout. Myrm keeps diagrams in chat and optional MCP extensions. Our Advantage: Mermaid artifacts and render_ui (A2UI v3.1) cover most cases in-conversation. Editable whiteboards use Excalidraw/tldraw MCP — same industry pattern as Goose/Codex plugins, zero core maintenance.5. Ultimate Omnichannel IM Coverage
Many competitors only support a few mainstream channels, with fragmented message formats and media processing capabilities. Our Advantage: Natively supports 15+ channels (Slack, Discord, Feishu, iMessage, WeChat, Telegram, etc.), truly achieving “ubiquity.” A unified underlying media processing pipeline ensures consistent Markdown rendering and interactive experiences across any platform.6. Natural Language Driven Unattended Scheduling
The vast majority of Agents can only act as “passive response tools,” requiring active user triggers to work. Our Advantage: Built-in powerful unattended scheduling engine. With just a natural language sentence (e.g., “Send a briefing every morning at 8”), the Agent transforms into a proactive digital employee. Supports Webhook push and omnichannel reach.7. Memory Deduplication for a Forever-Lean Brain
Over long conversations, Agents easily accumulate duplicate facts, causing the memory bank to bloat and retrieval accuracy to plummet. Our Advantage: The underlyingdedup_semantics mechanism strictly filters duplicate facts during writing (auto-deduplication for similarity ≥0.95). Keeps the memory bank forever lean and efficient, ensuring the Agent gets smarter without bloating.
8. Full-Chain Feedback Rejecting Silent Failures
Bulk operations without progress bars, network errors failing silently, and files deleted instantly upon “zero references” cause extreme user anxiety. Our Advantage: Full-chain UI progress feedback, explicit exception Toast notifications, and retry mechanisms. Any action that jumps out of the current context has strong visual feedback. A built-in file system “Trash” soft-delete mechanism completely eliminates data loss anxiety.9. Line-Level Parallel File Conflict Protection
When multiple sub-agents work on the same codebase simultaneously, file conflicts can silently corrupt code. AWS uses coarse file-path mutex waves; OpenAI Codex has no application-level protection at all. Our Advantage: Three-layer defense — L1 line-level conflict detection (check_conflict_pre_write) blocks overlapping edits with precise line range info; L2 FileActivityTracker records every agent’s write ranges; L3 apply_parallel_write_isolation auto-switches to ISOLATED_COPY workspace + deferred merge when multiple writers detected. Non-overlapping regions of the same file can be written concurrently (2x throughput vs AWS blocking). Automatic lifecycle cleanup on task completion — zero memory leaks. 1,199 tests verified, 0 failures.
10. Industry’s Most Complete HITL Security Approval System
SDK frameworks like CopilotKit only provide a simple Promise hook for approvals (respond to continue), with no risk assessment, no rollback, no batch approval. Our Advantage: Myrm features a fully engineered HITL approval system — per-segment command risk annotation (safe/unknown/dangerous) + natural language impact explanation + six-locale FE Humanize SSOT (progress + approval same dialect; local/external scope hints; save-skill preview — 49 vitest Aug 2026) + screenshot BBox visual context + file-level snapshot rollback (pre_rollback trigger) + external side-effect warnings + batch approval + command edit with secondary verification + 4-level Allow-always granularity (permission/tool/exact/pattern) + timeout policies + cross-platform notifications + disconnection recovery + AI reviewer smart denial. Covers Desktop, Mobile, and Web. 2,036 security tests passed (80 approval logic + 42 command parsing + 16 component-level + 33 framework approval + 1,865 security engine).11. Complete Developer Diagnostics Toolchain (Not a Simple Debug Panel)
CopilotKit provides a simple floating debug panel + console.log; Myrm provides a complete product-grade diagnostic ecosystem. Our Advantage: 66+ SSE events all rendered as visual React UI + Browser/Desktop live view with element inspection + Eval Lab (8 assertion types + Recharts history) + distributed tracing (W3C traceparent) + session analytics + memory health dashboard + audit logs + session recording/replay + rate limit monitor + context health panels. 84 diagnostic tests verified.12. Four-Mechanism Extensibility Without a Plugin SDK
Hermes v0.19.1 introduced a Desktop Plugin SDK (@hermes/plugin-sdk) with ESM hot-reload and 10 contrib areas (panes, statusbar, titlebar, routes, etc.). However, it only works on Electron Desktop, has zero third-party plugins, and runs without sandboxing.
Our Advantage: Instead of a monolithic Plugin SDK, Myrm provides four extension mechanisms that work across all three deployment modes (Local WebUI, Tauri Desktop, Cloud Control Plane):
Additionally, Myrm’s global
⌘K search dialog and Slash command palette provide the same quick-access UX as Hermes’ PALETTE_AREA, while NavBar’s four status components (Extension Bridge, Remote Gateway, Background Tasks Panel, Notification Bell) surpass Hermes’ single Gateway Pill. 142+ extension tests verified.