<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>llama-cpp on tomrochette.com</title>
    <link>https://tomrochette.com/tags/llama-cpp/</link>
    <description>Recent content in llama-cpp on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Sun, 11 Oct 2026 02:01:18 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/llama-cpp/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>Atomic Agent</title>
      <link>https://tomrochette.com/agents/harnesses/atomic-agent/</link>
      <pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/harnesses/atomic-agent/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>harnesses</category><category>local-first</category><category>llama-cpp</category><category>open-source</category>
      <description>&lt;p&gt;Atomic Agent is an MIT local-first agent (TUI, CLI, and a Tauri desktop shell) that runs open-weight models through its own TurboQuant llama.cpp fork and drives your browser, files, shell, git, and MCP tools with the control loop and all state on your machine.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;The bet is that small quantized models stay useful for long, tool-heavy work if the harness does the engineering: grammar-constrained (GBNF) tool calls, a byte-stable prompt prefix for KV-cache reuse, a bounded context tail, and externalized state, so a 9B model clears half of GAIA Level 1 through the loop rather than the model.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A developer-preview Node.js/TypeScript agent installed by curl script with self-update and a documented uninstall flow, released under MIT (v0.6.7 on 2026-10-07, plus a first desktop-v0.0.1 Tauri build the same day).&#xA;One inference produces a JSON array of tool calls (grammar-constrained on a local llama-server), independent reads run in parallel, risky actions ask first, and long jobs continue past 25-step checkpoints to a 1,000-step or 2-hour ceiling.&#xA;The maker maintains a TurboQuant llama.cpp fork (claimed up to roughly 6.4x KV-cache compression and 30-50% throughput gains from speculative decoding) and a managed mode that downloads and runs the backend for you.&#xA;Surfaces go beyond the terminal: an HTTP API whose &lt;code&gt;/v1/chat/completions&lt;/code&gt; maps one request to one full macro-turn, a Tauri sidecar for embedding, Telegram and Discord bots with approval buttons, and Fusion mode where one model plans and a pool of throwaway workers executes via &lt;code&gt;fusion.delegate&lt;/code&gt; (workers cannot delegate, reach you, schedule tasks, or write memory).&#xA;It imports skills, memory, sessions, and (opt-in) keys from Claude Code, Codex, Pi, Oh My Pi, Hermes, and OpenClaw.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Active and quick-growing: 3,230 stars and 265 forks as of 2026-10-10, created 2026-04-21, pushed 2026-10-09, with a steady weekly release train and a developer-preview warning that APIs and behavior are still moving.&#xA;The headline benchmark is self-run: on the public GAIA validation Level 1 split (53 tasks), Atomic Agent scored 69.8% (37/53) against Hermes at 58.5%, both driving the same local &lt;code&gt;qwen-3.6-35b-a3b&lt;/code&gt; on one M4 Max with the same step budget, with per-task matrices, NDJSON traces, and logs published on the release tag; a model-scaling table shows 52.8% at 9B and 45.3% at 12B on the same split.&#xA;No independent reproduction exists, and a Hacker News search finds no thread about the project as of 2026-10-10, so the adoption evidence is the star curve and the published artifacts, not independent technical discussion.&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=AtomicBot-ai/atomic-agent&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=AtomicBot-ai/atomic-agent&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=AtomicBot-ai/atomic-agent&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The most complete local-model engineering in this section: grammar-constrained tool calls, cache-friendly prompt stability, and a compression step are aimed squarely at the failure modes of small models on long tasks.&lt;/li&gt;&#xA;&lt;li&gt;Fusion is a pragmatic subagent story: a planner delegates wide reads and drafts to disposable workers and merges the results, with the orchestrator barred from mutating tools.&lt;/li&gt;&#xA;&lt;li&gt;The vendor-run benchmark publishes its artifacts (matrices, traces, logs on the release tag) and holds the model and hardware constant against a named rival, which is better evidence practice than most self-reports.&lt;/li&gt;&#xA;&lt;li&gt;Full desktop tool surface with approval gating on every dangerous action, plus verify checks that never report an unchecked file as passing.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Developer preview: the README says APIs, commands, config, and behavior are still moving, so any integration should pin a release.&lt;/li&gt;&#xA;&lt;li&gt;The benchmark is vendor-run on one model, one machine, and one 53-task split, and the loser (Hermes) is also the comparison the vendor chose; treat 69.8% as a reproducible claim, not an independent one.&lt;/li&gt;&#xA;&lt;li&gt;The ecosystem around the agent is broad consumer surface (Atomic Mail, Atomic Chat, Atomic Wallet, Sigma Browser, Atomic VPN on the vendor site), which is an odd neighborhood for a dev tool and worth factoring into trust.&lt;/li&gt;&#xA;&lt;li&gt;Category adjacency: Telegram and Discord channels plus ClawHub skill installs overlap the assistant-runtimes family, so teams should decide whether they want a harness that answers your phone.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free, MIT, with no paid tier recorded as of 2026-10-10: local models cost nothing, and cloud models bill your own provider keys, including Claude Code and OpenAI Codex subscriptions driven through their signed-in CLIs.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/harnesses/ante/&#34; &gt;Ante&lt;/a&gt;: the other embedded-llama.cpp bet, roughly 15MB with the engine inside the binary; choose Ante for footprint, Atomic Agent for the fuller tool surface and Fusion.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/assistant-runtimes/hermes/&#34; &gt;Hermes&lt;/a&gt;: the rival it benchmarks against, a much larger assistant runtime with a learning loop; choose Hermes for channels and self-improvement, Atomic Agent for a local coding-and-desktop agent under your keys.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/harnesses/nanocoder/&#34; &gt;Nanocoder&lt;/a&gt;: the community-built local-first harness; Nanocoder has the wider documented local-engine list, Atomic Agent ships the deeper local-inference engineering.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for running actual work on small local models, where the harness-side engineering (GBNF calls, cache-stable prompts, externalized memory) is the difference between a demo and a workday.&lt;/strong&gt;&#xA;Not for anyone needing a stable API contract today (developer preview), and not for teams who want independently reproduced benchmark numbers before adopting.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-10 - Created from the harnesses resolution pass (two of three category workers placed it here); the assistant-runtimes adjacency (chat channels, ClawHub skills) is recorded in the cautions.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/harnesses/ante/&#34; &gt;Ante&lt;/a&gt; - the minimal embedded-llama.cpp contrast&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/assistant-runtimes/hermes/&#34; &gt;Hermes&lt;/a&gt; - the benchmarked rival across the category line&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/harnesses/nanocoder/&#34; &gt;Nanocoder&lt;/a&gt; - the other community local-first harness&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/harnesses/harness-feature-matrix/&#34; &gt;Harness Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/evaluation-review/frontierharness-eval/&#34; &gt;FrontierHarness Eval&lt;/a&gt; - the independent harness benchmark to read before believing any vendor self-report&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/AtomicBot-ai/atomic-agent&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/AtomicBot-ai/atomic-agent&lt;/a&gt; - repository: README, architecture, tool surface, license&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/AtomicBot-ai/atomic-agent&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/AtomicBot-ai/atomic-agent&lt;/a&gt; - stars, forks, issues, creation and push dates as of 2026-10-10&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/AtomicBot-ai/atomic-agent/releases&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/AtomicBot-ai/atomic-agent/releases&lt;/a&gt; - v0.6.7 and desktop-v0.0.1 published 2026-10-07&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://atomicagent.io&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=atomicagent.io&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://atomicagent.io&lt;/a&gt; - product site: installer, benchmark banner, and the consumer-ecosystem surface&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/AtomicBot-ai/atomic-agent/main/eval-agents/docs/GAIA-L1-EXPERIMENT.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/AtomicBot-ai/atomic-agent/main/eval-agents/docs/GAIA-L1-EXPERIMENT.md&lt;/a&gt; - the benchmark write-up: environment, dataset, and artifact publication&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/AtomicBot-ai/atomic-agent/main/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/AtomicBot-ai/atomic-agent/main/README.md&lt;/a&gt; - the full README: loop design, TurboQuant claims, import matrix, privacy and egress&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/search?query=atomic-agent&amp;amp;hitsPerPage=5&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/search?query=atomic-agent&amp;hitsPerPage=5&lt;/a&gt; - the no-thread community-footprint check as of 2026-10-10&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
