<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>agentic-llm on tomrochette.com</title>
    <link>https://tomrochette.com/tags/agentic-llm/</link>
    <description>Recent content in agentic-llm on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Sun, 11 Oct 2026 02:01:18 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/agentic-llm/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>ARIS</title>
      <link>https://tomrochette.com/agents/automated-research/aris/</link>
      <pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/automated-research/aris/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>automated-research</category><category>agentic-llm</category><category>machine-learning</category><category>open-source</category><category>skills</category>
      <description>&lt;p&gt;ARIS (Auto-Research-In-Sleep) is wanshuiyin&amp;rsquo;s MIT-licensed set of markdown-defined skills that turns Claude Code, Codex CLI, or another coding agent into an autonomous ML research loop, with a reviewer model from a different family critiquing every artifact at gated passes.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;ARIS is this category&amp;rsquo;s first member whose judge is required to be a different model family: the executor drives and an independent reviewer (Codex MCP by default) demands revisions, an arrangement its technical report builds against the category&amp;rsquo;s shared failure mode, the plausible unsupported success.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A methodology shipped as 82 markdown skills (83 in the standalone CLI bundle) with no framework and no lock-in, created 2026-03-10.&#xA;The technical report (arXiv 2605.03042, May 2026, by Ruofeng Yang, Yongcan Li, and Shuai Li) describes three layers: an execution layer with more than 65 reusable markdown skills, MCP model integrations, and a persistent research wiki; a review layer where a reviewer from a different model family critiques intermediate artifacts and requests revisions; and deployment experience from long unattended runs.&#xA;The skills cover the research arc end to end: idea discovery, experiment queues (including SSH multi-seed job queues), paper writing, citation audits, integrity forensics, and Overleaf sync.&#xA;Adapters exist for Claude Code, Codex CLI, Cursor, Trae, Antigravity, GitHub Copilot CLI, OpenClaw, and DeepSeek Harness.&#xA;Beyond the skills, the project ships a standalone Rust CLI (ARIS-Code, v0.4.28 on 2026-09-28, mid a 24-release train) and Claude Code, Codex CLI, and DeepSeek Harness plugins.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Active: 17,250 stars, 1,446 forks, created 2026-03-10, last push 2026-10-07, as of 2026-10-10.&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=wanshuiyin/Auto-claude-code-research-in-sleep&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=wanshuiyin/Auto-claude-code-research-in-sleep&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=wanshuiyin/Auto-claude-code-research-in-sleep&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&lt;p&gt;&lt;strong&gt;The Western discussion footprint is zero: a Hacker News search for the repository name returns no stories as of 2026-10-10, and adoption runs through a PaperWeekly feature and the Chinese-language community, the channel pattern DeepAnalyze shows; the Agon paper&amp;rsquo;s own count puts ARIS at 79 roles and 1,157.4 KiB of prompts, roughly four times Agon&amp;rsquo;s surface.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The cross-model rule attacks a failure the kernel columns never face: a reviewer from a different family is less likely to inherit the executor&amp;rsquo;s framing, which is where the report locates unsupported claims.&lt;/li&gt;&#xA;&lt;li&gt;The skill set covers the whole arc, including integrity tooling (citation audit, fabrication forensics) most research loops leave to the human.&lt;/li&gt;&#xA;&lt;li&gt;No lock-in: the same markdown skills port across eight harnesses, so the loop survives any single vendor&amp;rsquo;s churn.&lt;/li&gt;&#xA;&lt;li&gt;The maintenance cadence is exceptional: 24 CLI releases between May and September 2026, including a same-week bridge when codex-cli 0.154 removed the reviewer&amp;rsquo;s entry point.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The report&amp;rsquo;s deployment claims are self-published, and no independent replication surfaced in this run&amp;rsquo;s searches.&lt;/li&gt;&#xA;&lt;li&gt;The judge is still a model: cross-model review raises the bar but is not a kernel, a grader, or a benchmark.&lt;/li&gt;&#xA;&lt;li&gt;The prompt surface is large (79 roles, 1,157.4 KiB by the Agon paper&amp;rsquo;s count), which makes behavior harder to inspect than Agon&amp;rsquo;s 18 roles.&lt;/li&gt;&#xA;&lt;li&gt;The zero-story Hacker News footprint says the 17.2k stars measure attention in one language community, not broad engineering adoption.&lt;/li&gt;&#xA;&lt;li&gt;The loop depends on moving harness internals, as the codex-cli 0.154 breakage demonstrated.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and MIT; there is no paid tier, so pricing does not apply.&#xA;Costs are your own model subscriptions or API keys (Claude executor plus a Codex reviewer by default, with ModelScope-hosted and local-model combinations documented) and your compute.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/agon/&#34; &gt;Agon&lt;/a&gt;: the Claude Code plugin running producer-critic factories; ARIS is the skills-based loop whose reviewers come from other model families and which ports across harnesses.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/karpathy-autoresearch/&#34; &gt;Karpathy Autoresearch&lt;/a&gt;: the founding keep-or-revert loop whose judge is a measured loss; ARIS&amp;rsquo;s judge is another model&amp;rsquo;s critique.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/openresearch/&#34; &gt;OpenResearch&lt;/a&gt;: the workspace that turns your own agents into researchers with an archived evidence tree; ARIS adds adversarial review gates to the same idea.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for ML researchers who want an unattended experiment loop with cross-model review, on the harness they already use, and for studying review-gate design against plausible unsupported successes.&lt;/strong&gt;&#xA;Not for anyone who needs a machine verifier behind results, a small inspectable prompt surface, or an independently replicated evaluation.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-10 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/agon/&#34; &gt;Agon&lt;/a&gt; - the producer-critic plugin that measured ARIS&amp;rsquo;s prompt surface&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/karpathy-autoresearch/&#34; &gt;Karpathy Autoresearch&lt;/a&gt; - the keep-or-revert ancestor of the loop family&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/deepanalyze/&#34; &gt;DeepAnalyze&lt;/a&gt; - the other thin-Western-footprint member, judged by itself where ARIS is judged by another model&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/automated-research-feature-matrix/&#34; &gt;Automated Research Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep&lt;/a&gt; - repository and README: skill set, harness adapters, executor-reviewer split, release train (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/wanshuiyin/Auto-claude-code-research-in-sleep&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/wanshuiyin/Auto-claude-code-research-in-sleep&lt;/a&gt; - stars, forks, created and pushed dates, MIT license, and topics for the as-of status (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/wanshuiyin/Auto-claude-code-research-in-sleep/main/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/wanshuiyin/Auto-claude-code-research-in-sleep/main/README.md&lt;/a&gt; - the README source: 82 skills, adapter list, ARIS-Code v0.4.28 details, and the codex-cli 0.154 bridge note (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2605.03042&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2605.03042&lt;/a&gt; - the technical report: cross-model adversarial collaboration, the plausible-unsupported-success failure mode, and the three-layer architecture (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/wanshuiyin/Auto-claude-code-research-in-sleep/releases?per_page=5&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/wanshuiyin/Auto-claude-code-research-in-sleep/releases?per_page=5&lt;/a&gt; - the release train through v0.4.28 (2026-09-28) (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/search?query=claude-code-research-in-sleep&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/search?query=claude-code-research-in-sleep&lt;/a&gt; - the zero-story Hacker News footprint (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>AutoResearchClaw</title>
      <link>https://tomrochette.com/agents/automated-research/autoresearchclaw/</link>
      <pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/automated-research/autoresearchclaw/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>automated-research</category><category>agentic-llm</category><category>open-source</category><category>paper-generation</category>
      <description>&lt;p&gt;AutoResearchClaw is the MIT-licensed 23-stage pipeline from the aiming-lab organization that turns a one-line research idea into a compile-ready paper, with sandbox experiments, multi-agent debate, a four-layer citation-verification layer, and seven human-in-the-loop modes.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;AutoResearchClaw is the idea-to-paper pipeline that treats failure and fabrication as first-class problems: its executor heals through Pivot/Refine loops, its citations pass arXiv, CrossRef, DataCite, and LLM checks before delivery, and its 54.7 percent win over AI Scientist v2 is measured on its own ARC-Bench, a self-run benchmark the README has since widened from 25 to 55 topics.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A Python pipeline (created 2026-03-15) run from the &lt;code&gt;researchclaw&lt;/code&gt; CLI, standalone, through an OpenClaw bridge to Discord, Telegram, Lark, or WeChat, or on any ACP-compatible agent backend (Claude Code, Codex CLI, Copilot CLI, Gemini CLI, Kimi CLI).&#xA;The arXiv paper (2605.20025, by Jiaqi Liu, Shi Qiu, Mairui Li, Bingzhou Li, Haonian Ji, Siwei Han, Xinyu Ye, and Peng Xia) presents five mechanisms: structured multi-agent debate, a self-healing executor with a Pivot/Refine decision loop, verifiable result reporting against fabricated numbers and hallucinated citations, human-in-the-loop collaboration across seven intervention modes, and cross-run evolution that converts past failures into future safeguards.&#xA;Deliverables per run: a paper draft, conference-ready LaTeX (NeurIPS, ICML, ICLR templates), a BibTeX file with references pulled from OpenAlex, Semantic Scholar, and arXiv, a verification report, sandbox experiment code and metrics, charts, multi-agent reviews, and evolution lessons.&#xA;MetaClaw adds cross-run learning (pipeline failures become structured lessons injected into later runs), and the skills library loads 20 preloaded skills plus community contributions.&#xA;The companion ARC-Bench dataset ships on Hugging Face under the AIMING-Lab-UNC account, widened at v0.5.0 (May 2026) to 55 topics across machine learning, high-energy physics, quantum, biology, and statistics.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Dormant since August: 14,618 stars, 1,706 forks, created 2026-03-15, last push 2026-08-19, latest release v0.5.0 on 2026-05-20, as of 2026-10-10.&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=aiming-lab/AutoResearchClaw&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=aiming-lab/AutoResearchClaw&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=aiming-lab/AutoResearchClaw&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&lt;p&gt;&lt;strong&gt;The repository has gone quiet while attention held: no commits in eight weeks and no release in five months as of 2026-10-10, a Hacker News footprint of two stories at two and one points with zero comments, and the eight showcase papers are self-published, so the 14.6k stars currently measure attention the repository is no longer feeding.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Citation integrity is a design point, not an afterthought: the four-layer check (arXiv, CrossRef, DataCite, LLM) kills fabricated references before delivery, the failure class the no-kernel columns are most exposed to.&lt;/li&gt;&#xA;&lt;li&gt;Failure handling is architectural: Pivot/Refine turns failed experiments into information, and MetaClaw accumulates them into reusable skills.&lt;/li&gt;&#xA;&lt;li&gt;The seven-mode human-oversight ladder (full-auto through step-by-step) is the widest in this category, an answer to the trust problem rather than a denial of it.&lt;/li&gt;&#xA;&lt;li&gt;ARC-Bench gives the loop a rubric-scored eval set spanning five domains instead of a single ML sandbox.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The 54.7 percent improvement over AI Scientist v2 is self-reported on the project&amp;rsquo;s own benchmark, with no independent replication surfaced in this run&amp;rsquo;s searches.&lt;/li&gt;&#xA;&lt;li&gt;Development is quiet: last push 2026-08-19 and last release 2026-05-20, which is a poor sign for a 23-stage pipeline whose stages depend on external APIs and agent backends.&lt;/li&gt;&#xA;&lt;li&gt;The two zero-comment Hacker News stories say there is no practitioner debate to check the claims against.&lt;/li&gt;&#xA;&lt;li&gt;A 23-stage pipeline is heavy to adopt and heavier to fork: the runtime surface (LLM providers, Docker, LaTeX, five agent backends) is wide.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and MIT; there is no paid tier, so pricing does not apply.&#xA;Costs are your own LLM API keys and compute (a Docker executor with GPU, MPS, or CPU auto-detection).&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/agon/&#34; &gt;Agon&lt;/a&gt;: the other idea-to-paper column; Agon is a Claude Code plugin with adversarial critics and a failure taxonomy, AutoResearchClaw a standalone pipeline whose added layer is citation verification and a human-oversight ladder.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/aris/&#34; &gt;ARIS&lt;/a&gt;: the markdown-skill loop ported across harnesses with cross-model review gates; AutoResearchClaw is a fixed 23-stage pipeline you run rather than a method you install.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/gpt-researcher/&#34; &gt;GPT Researcher&lt;/a&gt;: the cited-report baseline with no experiments; AutoResearchClaw runs sandbox experiments and ships LaTeX.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for studying how an autonomous pipeline can make citation fabrication and silent failure first-class problems, and for teams that want idea-to-paper runs with explicit human-oversight modes.&lt;/strong&gt;&#xA;Not for anyone who needs an actively maintained tool (eight weeks without a commit) or independently verified benchmark claims.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-10 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/agon/&#34; &gt;Agon&lt;/a&gt; - the producer-critic idea-to-paper column and the Agon paper&amp;rsquo;s comparator&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/aris/&#34; &gt;ARIS&lt;/a&gt; - the skills-based cross-model loop in the same family&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/gpt-researcher/&#34; &gt;GPT Researcher&lt;/a&gt; - the cited-report baseline AutoResearchClaw extends to experiments&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/automated-research-feature-matrix/&#34; &gt;Automated Research Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/aiming-lab/AutoResearchClaw&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/aiming-lab/AutoResearchClaw&lt;/a&gt; - repository and README: 23-stage pipeline, deliverables, OpenClaw and ACP-backend support, release news (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/aiming-lab/AutoResearchClaw&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/aiming-lab/AutoResearchClaw&lt;/a&gt; - stars, forks, created and pushed dates, MIT license, and topics for the as-of status (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/aiming-lab/AutoResearchClaw/main/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/aiming-lab/AutoResearchClaw/main/README.md&lt;/a&gt; - the README source: stage outputs, citation-verification layers, HITL system, ARC-Bench v0.5.0 widening (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2605.20025&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2605.20025&lt;/a&gt; - the paper: five mechanisms, the 54.7 percent AI Scientist v2 comparison on ARC-Bench, and the seven intervention modes (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/aiming-lab/AutoResearchClaw/releases?per_page=5&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/aiming-lab/AutoResearchClaw/releases?per_page=5&lt;/a&gt; - the release train through v0.5.0 (2026-05-20) (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/search?query=AutoResearchClaw&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/search?query=AutoResearchClaw&lt;/a&gt; - the two zero-comment Hacker News stories behind the discussion-footprint claim (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://huggingface.co/datasets/AIMING-Lab-UNC/ARC-Bench&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=huggingface.co&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://huggingface.co/datasets/AIMING-Lab-UNC/ARC-Bench&lt;/a&gt; - the ARC-Bench dataset page under the AIMING-Lab-UNC account (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
    <item>
      <title>DeepAnalyze</title>
      <link>https://tomrochette.com/agents/automated-research/deepanalyze/</link>
      <pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/automated-research/deepanalyze/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>automated-research</category><category>agentic-llm</category><category>data-science</category><category>open-source</category>
      <description>&lt;p&gt;DeepAnalyze is Renmin University of China and Tsinghua&amp;rsquo;s MIT-licensed 8B-parameter agentic LLM, trained to run the data-science research loop end to end, from raw files to an analyst-grade report, with no harness beyond a code sandbox.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;DeepAnalyze is this category&amp;rsquo;s first member whose loop is carried by a trained model rather than a prompted harness: the plan, act, and self-reflect structure is learned in the weights, which makes it the cheapest loop here to run and the only one with no external verifier anywhere in it.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;DeepAnalyze-8B, released October 21, 2025 with a paper (arXiv 2510.16872, online October 19), open weights, code, and a 500,000-example instruction dataset (DataScience-Instruct-500K).&#xA;Training uses a curriculum-based agentic paradigm that mimics a data scientist&amp;rsquo;s learning trajectory, with training trajectories synthesized from data sources, so the model learns to question, prepare, analyze, model, visualize, and report without workflow code.&#xA;Given structured data (databases, CSV, Excel), semi-structured files (JSON, XML, YAML), or unstructured text, it runs open-ended data research and produces analyst-grade reports.&#xA;Deployment is &lt;code&gt;vllm serve DeepAnalyze-8B&lt;/code&gt; plus a WebUI or a Docker-sandboxed WebUI v2, with a hosted API available through HeyWhale key applications.&#xA;The authors are Shaolei Zhang, Ju Fan, Meihao Fan, Guoliang Li, and Xiaoyong Du (Renmin University of China with Tsinghua), and the ecosystem has widened: DA-Studio, the system behind WebUI v2, was accepted to the VLDB 2026 demonstration track, CoDA-Bench evaluates code agents on data-intensive analytical tasks, DeepPrep extends the loop to data preparation, and EvoOntology adds a self-evolving ontology layer.&#xA;It served as the official agent for the 2026 China Collegiate Computer Design Contest&amp;rsquo;s Big Data track.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Active: 4,678 stars, 739 forks, created 2025-10-11, last push 2026-09-23, as of 2026-10-10.&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=ruc-datalab/DeepAnalyze&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=ruc-datalab/DeepAnalyze&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=ruc-datalab/DeepAnalyze&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&lt;p&gt;&lt;strong&gt;The usage footprint is far thinner than the star count: the Hugging Face weights show 203 downloads and 93 likes as of 2026-10-10, and a Hacker News search returns zero stories, so adoption runs through the Chinese data-science community (PaperWeekly, WeChat, the collegiate contest) rather than the Western harness ecosystem.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The full loop is open and reproducible: weights, code, and training data, at a size one GPU serves.&lt;/li&gt;&#xA;&lt;li&gt;The paper claims it outperforms workflow-based agents built on the most advanced proprietary LLMs at data tasks.&lt;/li&gt;&#xA;&lt;li&gt;The ecosystem compounds beyond one repo: a benchmark (CoDA-Bench), a preparation companion (DeepPrep), a VLDB-accepted serving system (DA-Studio), and an ontology layer (EvoOntology).&lt;/li&gt;&#xA;&lt;li&gt;Deployment is genuinely self-hosted: vllm on your own GPU, sandboxed execution, no external service required.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The headline claim is self-reported in the paper&amp;rsquo;s own evaluation, and no independent replication surfaced in this run&amp;rsquo;s searches.&lt;/li&gt;&#xA;&lt;li&gt;A single 8B model plans, executes, and judges its own work inside the sandbox, so there is no external verifier anywhere in the loop.&lt;/li&gt;&#xA;&lt;li&gt;The project homepage ships with its template placeholders unfilled, thin for a project of the claimed importance.&lt;/li&gt;&#xA;&lt;li&gt;4,678 stars against 203 model downloads and zero Hacker News stories says the stars measure attention, not usage.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and open source: MIT-licensed code, open weights on Hugging Face, and no paid tier, so pricing does not apply.&#xA;Costs are your own GPU, or a HeyWhale-hosted API key, which is granted by application with no public prices on the fetched pages.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/gpt-researcher/&#34; &gt;GPT Researcher&lt;/a&gt;: the prompted web-research loop that cites the open web; DeepAnalyze is data-grounded and carried by a trained model.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/karpathy-autoresearch/&#34; &gt;Karpathy Autoresearch&lt;/a&gt;: the keep-or-revert ancestor whose judge is a measured loss; DeepAnalyze&amp;rsquo;s judge is itself.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/openresearch/&#34; &gt;OpenResearch&lt;/a&gt;: the workspace that turns prompted coding agents into researchers; DeepAnalyze collapses the loop into one model you can fine-tune.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for teams that want a self-hosted data-analysis agent they can fine-tune, and for researchers studying learned agentic loops against prompted ones.&lt;/strong&gt;&#xA;Not for audit-sensitive analysis: the loop&amp;rsquo;s only judge is the model that wrote the analysis.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-10 - Created.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/karpathy-autoresearch/&#34; &gt;Karpathy Autoresearch&lt;/a&gt; - the measured-judge ancestor of the loop family&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/gpt-researcher/&#34; &gt;GPT Researcher&lt;/a&gt; - the prompted research-report baseline&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/openresearch/&#34; &gt;OpenResearch&lt;/a&gt; - the workspace implementation of the loop on your own agents&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/automated-research/automated-research-feature-matrix/&#34; &gt;Automated Research Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/ruc-datalab/DeepAnalyze&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/ruc-datalab/DeepAnalyze&lt;/a&gt; - repository, README, task surface, deployment paths, and the ecosystem news timeline (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/ruc-datalab/DeepAnalyze&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/ruc-datalab/DeepAnalyze&lt;/a&gt; - stars, forks, created and pushed dates, and the MIT license for the as-of status (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2510.16872&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2510.16872&lt;/a&gt; - the paper: title, authors, October 19 2025 online date, curriculum-based agentic training, and the frontier-workflow-agents comparison claim (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://huggingface.co/RUC-DataLab/DeepAnalyze-8B&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=huggingface.co&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://huggingface.co/RUC-DataLab/DeepAnalyze-8B&lt;/a&gt; - the open weights page (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://huggingface.co/api/models/RUC-DataLab/DeepAnalyze-8B&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=huggingface.co&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://huggingface.co/api/models/RUC-DataLab/DeepAnalyze-8B&lt;/a&gt; - the 203 downloads and 93 likes usage footprint (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://ruc-deepanalyze.github.io/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=ruc-deepanalyze.github.io&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://ruc-deepanalyze.github.io/&lt;/a&gt; - the project homepage, fetched with its template placeholders unfilled (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/search?query=DeepAnalyze&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/search?query=DeepAnalyze&lt;/a&gt; - the zero-story Hacker News footprint (fetched 200, 2026-10-10)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
