Newsletter image

Subscribe to the Newsletter

Join 10k+ people to get notified about new posts, news and tips.

Do not worry we don't spam!

By pressing the Subscribe button, you confirm that you have read and are agreeing to our Privacy Policy and Terms of Use

Search

GDPR Compliance

We use cookies to ensure you get the best experience on our website. By continuing to use our site, you accept our use of cookies, Privacy Policy, and Terms of Service.

Nous Research - AI Agent, Conversational AI

Hermes Agent

Nous Research Hermes Agent writes its own skills, refines them in use and models you across sessions. MIT, model-agnostic, seven sandbox backends. And it corrects a claim we made.

License MIT
License MIT
TL;DR
  • Writes its own skills from experience and refines them during use
  • Model-agnostic: any provider, switch with one command, 300+ via Nous Portal
  • Seven execution backends from local to Docker, SSH, Modal and Vercel Sandbox

Hermes Agent has 250,195 stars, which makes it larger than DeepSeek Harness and OpenCode, and we have never written about it. It is MIT, it is from Nous Research, and it is not a coding agent. It is a general agent whose distinguishing claim is a learning loop: it writes its own skills from experience, refines them as it uses them, and builds a model of you across sessions. That is a different bet from every other project in this series, and it deserves more scrutiny than a star count.

What makes it different

Nous describes it as "the self-improving AI agent built by Nous Research", and claims "it's the only agent with a built-in learning loop". The specifics behind that phrase are what matter, so here they are as the project states them: it "creates skills from experience, improves them during use, nudges itself to persist knowledge, searches its own past conversations, and builds a deepening model of who you are across sessions."

Unpack that into four mechanisms. After finishing a complex task it can write a reusable procedure and keep it, which is skill creation rather than a static tool list. Those procedures then get refined as they are used again. A memory system with agent-driven curation decides what is worth keeping rather than logging everything. And a cross-session user model, built with full-text search over past conversations plus LLM summarisation, accumulates a picture of how you work.

Whether that produces genuine compounding improvement or an elaborate cache is exactly the question, and nobody outside Nous has published an answer. Treat the learning loop as an interesting architecture with no independent evaluation, not as a demonstrated result.

It is not a coding agent, and that matters here

Everything else in this series wraps a model to edit your repository. Hermes runs in a terminal but also in Telegram, Discord, Slack, WhatsApp and Signal, plus a desktop interface. It has a cron scheduler for unattended automations and can spawn subagents for parallel work.

The execution model is the part that will interest anyone self-hosting. It supports seven terminal backends: local, Docker, SSH, Singularity, Modal, Daytona and Vercel Sandbox. That range says the design goal is running untrusted agent-generated commands somewhere contained, which is a more serious posture than most agents take.

It ships more than forty built-in tools, connects to any MCP server, and its skills are compatible with the agentskills.io open standard. That last detail is worth flagging against the pattern this series keeps finding: while four coding harnesses each invented an incompatible plugin system, Hermes adopted an existing open standard for the same job.

No model lock-in

The obvious suspicion with an agent from a lab that trains models is that it will steer you toward them. It does not. The project is explicit: "Use any model you want, Nous Portal, OpenRouter, OpenAI, your own endpoint, and many others. Switch with hermes model, no code changes, no lock-in." Nous Portal alone reaches 300-plus models.

Compare ZCode, where we found official-GLM concepts baked into shared code and a database migration dedicated to the vendor's own models. Hermes is from a lab with its own well-known model family and has less vendor gravity in its code than a harness from a lab that mostly sells API access. That is not what you would predict.

The Issues test, and a correction

In the harness wars piece we wrote that every harness built by a frontier lab has its issue tracker closed, and that the only one you can file a bug against is the one no lab owns. Hermes complicates that.

ProjectWhoIssuesOpenForks
Hermes AgentNous ResearchEnabled47,23553,425
OpenCodeIndependentEnabled6,28127,948
HerdrIndependentEnabled3893,201
DeepSeek HarnessDeepSeekDisabled028,907
grok-buildxAIDisabled05,114
ZCodeZ.aiDisabled02,191

Nous Research is a lab, and its tracker is not merely open, it is carrying 47,235 open issues against 53,425 forks. So the honest version of our earlier finding is narrower than we wrote it: the labs that close their trackers are the ones whose main business is selling access to a hosted frontier model. An open-weights research collective behaves like the independents. That is a distinction about commercial posture, not about being a lab, and our original sentence was too broad.

One caveat in the other direction: 47,235 open issues is not obviously a healthy number. It can mean an engaged community, or a tracker nobody is triaging. We cannot tell which from the outside, and neither can you, so do not read an open tracker as a support guarantee.

Running it

Installation is a single script, and the first thing worth doing is pointing it at a model you already pay for.

curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes

Then hermes model switches providers. As always with a piped install script, read it before you run it.

If you want the memory and skill machinery to be more than a demo, give it real work over real time. The entire premise is accumulation across sessions, so a ten-minute evaluation will show you a competent tool-using agent and none of the thing that makes it interesting.

Where it sits in this series

Three layers have now appeared, and they are not competing with each other.

LayerWhat it ownsExamples
RuntimeThe terminals agents live inHerdr
HarnessThe agent loop over your repositoryOpenCode, DeepSeek Harness, ZCode, grok-build
General agentLong-running assistant with memoryHermes Agent

You can run Hermes inside Herdr. You can run a coding harness beside it. The layer that has become a vendor battleground is the middle one, and it is the middle one where nothing interoperates.

Who should run it

Run Hermes if you want a persistent assistant rather than a per-repository coding tool, if the messaging-platform surfaces fit how you work, or if you want a sandboxed execution story with seven backends to choose from. Run it if you are curious whether a learning loop compounds, and are willing to be the experiment.

Do not reach for it to edit a codebase. That is what the harnesses in this series are for, and Hermes is not trying to be one.

Sources and further reading

Ten minutes: install it, connect it to a model you already have, and ask it to do something it will need to repeat next week. Then come back in a week and see whether it wrote itself a skill for it. That single test tells you more about whether the learning loop is real than any benchmark currently available.

Tested on: not independently tested. Architecture, the learning-loop description, backend list, tool counts and model-agnosticism are quoted from the project's own README. No independent evaluation of the self-improvement claim exists at the time of writing. Star, fork and issue figures for all six projects in the comparison table were read from the GitHub API on the date below.
Date checked: 2026-09-30

Prev Article
Herdr
Next Article
Jeff

Related to this topic: