Independent agent benchmarks

I run reproducible benchmarks on AI agent systems.

I measure what agents keep and retrieve, what they hand to the prompt, how much it costs, and what breaks once MCP enters the execution path. I am an independent tester with no affiliation to the systems I measure. Memory is where I started, comparing options like Honcho, Mem0, Hindsight, Mnemosyne and Sibyl, and the results, caveats and rerun files are all below.

Systems compared5Honcho, Mem0, Hindsight, Mnemosyne and Sibyl
Retrieval questions350Same question set asked to every system
Test window365 daysSynthetic business year of memory writes
Public artifacts75Reports, scripts and raw run files, all rerunnable

No sponsor, no affiliation with any system tested. The results live further down, next to the settings, caveats and rerun files that produced them.

Tested stack

Memory is tested across agents, plugins and MCP surfaces.

Benchmark proof

The claim is backed by a corpus, competitors and security runs.

Final records191kScale365 memory corpus
Write calls209kSynthetic business year
Competitors4Honcho, Mem0, Hindsight, Mnemosyne
Security baseline31 runs4 confirmed symlink issues

Narrative tests

Start with the story, then open the files.

The main pages explain what is being compared, why memory, MCP and plugin settings matter, and where the caveats are. The archive underneath keeps the raw reports and runners close to the claim.

Scale365 snapshot

The chart is readable, the caveat stays visible.

Each system is tested in a documented mode: some in a full structured-state stack, others in reproducible local modes, and every page explains what a more expensive or smarter configuration would change.

Read the methodology note

Retrieval benchmark

Questions passed

350Scale365
Sibyl Scale365 retrieval350/350
Hindsight retrieval152/350
Mem0 retrieval92/350
Mnemosyne retrieval5/350

Test archive

75 public entries across benchmarks, security, migration and rerun files.

The old dashboard is no longer the only place to find the work. The archive lists the reports, scripts and raw artifact index behind the public pages.

View full archive

Long-form writing

The context before the claim.

MCP and Hermes security

Agent tooling is also an attack surface.

Research paths

Follow the threads behind the benchmarks.