DeepSeek V4.1 Flash manages agent memory 10× less efficiently than V4 Flash (Hermes, MarathonMemBench)

#46
by ovsale - opened

MarathonMemBench is a black-box benchmark for agent long-term memory: an LLM persona talks to the agent for about an hour, tells it ~150 facts across several life topics — renovation, car, trip, health, family — keeps returning to old topics with edits, then quizzes it on what it remembers. I run Hermes, OpenClaw and my own agent on it. Repo: https://github.com/ovsale/marathon-mem-bench. Write-up with baseline numbers, published before V4.1 came out: https://dev.to/alexander_ovsov_f85e6b68f/marathonmembench-a-new-kind-of-benchmark-for-testing-your-llm-agents-memory-5a3l

I ran the same scenario on Hermes twice: on V4.1 Flash and, as a control, on V4 Flash 0731 (or 0423). Same Hermes, same config, same prompts. The persona says ~7.2k tokens over the whole run. Both runs scored 36/36 on the exam. Here is what Hermes wrote to memory:

V4 Flash V4.1 Flash
Topic memory files 9 11
Total size 13 KB (~4.6k tokens) 161 KB (~42k tokens)
Largest file renovation-ledger.md, 5.5 KB renovation.md, 70 KB

Twelve times more memory for the same conversation, and the extra is not facts.

What's inside. The old model keeps a table of current values with short notes; edit history lives in one line inside a cell: $42→$55/m², was 80 L. The new model keeps a chronicle: the same table plus a "Decisions log" where every step is retold as a paragraph — committed before and after, zone subtotal, projection — plus "open questions" and advice to itself. The user said "Shimano 105 12-speed"; the file has eight paragraphs on groupset generations and bottom-bracket standards. Calories per 85 km, watts at 25 km/h. 109 "TBC" placeholders for facts the user never gave, including a "Fit" section of eight empty lines. Preambles like "why this file exists" and "how to keep it".

This isn't memory. The model writes down everything it knows and thinks. Then it chokes on it: reads its own 70 KB back into context over and over, writes more. Over the run it lives through roughly 10× more context than V4 Flash on the same conversation, and the run took 99 minutes instead of 53–70.

Reproduced on my own agent — different prompt, different memory layout: 21 KB → 350–500 KB, one hour → two to three, 788k output tokens instead of 141k. The files look the same, down to identical conclusions drawn from the same facts — both agents decided from "$1,900 for a carbon Domane" that it must be an older generation, both decided 54 cm at 183 cm is two sizes below Trek's chart.

It also researches what nobody asked for. The scenario needs no internet at all: the persona tells facts about her own life and never asks to look anything up. V4 Flash never went online. 4.1 did about ten searches, unprompted and for the future: Georgian excise rates by engine volume for a car not yet chosen, city climate, eight engine options under a $15k budget. All of it landed in memory as tables.

The prompt doesn't help. My agent's system prompt already had rules of this kind — don't save reasoning, don't save what can be found on the internet. What the benchmark exposed was so blatant that I tried to patch exactly that, with a dedicated section on top:

## Important behavioral and information saving rules
- Don't research what the user didn't ask for.
- Don't save what can be found again, unless the user asks.
- Don't save your own intermediate ideas, evaluations, calculations, or a log of your work.

Zero effect. Not less — the same.

Discussion #39 here reports the same shape on a coding task: the model has the answer by turn 6–12, then spends 20–40 more turns verifying. Different task, same missing stopping criterion.

I can share the Hermes memory exports from both runs. Given that V4 Pro was just kept alive on user demand, keeping deepseek-v4-flash-0731 available as an explicit model ID would help everyone whose agents were tuned on it.

Anyone else seeing this on 4.1?

yeah, we love deepseek ^^ but i feel like the visual reasoning quality has regressed compared to the dedicated vision model :C

I'm seeing similar run-away train type behavior as well.

This model has severe ADHD, and has a tendency to be overly eager with diving into work that wasn't asked of it.

I love the direction these new architectures are going - hopefully the less than stellar results of this model aren't telling of the architecture.

Thanks — "eager to dive into work that wasn't asked" is exactly it. Do you have any numbers on it, even rough — tool calls or tokens per task on 4.1 vs 4.0? A second independent measurement would help a lot here. On the architecture: I'd guess this is post-training, not the encoder-decoder itself — the shape of the output (decision logs, "what I'd do in his position") looks like a reward for thoroughness, not a capability limit.

Sign up or log in to comment