80.5% of revisions with no overall gain still improve at least one capability.
Existing self-evolving memory methods revise a memory program using one overall score. But memory is multidimensional, and that score blurs what each revision actually achieved.
Figure 1. Why capability-driven evolution matters. Analysis of M⋆ and EvolveMem evolution traces on BEAM.
We derive five representative capabilities from a broad review of agent memory benchmarks and memory systems. Then we annotate every question with the one it primarily requires, plus up to two secondary ones. These labels only organize feedback during evolution and are never given to the memory program at test time.
F
Factual retrieval
Preserve and retrieve explicit facts, values, and other localized details.
T
Temporal tracking
Order events, determine recency, track states and knowledge updates.
P
Preference extraction
Recover and apply user preferences, instructions and constraints.
M
Multi-session synthesis
Aggregate and synthesize evidence across multiple sessions or events.
A
Adversarial
Abstain when memory lacks supporting evidence, or cautiously resolve contradictory or misleading evidence.
02 / THE METHOD
Explore separately. Improve together.
PrisMem evolves an executable memory program in three stages: build a base, grow capability specialists, then integrate them.
Figure 2. Overview of PrisMem: cold start, capability refinement, and trace-guided integration.
ZOOM IN · STAGE II
Every round makes two choices before any code changes.
Which capability to refine, and which of its failures to learn from.
CHOICE 1 · DEPENDENCY-AWARE SELECTION
Which capability to refine?
urgency = own error + strongest pressure on another capability1 + times already picked
F Factual
picked 1× → ÷2
T Temporal
not picked yet
P Preference
picked 1× → ÷2
M Synthesis
picked 2× → ÷3
A Adversarial
not picked yet
A capability is chosen because it is weak itself, or because it holds another capability back. Each past pick lowers its priority, so plateaued capabilities step aside: here T wins although M has the larger raw urgency (dashed).
CHOICE 2 · HISTORY-GUIDED DIAGNOSIS
Which failures to learn from?
priority = how well earlier programs answered ittimes it, or a similar case, was already picked
A case that used to be solved and now fails shows exactly what a recent revision broke. Discounting repeats keeps the five cases diverse.
Schematic diagrams with illustrative values. They show the idea behind the two selection rules (Eqs. 1–7 in the paper).
ZOOM IN · STAGE III
Every integration rests on evidence and a plan.
Paired differential cases show where programs behave differently; the integration plan decides what to carry over.
EVIDENCE · PAIRED DIFFERENTIAL CASES
Where do two programs really differ?
SPECIALISTvsBASE
F FactualP131vsP170.5+ bring in
T TemporalP70.4vsP171◆ protect
P PreferenceP151vsP170+ bring in
M SynthesisP101vsP170+ bring in
A AdversarialP70.25vsP171◆ protect
Specialist wins → bring in. Base P17 wins → protect.
PLAN · LINEAGE-ORDERED INTEGRATION
What to take, and what to guard?
While contrasting cases serve as diagnostic indicators targeting specific program components, the plan leverages these insights to localize code modules and formulate a step-by-step integration strategy.
03 / LIVE DEMO
Watch memory evolve.
One real 20-round run on BEAM, replayed step by step: which program was revised, what failed, and what changed.
+0+10% vs P5new specialist
04 / THE EXPERIMENTAL RESULTS
Built on shorter histories. Tested at a million tokens.
Static memorySelf-evolving baselinesPrisMem
Table 2 visualized · Qwen3.8-27B · 3 runs · Bars show means; whiskers show ±1 standard deviation.
BEAM-1M
Capability scores (%) ↑
Factual retrieval
PrisMem: 74.17 ± 4.87 %
Temporal tracking
PrisMem: 57.26 ± 2.15 %
Preference extraction
PrisMem: 75.58 ± 3.91 %
Multi-session synthesis
PrisMem: 54.12 ± 3.37 %
Adversarial
PrisMem: 64.32 ± 6.51 %
Overall
% · higher is better
PrisMem: 63.39 ± 1.02 %
Average token usage per question (K) ↓
lower is better
PrisMem: 100.59 ± 5.02 K tokens / question
PrisMem leads all five capability metrics on BEAM-1M and improves Overall by 11.53 pp over the strongest static baseline (10.54 pp over the strongest self-evolving baseline), while using fewer tokens per question than Mem0, A-MEM, and HippoRAG2.
LongMemEval-M
Capability scores (%) ↑
Factual retrieval
PrisMem: 82.29 ± 3.93 %
Temporal tracking
PrisMem: 75.00 ± 3.13 %
Preference extraction
PrisMem: 47.78 ± 1.92 %
Multi-session synthesis
PrisMem: 82.29 ± 3.61 %
Adversarial
PrisMem: 96.67 ± 5.77 %
Overall
% · higher is better
PrisMem: 75.50 ± 1.80 %
Average token usage per question (K) ↓
lower is better
PrisMem: 2232.65 ± 32.72 K tokens / question
PrisMem leads or ties four of five capability metrics on LongMemEval-M, improves Overall by 7.83 pp over the strongest baseline, EvolveMem, and uses fewer tokens per question than EvolveMem.
Why it works
Integration keeps the peaks
Capability scores on BEAM · dashed: best ever reached · solid: final program
PrisMem reaches further than holistic evolution on every capability, and its final program closely matches its peaks.
Every design choice matters
Overall score lost when a component is removed (pp) · BEAM-1M · full PrisMem: 63.39
Capability-driven evolution−10.15
Random case selection−5.26
Trace-guided integration−3.16
Round-robin target selection−2.55
Removing capability-driven evolution yields the largest drop by blurring potential optimization directions when focusing solely on overall performance. Meanwhile, substituting with random case selection, omitting trace-guided integration, and simplifying to round-robin target selection degrade results by removing informative failure feedback, program fusion, and priority targeting.
05 / CITE
BibTeX
@misc{chen2026prismem,
title = {Capability-Driven Self-Evolution of Agent Memory},
author = {Chen, Yaoqi and Feng, Yuru and Zhang, Qianxi and Lu, Baotong and Lu, Jianan and Wang, Zhirui and Xu, Shusen and Jin, Zewen and Li, Zengzhong and Li, Cheng and Chen, Qi},
year = {2026},
eprint = {2610.06361},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2610.06361}
}