PrisMemRead the paper
Microsoft
University of Science and Technology of China
University of California, San Diego

AGENT MEMORY / SELF-EVOLUTION

PrisMem
Capability-Driven Self-Evolution of Agent Memory

Yaoqi Chen1,2,*Yuru Feng2,3,*Qianxi Zhang2Baotong Lu2Jianan Lu2Zhirui Wang2Shusen Xu2Zewen Jin2Zengzhong Li2Cheng Li2Qi Chen2
1 University of Science and Technology of China2 Microsoft3 University of California, San Diego

* Equal contribution

+10.54pp

BEAM-1M

Over strongest baseline
+7.83pp

LongMemEval-M

Over strongest baseline
1M+

Tokens of history

Generalization to longer contexts

01 / THE INSIGHT

80.5% of revisions with no overall gain still improve at least one capability.

Existing self-evolving memory methods revise a memory program using one overall score. But memory is multidimensional, and that score blurs what each revision actually achieved.

Four panels showing hidden capability gains, uneven capability changes, and conceptual exploration boundaries
Figure 1. Why capability-driven evolution matters. Analysis of M⋆ and EvolveMem evolution traces on BEAM.

We derive five representative capabilities from a broad review of agent memory benchmarks and memory systems. Then we annotate every question with the one it primarily requires, plus up to two secondary ones. These labels only organize feedback during evolution and are never given to the memory program at test time.

F

Factual retrieval

Preserve and retrieve explicit facts, values, and other localized details.

T

Temporal tracking

Order events, determine recency, track states and knowledge updates.

P

Preference extraction

Recover and apply user preferences, instructions and constraints.

M

Multi-session synthesis

Aggregate and synthesize evidence across multiple sessions or events.

A

Adversarial

Abstain when memory lacks supporting evidence, or cautiously resolve contradictory or misleading evidence.

02 / THE METHOD

Explore separately.
Improve together.

PrisMem evolves an executable memory program in three stages: build a base, grow capability specialists, then integrate them.

PrisMem method overview with cold start, dependency-aware selection, history-guided diagnosis, specialist refinement and trace-guided integration
Figure 2. Overview of PrisMem: cold start, capability refinement, and trace-guided integration.

ZOOM IN · STAGE II

Every round makes two choices before any code changes.

Which capability to refine, and which of its failures to learn from.

CHOICE 1 · DEPENDENCY-AWARE SELECTION

Which capability to refine?

urgency = own error + strongest pressure on another capability1 + times already picked

F Factual
picked 1× → ÷2
T Temporal
not picked yet
P Preference
picked 1× → ÷2
M Synthesis
picked 2× → ÷3
A Adversarial
not picked yet

A capability is chosen because it is weak itself, or because it holds another capability back. Each past pick lowers its priority, so plateaued capabilities step aside: here T wins although M has the larger raw urgency (dashed).

CHOICE 2 · HISTORY-GUIDED DIAGNOSIS

Which failures to learn from?

priority = how well earlier programs answered ittimes it, or a similar case, was already picked

ALL TRAINING CASES OF THE TARGET CAPABILITY✕ already answered well (score > 0.8)REMAINING FAILURES, RANKED BY PRIORITY5 cases → diagnosis

A case that used to be solved and now fails shows exactly what a recent revision broke. Discounting repeats keeps the five cases diverse.

Schematic diagrams with illustrative values. They show the idea behind the two selection rules (Eqs. 1–7 in the paper).

ZOOM IN · STAGE III

Every integration rests on evidence and a plan.

Paired differential cases show where programs behave differently; the integration plan decides what to carry over.

EVIDENCE · PAIRED DIFFERENTIAL CASES

Where do two programs really differ?

ONE CASETWO PROGRAMSTWO OUTCOMES ? program A program B 1.0 0.0 paired caseΔ = 1.0
F FactualP131vsP170.5+ bring in
T TemporalP70.4vsP171◆ protect
P PreferenceP151vsP170+ bring in
M SynthesisP101vsP170+ bring in
A AdversarialP70.25vsP171◆ protect

Specialist wins → bring in. Base P17 wins → protect.

PLAN · LINEAGE-ORDERED INTEGRATION

What to take, and what to guard?

SPECIALISTS · CLOSEST LINEAGE FIRST BaseTA62.5FactualspecialistFsynthesis rulesPreferencespecialistPpreference ruleSynthesisspecialistMtopic metadataUnified65.8

While contrasting cases serve as diagnostic indicators targeting specific program components, the plan leverages these insights to localize code modules and formulate a step-by-step integration strategy.

03 / LIVE DEMO

Watch memory evolve.

One real 20-round run on BEAM, replayed step by step: which program was revised, what failed, and what changed.

04 / THE EXPERIMENTAL RESULTS

Built on shorter histories.
Tested at a million tokens.

Static memorySelf-evolving baselinesPrisMem

Table 2 visualized · Qwen3.8-27B · 3 runs · Bars show means; whiskers show ±1 standard deviation.

BEAM-1M

Capability scores (%) ↑

Factual retrieval

BEAM-1M: Factual retrieval050100Mem0: 54.74 ± 3.06 %Mem0A-MEM: 56.21 ± 1.83 %A-MEMHippoRAG2: 65.24 ± 1.26 %HippoRAG2SimpleMem: 48.55 ± 3.33 %SimpleMemLightMem: 46.66 ± 0.82 %LightMemM⋆: 61.29 ± 5.14 %M⋆EvolveMem: 41.80 ± 4.77 %EvolveMemPrisMem: 74.17 ± 4.87 %PrisMem
PrisMem: 74.17 ± 4.87 %

Temporal tracking

BEAM-1M: Temporal tracking050100Mem0: 46.51 ± 0.79 %Mem0A-MEM: 46.22 ± 0.62 %A-MEMHippoRAG2: 47.06 ± 1.85 %HippoRAG2SimpleMem: 36.82 ± 0.47 %SimpleMemLightMem: 33.40 ± 0.65 %LightMemM⋆: 48.98 ± 5.87 %M⋆EvolveMem: 39.74 ± 3.15 %EvolveMemPrisMem: 57.26 ± 2.15 %PrisMem
PrisMem: 57.26 ± 2.15 %

Preference extraction

BEAM-1M: Preference extraction050100Mem0: 60.48 ± 2.39 %Mem0A-MEM: 65.39 ± 0.52 %A-MEMHippoRAG2: 66.43 ± 1.57 %HippoRAG2SimpleMem: 48.85 ± 1.19 %SimpleMemLightMem: 36.64 ± 0.63 %LightMemM⋆: 71.91 ± 3.40 %M⋆EvolveMem: 52.21 ± 1.53 %EvolveMemPrisMem: 75.58 ± 3.91 %PrisMem
PrisMem: 75.58 ± 3.91 %

Multi-session synthesis

BEAM-1M: Multi-session synthesis050100Mem0: 38.02 ± 0.63 %Mem0A-MEM: 53.87 ± 0.86 %A-MEMHippoRAG2: 44.83 ± 0.60 %HippoRAG2SimpleMem: 21.87 ± 2.00 %SimpleMemLightMem: 16.70 ± 0.89 %LightMemM⋆: 42.19 ± 4.88 %M⋆EvolveMem: 21.67 ± 3.07 %EvolveMemPrisMem: 54.12 ± 3.37 %PrisMem
PrisMem: 54.12 ± 3.37 %

Adversarial

BEAM-1M: Adversarial050100Mem0: 49.11 ± 1.24 %Mem0A-MEM: 41.13 ± 0.51 %A-MEMHippoRAG2: 44.82 ± 2.28 %HippoRAG2SimpleMem: 50.68 ± 0.40 %SimpleMemLightMem: 55.45 ± 0.94 %LightMemM⋆: 46.02 ± 7.52 %M⋆EvolveMem: 45.67 ± 3.50 %EvolveMemPrisMem: 64.32 ± 6.51 %PrisMem
PrisMem: 64.32 ± 6.51 %

Overall

% · higher is better

BEAM-1M: Overall050100Mem0: 48.95 ± 0.61 %Mem048.95A-MEM: 51.56 ± 0.62 %A-MEM51.56HippoRAG2: 51.86 ± 0.87 %HippoRAG251.86SimpleMem: 40.18 ± 0.76 %SimpleMem40.18LightMem: 36.44 ± 0.38 %LightMem36.44M⋆: 52.85 ± 2.49 %M⋆52.85EvolveMem: 40.01 ± 1.46 %EvolveMem40.01PrisMem: 63.39 ± 1.02 %PrisMem63.39
PrisMem: 63.39 ± 1.02 %

Average token usage per question (K) ↓

lower is better

BEAM-1M: Average token usage0250500Mem0: 190.28 ± 0.20 K tokens / questionMem0190.28A-MEM: 428.95 ± 3.58 K tokens / questionA-MEM428.95HippoRAG2: 207.70 ± 0.19 K tokens / questionHippoRAG2207.70SimpleMem: 95.72 ± 0.24 K tokens / questionSimpleMem95.72LightMem: 86.20 ± 0.05 K tokens / questionLightMem86.20M⋆: 89.47 ± 5.11 K tokens / questionM⋆89.47EvolveMem: 98.40 ± 4.67 K tokens / questionEvolveMem98.40PrisMem: 100.59 ± 5.02 K tokens / questionPrisMem100.59
PrisMem: 100.59 ± 5.02 K tokens / question

PrisMem leads all five capability metrics on BEAM-1M and improves Overall by 11.53 pp over the strongest static baseline (10.54 pp over the strongest self-evolving baseline), while using fewer tokens per question than Mem0, A-MEM, and HippoRAG2.

Why it works

Integration keeps the peaks

Capability scores on BEAM · dashed: best ever reached · solid: final program

Figure 3(b): radar comparison of M star, EvolveMem and PrisMem across factual, temporal, preference, multi-session and adversarial capabilities. Dashed lines show historical capability-wise peaks and solid lines the final selected program.

PrisMem reaches further than holistic evolution on every capability, and its final program closely matches its peaks.

Every design choice matters

Overall score lost when a component is removed (pp) · BEAM-1M · full PrisMem: 63.39

Capability-driven evolution−10.15
Random case selection−5.26
Trace-guided integration−3.16
Round-robin target selection−2.55

Removing capability-driven evolution yields the largest drop by blurring potential optimization directions when focusing solely on overall performance. Meanwhile, substituting with random case selection, omitting trace-guided integration, and simplifying to round-robin target selection degrade results by removing informative failure feedback, program fusion, and priority targeting.

05 / CITE

BibTeX

@misc{chen2026prismem,
  title = {Capability-Driven Self-Evolution of Agent Memory},
  author = {Chen, Yaoqi and Feng, Yuru and Zhang, Qianxi and Lu, Baotong and Lu, Jianan and Wang, Zhirui and Xu, Shusen and Jin, Zewen and Li, Zengzhong and Li, Cheng and Chen, Qi},
  year = {2026},
  eprint = {2610.06361},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2610.06361}
}