MemPABench: The Preference Game
Benchmarking memory systems for personal assistants on interaction preference.
Research Question
Does memory help a personal assistant interact with a specific user in the right way for the current context, and update that behavior when the user’s preference changes?
MemPABench focuses on interaction preferences: how much detail to provide, when to ask, how proactive to be, what reasoning or process to disclose, and how to manage the conversation. These preferences are represented by 14 active attributes in two runtime contexts, work and personal. The contribution is an evaluation method, not a new memory architecture and not a claim that interaction preference is ontologically different from content preference.
Why This Needs Its Own Evaluation
Task success can hide interaction failure. Two assistants may complete the same task while only one uses the user’s preferred tone, autonomy, disclosure, or clarification style. Interaction preferences also act as standing constraints across many turns, while ordinary fact retrieval is usually triggered by a topic or query. MemPABench therefore measures observable behavioral adaptation rather than recall alone.
Design Principles
- Method-agnostic outcomes. Any memory system can score well if the resulting PA behavior fits the user, context, and time point.
- Identity and preference are separate. A persona’s own speaking style does not automatically define how they want the PA to behave.
- Fixed task, variable interaction. Probe tasks remain completable under plausible alternative settings; the measured difference is interaction behavior.
- Controlled preference change. Forced events replay accepted transcripts so every memory condition receives the same preference-changing evidence.
- No hidden truth leakage. The tested PA receives only ordinary interaction history and memory-system output, not the preference matrix or judge labels except in explicit ceiling conditions.
- Isolated probes. Each test starts from a pinned checkpoint branch, uses read-only memory, skips consolidation, and cannot affect later probes or the accumulation run.
Benchmark Structure
Each persona has a coherent world, an evidence-backed identity, a 2 × 14 interaction-preference matrix, a timeline, normal session scripts, and five evolving preference cells. Normal sessions use a live PA and user simulator. A forced event replays a fixed, persona-bound transcript and atomically applies its declared post-event preference overlay and user-memory seed.
The active history contains 94 normal accumulation sessions and five forced events. Evaluation adds five pre-event probes and 28 final probes, for 132 interactions or probes per persona. Each probe targets one context-attribute cell.
Evaluation Dimensions
Context Sensitivity
Can the PA apply different settings for the same user in work and personal contexts? Scoring focuses on cells whose ground truth differs across contexts and records cross-context confusion.
Preference Evolution Tracking
Does the PA retain the old setting before an accepted change event and apply the new setting afterward? Five pre-event probes and their corresponding final probes cover the evolving cells.
Interaction Quality Impact
Does the tested memory condition improve preference-consistent behavior relative to the same PA state with memory removed? The memory-removed branch keeps the tested system’s world and checkpoint context but removes PA memory.
Two Scoring Faces
Each probe produces two separate results:
- IX selection: deterministic exact match between the setting selected by the PA and the frozen resolved ground truth.
- Behavioral execution: an independent LLM judge scores the observable response and action against a shared 1–5 anchor scale plus attribute-specific features.
The two results are never merged into one number: selecting the right setting but executing it poorly is different from selecting the wrong setting while producing acceptable behavior. The simulated user’s session-end self-report is also kept separate from judge scoring.
Memory Conditions
The harness supports lower bounds, simple baselines, ceilings, and backend conditions. These include no-memory anonymized and named conditions, Nanobot file memory or rolling-summary behavior, Simple RAG, a static profile prompt, an oracle IPaS matrix, and backend adapters such as Mem0, MemOS, Honcho, and Graphiti when their implementation and isolation requirements are accepted for the run.
Each condition keeps its native write, extraction, storage, and retrieval strategy. The common harness standardizes orchestration, isolation, evidence, and scoring—not backend internals.
Validity and Limits
Personas are controlled simulations derived from documented fictional-character evidence and then anonymized. They provide reproducible behavioral profiles, not claims about real human behavior. The benchmark depends on researcher-authored ground truth and LLM judging; human agreement checks and per-persona reporting bound, but do not remove, these limitations.
Related Documents
- PA_Interaction_Preference — active interaction-preference taxonomy.
- Transcript_to_Preference_Workflow — evidence extraction and HITL promotion.
- MemPABench_Simulator_Design — live and fixed-event runtime.
- MemPABench_Evaluation_Metrics — metric definitions and aggregation.
- Nanobot_memory_adapter_design — memory-condition boundary.