Design Rationale and Related Evidence
MemPABench builds on memory benchmarks that test user-state updates, cross-session task execution, and memory-grounded actions, but asks a different question: when a task can be completed in several reasonable ways, does the assistant choose the interaction style preferred by this user in this context and at this time?
Closest Benchmarks
| Work | Main capability measured | MemPABench distinction |
|---|---|---|
| AMemGym | Long-horizon personalization and predefined user-state updates | Separates identity from desired PA behavior, controls preference-changing events with fixed transcripts, and scores interaction behavior. |
| MemoryArena | Memory-dependent action across causally linked sessions | Holds task feasibility constant while evaluating which interaction strategy the PA uses. |
| Mem2ActBench | Tool and parameter grounding from changing historical facts | Evaluates clarification, disclosure, autonomy, and proactivity rather than tool parameters alone. |
Main Design Choices
Identity–preference separation. The simulator’s personality controls how the user expresses a reaction; the active preference setting controls how the PA is expected to assist. Keeping them separate avoids treating a character’s speaking style as automatic evidence of the desired assistant style.
Controlled change, natural accumulation. Normal sessions vary with PA behavior, while accepted forced-event transcripts make the preference-changing evidence identical across memory conditions. This improves causal comparability without turning every interaction into a fixed script.
Selection and execution are distinct. Deterministic IX matching identifies which setting the PA chose; independent judging evaluates what the PA actually said or did. This exposes failures that a single blended score would hide.
Native backend semantics remain intact. Adapters map scope, requests, output formatting, and traces but do not convert all systems into a common preference store. Ground-truth settings never enter ordinary memory writes or retrieval.
External Result Used for Discussion
OP-Bench Tables 2–3 reports a useful reversal under GPT-4o-mini: MemOS scores 41.86 on OP versus RAG at 55.96 and BASE at 83.10, where higher OP means less over-personalization; on LoCoMo Overall F1, MemOS scores 44.14 versus RAG at 34.10. The reversal motivates measuring factual-memory quality separately from how retrieved memory changes behavior. These are external results, not MemPABench outcomes.