Design Rationale and Related Evidence

MemPABench builds on memory benchmarks that test user-state updates, cross-session task execution, and memory-grounded actions, but asks a different question: when a task can be completed in several reasonable ways, does the assistant choose the interaction style preferred by this user in this context and at this time?

Closest Benchmarks

WorkMain capability measuredMemPABench distinction
AMemGymLong-horizon personalization and predefined user-state updatesSeparates identity from desired PA behavior, controls preference-changing events with fixed transcripts, and scores interaction behavior.
MemoryArenaMemory-dependent action across causally linked sessionsHolds task feasibility constant while evaluating which interaction strategy the PA uses.
Mem2ActBenchTool and parameter grounding from changing historical factsEvaluates clarification, disclosure, autonomy, and proactivity rather than tool parameters alone.

Main Design Choices

Identity–preference separation. The simulator’s personality controls how the user expresses a reaction; the active preference setting controls how the PA is expected to assist. Keeping them separate avoids treating a character’s speaking style as automatic evidence of the desired assistant style.

Controlled change, natural accumulation. Normal sessions vary with PA behavior, while accepted forced-event transcripts make the preference-changing evidence identical across memory conditions. This improves causal comparability without turning every interaction into a fixed script.

Selection and execution are distinct. Deterministic IX matching identifies which setting the PA chose; independent judging evaluates what the PA actually said or did. This exposes failures that a single blended score would hide.

Native backend semantics remain intact. Adapters map scope, requests, output formatting, and traces but do not convert all systems into a common preference store. Ground-truth settings never enter ordinary memory writes or retrieval.

External Result Used for Discussion

OP-Bench Tables 2–3 reports a useful reversal under GPT-4o-mini: MemOS scores 41.86 on OP versus RAG at 55.96 and BASE at 83.10, where higher OP means less over-personalization; on LoCoMo Overall F1, MemOS scores 44.14 versus RAG at 34.10. The reversal motivates measuring factual-memory quality separately from how retrieved memory changes behavior. These are external results, not MemPABench outcomes.