MemPABench Writing Caveats
Safe Framing
MemPABench evaluates how existing memory systems learn and apply interaction preferences. It does not propose a new memory architecture or claim to cover all preference learning.
Interaction and item preferences are operationally distinguishable for evaluation, not ontologically different kinds of information. Interaction preference is tractable as a stable, finite matrix of active attributes across work and personal contexts; item preference has an open-ended domain and usually needs much more behavioral evidence to establish ground truth.
Claims to Avoid
- Do not claim interaction preference is uniquely contextual, evolving, or implicit; item preferences can have all three properties.
- Do not call interaction preference the universal deployment bottleneck without supporting user evidence.
- Do not claim that retrieval is useless, that all existing memory systems share one limitation, or that a reflexive model component is required.
- Do not present the benchmark as a replacement for factual-memory or recommendation benchmarks.
Defensible Contribution
The central contribution is a reproducible evaluation setting for behavioral adaptation across context and time. Its ground truth is constructed from accepted persona and script evidence, while paired probes distinguish explicit preference recall from behavioral application. The benchmark compares memory-system behavior under a common harness without rewriting backend-native storage or retrieval mechanisms.
Evidence Standard
Keep literature claims specific: name the condition, metric direction, numerical result, and source. Separate published external evidence from MemPABench results. Treat speculative mechanisms as discussion questions unless empirical results support them.
References
- MemPABench_GAME for benchmark scope and protocol.
- PA_Interaction_Preference for the active taxonomy.
- MemPABench_Evaluation_Metrics for behavioral evaluation.