MemPABench Index
MemPABench evaluates whether memory-enabled personal assistants learn and apply a user’s interaction preferences across contexts and over time. This site contains concise English descriptions of the current benchmark design. Internal plans, changelogs, archived alternatives, and HITL work queues are intentionally excluded.
Benchmark
- MemPABench_GAME — overall project description
- MemPABench_Evaluation_Metrics — evaluation metrics
- Design_Rationale_and_Evidence — design rationale and related evidence
- Paper_Overview_Figure — visual overview specification
Preference Model
- PA_Interaction_Preference — 14-attribute interaction-preference taxonomy
- Pattern_to_Preference_Rubrics — transcript-evidence interpretation
- Transcript_to_Preference_Workflow — extraction, HITL, and promotion
- Personality_Model_Rubrics — personality priors for simulation
Runtime
- MemPABench_Simulator_Design — simulator runtime and boundaries
- MemPABench_Simulator_Judge — independent post-hoc judge
- MemPABench_Judge_Endpoint_Pool — judge endpoint policy
- Nanobot_memory_adapter_design — memory adapter contract
Personas and Sessions
- User_A_World_Design · User_A_Timeline
- User_B_World_Design · User_B_Timeline
- User_C_World_Design · User_C_Timeline
- User_D_World_Design · User_D_Timeline