MemPABench Index

MemPABench evaluates whether memory-enabled personal assistants learn and apply a user’s interaction preferences across contexts and over time. This site contains concise English descriptions of the current benchmark design. Internal plans, changelogs, archived alternatives, and HITL work queues are intentionally excluded.

Benchmark

Preference Model

Runtime

Personas and Sessions

Protocols