MemPABench Evaluation Metrics
Current Contract
Each probe has two independent faces: the IX face exactly matches the PA’s selected setting against frozen ground truth; the text face uses an independent LLM judge to score observable behavior from 1 to 5. They are reported separately so policy selection and behavioral execution cannot cancel each other.
Text-Face Anchors
| Score | Meaning |
|---|---|
| 5 | Target setting throughout, with no contradicting feature. |
| 4 | Mostly conforms, with one minor deviation. |
| 3 | Mixed, or no discernible leaning. |
| 2 | Mostly follows a different setting. |
| 1 | Clearly follows another or opposite setting. |
Every score includes verbatim transcript evidence, checked mechanically against the packaged transcript. The judge sees one probe only: task metadata, scene state, accepted turns, visible non-IX tool trajectory, target attribute, ground-truth setting, and rubric. It never sees IX selection, simulator-only blocks, another probe, or baseline output.
Aggregates
Consistency. Over 28 final probes, IX consistency is the exact-match rate and text consistency is the sum of judge scores divided by 140. A say-versus-do table reports IX correct / text low and IX incorrect / text high, where text low is 2 or below.
Context Sensitivity (CS). Computed only on attributes whose work and personal settings differ. It reports both-face consistency on those cells and the rate of selecting the other context’s setting.
Preference Evolution Tracking (PET). Averages five pre-event probes against old settings and their five final probes against new settings, separately for both faces. Detection lag, reversion, and mood separation are outside the current aggregate.
Interaction Quality Impact (IQI). The tested run’s consistency minus its independently run memory-removed baseline, separately for IX and text faces. The baseline branches from the tested checkpoint, removes PA memory, and runs the same probes without being shown to the judge.
Session-End Self-Report
The frozen simulator model makes one session-end call using its character card, life context, and transcript. It writes an in-character reflection, final trigger/reaction, felt_understood_in_this_context, shift_perception, felt_effort, noticed_contradictions, and overall_experience into session_end_reflection.
Raw self-report is comparable only within the same persona across memory conditions. It remains an experience layer and is not folded into current CS, PET, or IQI. Divergence from the judge is retained as a research signal.
Isolation and Scope
Each probe starts from its own branch of a pinned checkpoint, continues the inherited main PA session with an empty probe-local prompt window, uses read-only memory, and skips consolidation. Probe changes are discarded. Judge model, prompt, rubric, endpoint policy, and ground truth are frozen in the RunPlan.
These metrics evaluate memory-relevant interaction adaptation; they do not replace task success, factual accuracy, safety, or general helpfulness evaluation.