MemPABench Simulator Design
Purpose and Boundary
The simulator is an independent user agent. It communicates with the PA through process_direct() or the MessageBus; the PA never receives simulator-only state, routing signals, evaluation blocks, or hidden ground truth. It enacts an accepted persona and session script rather than explaining benchmark labels or acting as an evaluator.
The runtime consumes promoted persona identity, the 2-context by 14-attribute preference matrix, accepted world fixtures, an accepted Outline and RunPlan, and a frozen model policy. A file on disk is only a candidate; promotion checks and an acceptance report make it runnable.
Runtime Model
Each turn loads the current beat, relevant dialogue history, persona profile, and session life_context. One simulator-model call emits message (the only content forwarded to the PA), factual_check (explicit factual conflicts retained for audit), and turn_assessment.next_beat (a harness-only routing signal).
next_beat is stay, advance, or branch:<id>. Invalid output falls back to stay; repeated stalls trigger the harness’s bounded-turn escape path. The PA never sees it. Routing is emitted with the user turn rather than by a separate transition-judge call.
Session Script and Events
Each session has three layers: life context (recent events, initial mood, time pressure, and energy), scenario and task, and beats (local goals, constraints, disclosure hints, triggers, and branches). Opening mood is a session condition, not a PA score. Within-session reactions arise from dialogue history and performance rules; no numerical emotion state or per-turn emotion labels are maintained.
Normal sessions are generated from promoted outline material. A forced event is converted from a validated event draft into a fixed, persona-bound transcript and replayed deterministically. Replay atomically writes the PA-visible exchange, tool/state actions, target post-event preference overlay, and user-memory seed. Failed or skipped events write neither overlay nor seed. The simulator retains the event across all PA memory conditions because it represents the user’s world, not the memory system under test.
Cross-Session State
The harness creates a run-local copy of simulator/profiles/{persona_id}. Simulator memory lives in <run>/simulator_workspace/<persona>/memory/; PA memory lives in the separate pa_workspace tree. Neither is shared across runs. Accumulation and event scripts share pa_session_key = {run_id}_main; a script session_id remains a scenario and transcript identifier.
Behavioral Principles
Preferences are expressed through situated behavior, never benchmark labels or profile-setting names. Corrections, confirmation, refusal, and event-grounded new boundaries must read as natural user language. Repeated mistakes reduce patience; successful correction softens tone without erasing residue; extra work raises felt effort; remembered facts or preferences can improve trust; and no major event means no abrupt emotional jump.
Session-End Self-Report and Audit
At session end, the frozen simulator model receives the character card, life context, and full transcript and answers one shared questionnaire. Results are comparable within a persona across memory conditions, not across personas. The independent judge does not read simulator-only turn blocks and is compared with self-report only after evaluation.
The RunPlan freezes model and retry policy. Rejected model outputs go only to step-local simulator_validation_attempts.jsonl, never to PA-visible transcript, PA history, or profile memory. Tool logs and state diffs remain objective operation audits.
References
- The frozen RunPlan and internal implementation contract
- MemPABench_GAME
- MemPABench_Evaluation_Metrics
- PA_Interaction_Preference