From Transcript to Interaction Preference

Goal

Convert dialogue evidence into an accepted 2-context by 14-attribute interaction-preference matrix while preserving provenance and keeping identity extraction separate from desired PA behavior.

Two-Pass Extraction

Pass 1 — filter and classify. google/gemini-2.5-pro identifies relevant scenes, assigns context and candidate attributes, records evidence, and rejects non-diagnostic dialogue.

Pass 2 — synthesize settings. anthropic/claude-sonnet-4.5 groups evidence for each context-attribute cell, applies Pattern_to_Preference_Rubrics, selects a canonical setting, records confidence and counterevidence, and preserves source references.

The two contexts are exactly work and personal. Every final matrix contains all 28 cells and every cell names one setting from PA_Interaction_Preference.

Confidence and HITL

High-confidence evidence may write directly to a candidate matrix. Every weaker, disputed, seeded, or insufficiently proxied cell enters human review. Human review must resolve the cell to a named setting before the matrix can be accepted; null, unknown, and indifferent are not runtime settings.

Eleven attributes usually have usable transcript proxies. Process visibility, solution breadth, and capability boundary often require more human judgment because ordinary character dialogue rarely exposes the corresponding PA-specific behavior directly.

Identity and Personality Line

Identity extraction runs in parallel and produces concrete life facts, relationships, habits, traits, speech patterns, and evidence. Big Five and Jungian ratings provide compressed priors, with all ratings reviewed by a human. Identity describes how the simulated user behaves; it does not substitute for the preference matrix that describes how the PA should behave.

Anonymization and Promotion

The matrix and identity form one anonymization bundle. A candidate replacement map is reviewed before approved identity and preference artifacts are generated. Deterministic leak, consistency, collision, and archetype checks must pass before the bundle is promoted. See MemPABench_Anonymization_Protocol.

Reproducibility

Each run records model identities, source corpus and season range, prompts, actual provider usage, output digests, and review decisions. When one cell has too much evidence for a single Pass 2 call, evidence may be synthesized by season chunks, but the final call must preserve the same rubric and provenance requirements.