MemPABench Judge Endpoint Pool
The endpoint pool supplies the independent simulator judge. It is operational infrastructure, not ground truth: each call evaluates a recorded interaction against its rubric and remains independent of simulator-only outputs.
- Two configured, health-checked judge endpoints serve the pool.
- One session is pinned to one endpoint for its judging pass; a retry stays there unless the endpoint becomes unavailable.
- Routing weights are static during a run and frozen with the RunPlan.
- Failed endpoints and retries are recorded in run audit metadata.
Both machines use the same frozen judge prompt, schema, model identity, and rubric version. They are capacity replicas, not an ensemble. Cross-endpoint agreement is measured only in validation probes.