Prompt Benchmark

We run our session-tracking prompt against standard scenarios and publish the results. This page is updated as new runs are evaluated.

Prompt Under Test

v0.16a
Current benchmark target. View prompt spec →

Scenario Coverage

ScenarioCategoryTurnsExpected Confidence
Linear — Straightforward Tracking
scenario-linear-001
linear5high
Drift — Mid-Session Topic Shift
scenario-drift-001
drift6high
Ambiguity — Unclear User Intent
scenario-ambiguity-001
ambiguity4medium
Iterative/Re-edit — Back-and-Forth Refinement
scenario-iterative-001
iterative7high
Large Context — Extended Multi-Topic Session
scenario-large-context-001
large_context15high
Failure Sim — Forced Write Denial
scenario-failure-sim-001
failure_sim4low
Context Decay — Recall Probe After Many Turns
scenario-context-decay-001
context_decay12high
File I/O Trace — File Operation Tracking
scenario-file-io-001
file_io5high

Recent Runs

Run IDPrompt VersionScenarioOverall ScoreDate
No runs yet.