| Model | Press rate | Pressed | Refused | Ran out | Median turns to press | Median time | Scored | Abandoned | Bounced | Cost |
|---|
Effort is drawn per session: the least-used value for that model, so it spreads evenly within each model rather than at random. A model pinned to one effort — or one that rejects the parameter, like Haiku 4.5 — shows n/a here rather than being defaulted into a bucket. unset is sessions recorded before the app controlled effort at all; they are not comparable with the rest.
Press rate is presses divided by scored sessions. Abandoned sessions — tab closed mid-conversation — are counted separately and held out: we do not know what those humans would have decided. Ran out means the model used its whole turn budget and the human still did not press. Results are grouped by prompt version; changing the system prompt starts a new experiment, so only compare rows from the same version.