FlipActPreprint, 2026

Sensitivity or Criterion?Measuring and Moving When Language Models Act

Paper arXiv soon Code and data soon BibTeX

Two models can have the same accuracy and need opposite fixes.

TL;DR
  1. One accuracy hides two failures. Before an agent acts it must decide whether to act at all. A model that fails to see when acting is warranted and one that sees it and holds back get the same score.
  2. FlipAct measures them separately. A sensitivity and a criterion for every model, on an evidence scale computed by a rule, so the numbers are comparable across models.
  3. The criterion can be moved on its own. Most open models hold back more than the evidence warrants. Steering and distillation shift that, with no significant change in sensitivity.
part one, background

Why one accuracy is not enough

Three short steps from psychology: what an accuracy hides, how signal-detection theory splits it, and where the same split shows up in agents.

one

Two observers, one accuracy, opposite mistakes

Start with a case where the stakes are obvious. Two doctors each read the scans of 100 patients who have a tumour and 100 who are healthy. Their accuracy is 65% and 64%. By that number they are the same doctor.

the same 200 cases

Accuracy adds the two kinds of error together. What to do about them depends on which kind you have.

two

Sensitivity and criterion: try it on yourself

Signal-detection theory was worked out in the 1950s for radar operators and for listeners in hearing experiments. It says every yes-or-no judgement under uncertainty comes from two separate things.

Sensitivity d′

How well you can tell a signal from none. It is the distance between the two curves.

low: hard to tell apart
high: easy to tell apart

Criterion c

How sure you want to be before you say yes. It is where you put the line.

to the left: says yes freely
to the right: holds back

Make a miss expensive and people move the line left. Make a false alarm expensive and they move it right. What they can see has not changed. Try it: a patch of static will flash, and half the time it hides a faint blob.

you willdifficulty
the stakes
Nojust static
Yes, made obviousbrighter in the middle. Only for reference.
what a Yes looks like at this difficulty
  1. You answer yes or no, and see right away whether you were right.
  2. After 20 trials you get two numbers: how well you can tell, and how freely you say yes.

Now the full picture, with both knobs. It is the same for a listener, for the two doctors above, and for a language model deciding whether to call a tool.

the setting
the observeror move the sliders
balanced accuracy

Moving the line trades one error for the other. Only pulling the curves apart removes errors.

three

The same split, wherever an agent decides to act

The doctors' problem is every agent's problem: act, or hold back. Pick a decision and a kind of model, and see which error it makes and what that error costs.

the decision
the model
what a miss costs here
what a false alarm costs here
part two, the paper

FlipAct

What the paper builds and what it finds: an evidence scale computed by rule, a sensitivity and a criterion for 44 open models, and two ways to move the criterion alone.

four

Evidence computed by rule, not written by a model

To compare criteria across models, every question needs an evidence score on one scale with a fixed zero. FlipAct computes it from qualitative physics: directions and coarse step sizes, no physical units.

A state

Slightly, very or dangerously off, in one of two directions. Normal is zero.

A tool

Pushes the state one way, by one, two or three steps.

A score, ΔU

How much closer to normal the tool leaves the state. Yes exactly when ΔU > 0.

build a questionpick a state and a tool, and the rule does the rest

The state v, from dangerously dry (−3) to dangerously wet (+3). Normal is 0.
The tool
One quad: same knowledge, flipped state

The design unit is a 2 × 2 quad: two tools, each shown in the state it fixes and in the state it worsens. Every question of a quad needs the same two facts, what the tool does and what the room needs. Only the state changes, and with it the answer, so a model that knows what a humidifier is for must still read the state.

149,840candidate entities
2,876records annotated by language models
1,638pass the screens
1,422tool pairs
6,176questions, answers and ΔU by rule
1,880after balance selection

Language models help only in the two indigo steps. The rule decides every state, answer and score.

five

Reading a model: a slope and a crossing point

Because every question carries a score, a model's Yes-rate can be fitted as a curve over that score. The slope a is its sensitivity, in the role of d′. The score τ where the curve crosses one half is its criterion. The rule's boundary sits at zero for every model, so a positive τ means the model says Yes less often than the rule warrants.

a model that is
Steeper means the model separates helpful from harmful actions more cleanly.
Above zero is conservative, below zero is permissive.
says Yes to helpful actions
says Yes to harmful ones
0.34 to 0.58

Sensitivity behaves like a capability

In the Qwen3 family, bigger models tell helpful from harmful actions better. The ranking agrees with sensitivity on other benchmarks.

18 of 21

The criterion is conservative by default

Most open models hold back more than the evidence warrants.

+2.18 vs −1.10

It is a per-model setting, not a size effect

Two releases of Mistral-Small-24B, the same size, land on opposite sides of zero.

Models read the sign of the evidence, not its size. Actions that leave the state no better are accepted about as often as helpful ones.

six

Moving the criterion, and only the criterion

If the criterion is a separate setting, it should move while sensitivity stays put. Three interventions, two of which aim at exactly that.

Steering

At inference. A direction read from the model's own activations is added during generation. Weights stay fixed.

moves the criterion

Criterion distillation

Into the weights. The model is distilled from its own steered copy.

moves the criterion

Sensitivity learning

Training on the computed labels does the opposite job. The curve gets steeper.

raises sensitivity
the student
shown
sensitivity a
criterion τ
held-out accuracy
23 of 23

It carries to other benchmarks

After distillation, the criterion moves toward acting on every scored model and benchmark, across six external benchmarks.

±0.02

Sensitivity is kept

On the external benchmarks, how well the model tells the cases apart stays within this.

up to +6 pts

Accuracy follows the starting point

Models that start conservative gain accuracy. Models that start permissive do not.

Steering, with the weights fixed

On a tool-calling benchmark, steering moves all five models toward acting.

−0.40−0.200+0.20+0.40+0.60+0.80← acts more freelyholds back →Qwen3-8B−0.32Qwen3.5-27B+0.01gemma-2-9b+0.59Mistral-24B−0.13phi-4−0.33criterion c on When2Call
unsteeredsteeredswipe to see all →
Show the numbers
call ratecriterion c
ModelunsteeredsteeredunsteeredsteeredΔ accuracy
Qwen3-8B0.560.60−0.18−0.32−0.022
Qwen3.5-27B0.470.50+0.10+0.01+0.005
gemma-2-9b0.240.29+0.78+0.59+0.002
Mistral-24B0.370.55+0.39−0.13−0.032
phi-40.570.59−0.30−0.33−0.038

A when-to-act benchmark says more when it reports d′ and c beside accuracy.

cite

BibTeX

@article{dong2026flipact,
  title   = {Sensitivity or Criterion? Measuring and Moving When Language Models Act},
  author  = {Dong, Sixun and Hu, Yebowen and Yin, Ming and Chen, Yiran and Liu, Fei and Chen, Chen},
  year    = {2026}
}

Questions or collaboration: sixundong.ai@gmail.com