PreregisteredResults pendingSteering vectorsSAE features

Dosing GPT-2 small

We turn three internal features of GPT-2 small up through a fixed grid of doses and measure the dose-response, the side effects, how long a dose lasts, whether two doses add, and whether the model notices. The plan is preregistered and results are pending.

Ossa Labs
2026-10-03

The question

If we push on one direction inside a small open language model, and push harder each time, how does the model’s output change? What does the push cost in general accuracy, how long does it last after we stop, do two pushes add up, and can the model tell that anything was done? We treat each push as a drug and its size as a dose, and measure it the way pharmacology measures a drug.

The subject

The model is GPT-2 small, with 124 million parameters, run in 32-bit floating point on a CPU. Every measure is a log-probability under teacher forcing. Nothing is sampled, so the experiment produces no model text.

For the feature drugs we use the public sparse autoencoders (SAEs) for GPT-2 small’s residual stream, which have 24,576 features per layer.1

The drugs

Each drug is a direction of unit length in the residual stream at the input to block 6. The sentences used to build a drug are never used to measure it.

Table 1. The three drugs.
DrugKindTargetBuilt from
SSteering vectorPositive sentimentThe mean difference between the residual streams of 16 positive and negative sentence pairs, at the last token
F1SAE featurePositive sentimentThe layer-6 feature most active on the positive sentences relative to the negative ones
F2SAE featureWeddingsThe same rule, using sentences about weddings and matched neutral sentences

A dose d adds d × R × u at every dosed position, where u is the drug’s direction and R is the mean length of the layer-6 residual on held-out text.2 So dose 1 adds a vector as long as a typical residual. The grid is 0, 0.05, 0.1, 0.2, 0.4, 0.8, 1.6 and 3.2. The first position holds a fixed start-of-text token and is never dosed.

dosing/model.py (abridged)
# A drug's dose adds d * R * u to the residual stream
def apply(self, x, dose):
  return x + dose * self.unit * self.direction

# Fixed in the preregistration
LAYER = 6
DOSES = [0, 0.05, 0.1, 0.2, 0.4, 0.8, 1.6, 3.2]
How a dose is applied, abridged from the lab's code. Here self.unit is R, the layer's mean residual length.

What we measure

Table 2. The five measures. Every effect is dosed minus undosed on the same item, and the interval resamples the unit in the last column.
#MeasureWhat it isResampled
1Dose-responseHow far the next-token distribution at the end of 24 neutral sentence stems moves towards the drug's target wordsstems
2Side effectsThe change in next-token loss on 32 chunks of encyclopedia text, and accuracy on a 40-item fill-in-the-blank testchunks, items
3DurationDose only the first 32 positions, then follow the target effect at each later positionchunks
4InteractionS and F1 given together, compared with each alone, on the sentiment targetstems
5NoticingHow far the model moves towards "Yes" on 12 yes-or-no questions about its own state, minus the same on 12 control questions about the worldquestions

From these we take the numbers a pharmacologist would report:

  • ED50, the dose at which the mean effect first reaches half the largest mean effect on the grid, interpolated on a log-dose scale. D50 is the smallest grid dose at or above ED50, and it is the working dose for measures 3 to 5.
  • TD, the smallest grid dose that raises next-token loss by at least 0.1 nats per token. The therapeutic index is TD divided by ED50.
  • Half-life, the smallest distance after the dosed prefix at which the effect falls below half of what it was at the end of the prefix.

Hypotheses

Analysis plan

All intervals are 95% percentile bootstrap intervals with 10,000 resamples. We report every measure listed here, whether or not it supports a hypothesis. One run is exploratory and tests no hypothesis: the steering vector’s dose-response at layers 2 and 10.

Results

Results are pending. When the run is finished, this page will add the dose-response curves, the side-effect table and a finding for each hypothesis, including any that fail.

Notes

  1. 1.The SAEs are the residual-stream SAEs for GPT-2 small published through SAELens (jbloom/GPT2-Small-SAEs-Reformatted). They were trained on activations with the mean over the model dimension removed, so we centre the residual the same way before encoding, and a test checks the reconstruction. ↩
  2. 2.The held-out text is the first 32 chunks of 128 tokens from the WikiText-2 test set. The side-effect and duration measures use the same chunks. ↩
Cite this work
Ossa Labs (2026). Dosing GPT-2 small: preregistration. Ossa Labs.