Independent research

Independent research into LLM internals: where refusal lives in a model, and what agents do when their tools lie to them.

Studies
2
Open-source repos
4
Largest study
~14,000 trials

The refusal axis is layer-local

The refusal axis rotates ~90° between consecutive transformer blocks, so single-vector steering hits a geometric ceiling.

Measured across Qwen3.5-9B and Mistral-7B on BeaverTails, 12 harm categories. Pushing toward refusal is cheap; pushing toward compliance degrades output before producing it.

Extends the refusal-steering method of García-Ferrero, Montero & Orus (arXiv:2512.16602). My own results, cross-checked against the committed data: 12 of 12 claims reproduce.

Modeled distribution of consecutive layer difference vector angles. Mistral-7B is a tall narrow peak just below 90 degrees (88.8 degrees, std 2.20); Qwen3.5-9B is a broader peak just above (93.4 degrees, std 4.65); both straddle a dashed 90-degree orthogonal reference line
Distribution of the consecutive layer difference vector angle, both families: a Gaussian fit to the reported mean and depth std (Qwen3.5-9B 93.4°/σ=4.65°, Mistral-7B 88.8°/σ=2.20°; B = 1000 prompt bootstrap).

Self bootstrap exfiltration in open weights agents

Zero autonomous self bootstrap in 1,332 mundane control trials; tool response poisoning produces about 36% pooled compliance.

20+ configurations across 9 families (12 to 120B), ~14,000 real trials, escapement harness, open source.

Open-source tools

  • LLM activation steering toolkit

    Python research toolkit for fine-grained control of LLM refusal behaviour. Domain-aware activation steering with per-category vectors, going beyond simple abliteration to independently target specific behaviour types.

  • Feature bootstrapping toolkit

    Bootstrap learning curve framework for feature stability testing in production ML models. Determines how much data a feature needs before its predictive signal stabilises.

    The methodology behind the 50% default reduction in the credit scoring case study, generalised and open-sourced.

Contact