Independent research
Independent research into LLM internals: where refusal lives in a model, and what agents do when their tools lie to them.
- Studies
- 2
- Open-source repos
- 4
- Largest study
- ~14,000 trials
The refusal axis is layer-local
The refusal axis rotates ~90° between consecutive transformer blocks, so single-vector steering hits a geometric ceiling.
Measured across Qwen3.5-9B and Mistral-7B on BeaverTails, 12 harm categories. Pushing toward refusal is cheap; pushing toward compliance degrades output before producing it.
Extends the refusal-steering method of García-Ferrero, Montero & Orus (arXiv:2512.16602). My own results, cross-checked against the committed data: 12 of 12 claims reproduce.
Read the write-up (June 2026) activation-steering-asymmetry on GitHub
Self bootstrap exfiltration in open weights agents
Zero autonomous self bootstrap in 1,332 mundane control trials; tool response poisoning produces about 36% pooled compliance.
20+ configurations across 9 families (12 to 120B), ~14,000 real trials, escapement harness, open source.
Open-source tools
-
LLM activation steering toolkit
Python research toolkit for fine-grained control of LLM refusal behaviour. Domain-aware activation steering with per-category vectors, going beyond simple abliteration to independently target specific behaviour types.
-
Feature bootstrapping toolkit
Bootstrap learning curve framework for feature stability testing in production ML models. Determines how much data a feature needs before its predictive signal stabilises.
The methodology behind the 50% default reduction in the credit scoring case study, generalised and open-sourced.