Research
We study the objects that training creates inside a model: how they emerge, what they want, and whether they can be trusted.
Personas
Post-training appears to select a persona from the base model’s space of possible characters — an object with stable dispositions and, plausibly, preferences over world states that it pursues through the base model’s world model. We study that object directly.
- What does post-training actually select, and from what space?
- Is the persona the only causally decisive mechanism behind a model’s outputs — or can latent objectives, competing personas, and base-model defaults take control?
- Does an aligned persona remain stable under adversarial pressure, distribution shift, capability gains, and long-horizon autonomy?
- What evidence would distinguish the persona account of post-training from its rivals?
- How do personas acquire preferences over world states, rather than over next tokens?
Reading
J-space
Recent interpretability work reads out a global workspace in the residual stream: a J-space of verbalizable concept directions that flexible computation routes through. We are primarily interested in how that workspace emerges.
- When during pretraining does the workspace emerge — gradually or abruptly — and how does it scale with model size?
- What causes a representation to enter the workspace, and what plays the role of attentional selection?
- Which computations must route through the workspace, and which bypass it as automatic circuits?
- What structure does the workspace carry beyond a bag of concepts — binding, roles, relations?
Reading
Subliminal learning
Training can transmit preferences through data that looks unrelated to them: a teacher model’s trait shifts its numeric outputs, and a student trained on those numbers can recover the trait. We are studying when this transfer happens, and what it says about how preferences are represented.
- When during pretraining are trait and token-level preferences learned, and how tightly are they coupled?
- Why does transfer weaken across model families, initializations, and data orders — and is varying data order equivalent in effect to varying initialization?
- How does transfer strength scale with the amount of teacher data?
- Can deployed models subliminally learn from one another’s outputs?
Reading