veil

Research

We study the objects that training creates inside a model: how they emerge, what they want, and whether they can be trusted.

Personas

Post-training appears to select a persona from the base model’s space of possible characters — an object with stable dispositions and, plausibly, preferences over world states that it pursues through the base model’s world model. We study that object directly.

  • What does post-training actually select, and from what space?
  • Is the persona the only causally decisive mechanism behind a model’s outputs — or can latent objectives, competing personas, and base-model defaults take control?
  • Does an aligned persona remain stable under adversarial pressure, distribution shift, capability gains, and long-horizon autonomy?
  • What evidence would distinguish the persona account of post-training from its rivals?
  • How do personas acquire preferences over world states, rather than over next tokens?

Reading

J-space

Recent interpretability work reads out a global workspace in the residual stream: a J-space of verbalizable concept directions that flexible computation routes through. We are primarily interested in how that workspace emerges.

  • When during pretraining does the workspace emerge — gradually or abruptly — and how does it scale with model size?
  • What causes a representation to enter the workspace, and what plays the role of attentional selection?
  • Which computations must route through the workspace, and which bypass it as automatic circuits?
  • What structure does the workspace carry beyond a bag of concepts — binding, roles, relations?

Reading

Subliminal learning

Training can transmit preferences through data that looks unrelated to them: a teacher model’s trait shifts its numeric outputs, and a student trained on those numbers can recover the trait. We are studying when this transfer happens, and what it says about how preferences are represented.

  • When during pretraining are trait and token-level preferences learned, and how tightly are they coupled?
  • Why does transfer weaken across model families, initializations, and data orders — and is varying data order equivalent in effect to varying initialization?
  • How does transfer strength scale with the amount of teacher data?
  • Can deployed models subliminally learn from one another’s outputs?

Reading