incoherent scratchwork
Risks, first experiment, and thesis.
June 2026
Question 1
What do you think are the biggest risks in the project, and what might be some of the reasons we fail at producing any relevant results? What can we do to mitigate it?
A primary concern is that subliminal learning must be both established and measurable in 160m parameter language models. This must be confirmed by attempting a replication of the fine-tuning subliminal learning results on a polypythia by measuring completions1 for animal preference-related changes (in both teacher and student polypythias, sequentially).
Another primary concern would be that fixed initialization would predict "preference features" in a way such that negative results from data order's effect on subliminal learning become uninformative rather than a negative result23. This is harder to mitigate as an effect, but we could still examine and compare the features learned by the teacher and student polypythias to see if similar representations were learned regardless of data order.
Question 2
What do you would be the first experiment you would run in this project?
As informed by the biggest risks of the project, I would first attempt to replicate the fine-tuning subliminal learning results in our polypythia to ensure project validity. This would entail constructing a training set to elicit animal preference in a polypythia, training a teacher polypythia to exhibit animal preference, finding some completions-model (animal) preference monitoring technique, training a student model on teacher generated completions, and monitoring the student model for preference acquisition.
Thesis
My thesis: that this project will show that representations of preference are acquired in part due to data order and initialization, and that data order's specific effect on preference acquisition isomorphism is that early learned representations of preference will be useful gradient minimizers for downstream preference formation tasks, and therefore enforce a particular morphism for preference. Prediction being that, if this morphism is enforced through early training, then subliminal learning may occur due to steering vector transitivity.
I predict the effect of data order would be that different representations of preference may be learned according to data order, since features may differentially affect loss at the time of preference data introduction in training, and thus may result (unless the features to represent preference are locked in early, and remain stable) in a negative result for subliminal learning among fixed initialization, variable data order training runs, traceable to the features which are recruited to represent the preference.
The concerns are thus that subliminal learning is difficult to establish at the 160m parameter language model scale, that if initialization is doubly predictive of learning basis and feature placement (therefore also representation geometry), then a negative result of data order does not necessarily inform us to anything other than initialization's known results of predicting subliminal learning.
Basically I think that thinking of subliminal learning as a steering vector means that data order's effect on SL may be partly mediated by which early learned representations are available for later preference acquisition.
Footnotes
- In general, finding a way to monitor completions-model preferences seems difficult, but I think approaches like basic prompting, logit probability monitor, or SAE feature activation-style methods would be able to measure preference acquisition.
- Specifically, that initialization would predict the representations of preferences because initialization and early learning would lock the recruitable features available for building a preference representation to a stable set, thus removing differential loss at training time attributable to different features, and resulting in the same learned representations despite different data order.
- Further, a simple prediction may be that if final model weight differences are simply too great (between control and variable data order models), then SL breaking here is already predicted by weight differences, losing any mechanism of data order.