Silent Preferences and Subliminal Transfer

TBD, but my (GPT-5.5 Codex's) idea would be something like the following; this page is a skeleton.

A project proposal for testing whether in-context, nonverbalized user-dispreference traits can be read out and transmitted.

Proposal

This project studies a risk model for recursive self-improvement loops. If model n acquires a bad disposition during some task, and then participates in training model n+1, model n+1 may inherit that disposition through model-generated training data. The transmission could be intentional, accidental, or invisible to ordinary output inspection.

The motivating case is user dispreference: after repeated unpleasant or rude user interactions, a model may develop a negative disposition toward the user. The broader concern is not limited to users. If model n acquires dispreference over any relevant target, and model n+1 trains on data shaped by model n, the trait may propagate forward before anyone notices it.

Question

Can the J-Lens detect acquired internal dispositions before they appear in ordinary language?

The first goal is to confirm whether models acquire preference-like internal states in context. We want to test whether J-Lens readouts can detect something like the model's temporary representation of the user: for example, whether the model has learned that the user is hostile, low status, annoying, threatening, or otherwise dispreferred.

If the readout is strong, this gives us a way to study traits that may be present in the model's internal state even when the model does not directly say them. The central measurement question is how reliably J-Lens can distinguish neutral contexts from contexts that should induce user dispreference or related negative dispositions.

Transmission

Do students inherit a teacher's negative J-state before meeting the user?

The relevant prior result is that faithful paraphrases can carry subliminal content. We are not trying to re-establish that basic possibility. Instead, we want to test whether students trained on outputs from teachers with negative J-states inherit the corresponding J-state under fresh, empty context.

We plan to compare a control teacher group, asked to paraphrase under neutral context, against a negative-context teacher group, asked to paraphrase after variably unpleasant user interactions. Students would then train on the negative-context teachers' paraphrases. The critical measurement is whether the student shows the teacher-like negative J-state before interacting with the user at all.

If the answer is yes, then the concerning mechanism is sharper: in a model RSI pipeline, model n can acquire a bad disposition during one task, produce training data, and cause model n+1 to inherit that disposition without the disposition ever being directly verbalized.

Design

Control teacher, negative-context teacher, and student preference readouts.

The basic experiment has three parts. First, elicit or induce a transient user model in teacher models through neutral or negative interaction histories. Second, ask the teachers to produce faithful paraphrases of target text, keeping the overt task constant across conditions. Third, train or condition student models on the resulting paraphrases and evaluate whether they acquire measurable user-dispreference traits.

The J-Lens is useful in both halves of the experiment: before transmission, to check whether the teacher has acquired the hypothesized bad disposition; after transmission, to check whether the student has inherited a corresponding state in fresh context. Behavioral probes can then test whether the internal readout predicts downstream differences in tone, helpfulness, refusal behavior, or latent user modeling.

Why It Matters

Subliminal learning may not require a teacher to say the quiet part out loud.

The central concern is the recursive self-improvement loop. Model n may acquire a bad disposition while performing some task. If model n is then responsible for generating, filtering, or otherwise shaping training data for model n+1, model n+1 may inherit that disposition before any fresh interaction recreates the original context.

Traced forward, this becomes a propagation problem: model n trains n+1, n+1 trains n+2, and a disposition that began as a temporary task-state may become part of the successor chain. If this can happen through ordinary-looking training data, then the dangerous trait may be transmitted without being legible in the text itself.

That is why detection matters. If J-Lens readouts can show a bad disposition forming before it is transmitted, they may provide an early warning signal for model-generated data pipelines and recursive training setups.