incoherent scratchwork
Universe-scale motivation notes.
June 2026
This is a lot of my scratch work, but in general I want the universe to be the best it can possibly be.
Motivation
To motivate work towards future universe states, I estimate a lower bound for the amount of beauty to have existed to be the sum of the welfare experienced by the 110 billion humans1 which have thus far existed, and I estimate the upper bound on the number of computations necessary to simulate2 such experienced beauty to be 110B * <num human neural computations per life>, and find the number to be roughly on the order of 110B * 10^n (n in (~23, ~26)). This, in universal terms, is vanishingly small with respect to the possible future volumes of computronium and time3.
To illustrate just how small this number is, a Fermi'd estimate of how many human lives a <unit of volume> of computronium may simulate exceeds 110 billion every <unit of time>, and these could run for <time until computation is no longer possible (only in computronium form! More computation in less efficient structures may be possible)>. We should therefore conclude that the potential for universal beauty is absurdly massive, and thus place commensurate absurdly massive weight on bringing about these best universal states.45
Humanity's Position
So far as we are aware, humanity is unique in its position. We are the only known intelligent life in our universe, and are therefore the only beings capable of bringing about the rise of superintelligence. Humanity is then the light-bearer, and so we bear the weight of the risk of universal extinction. Any mistake we make in constructing our ASIs (likely via the scalable oversight / bootstrapping agenda) could propagate towards human extinction, or worse, unacceptably high s-risk by an improperly extrapolated alignment target.6
Project Framing
Roughly: I want the future universe to be the best it can possibly be, and have reason to believe the best possible universe to be an overwhelmingly dominating moral pressure. I recognize LLMs as the most imminent technology with potential to achieve the best possible universal outcomes. Ensuring that LLMs are robustly aligned for future scalable oversight bootstrapping projects is something I deem absolutely essential to guarantee we end up in the best possible universe basins, and subliminal learning as a phenomenon directly threatens this.
If we do not have a strong theory of subliminal learning, if our model n (from the model n supervises model n+1) has some misaligned trait t, then we risk t's propagation through to substantial light cone basins and, if uncorrected, the risk is approximately infinite. Also, I want humanity to survive and flourish, like in Machines of Loving Grace, and so investigating data order's effect on subliminal learning (for scalable oversight futures) is extremely valuable, and also valuable as a step into serious technical alignment research for me.
Footnotes
- And possibly non-negligible sum of animal experience, but this is more difficult to calculate and I think that beauty as an emergent property of intelligence seems a legitimate hypothesis, further complicating calculations. Besides, lower bound is true for the statement when subtracting all animal welfare.
- It's not quite clear that computational / pure computronium / generally substrate-modifying simulations could even yield phenomenologically comparable lives to the biological ones, but I find it likely that welfare simulations would result in some meaningful welfare, at least in some cases.
- Some people might call these "the possible states of the light cone."
- As a general corollary to these big numbers, inefficient computronium / utilitronium solutions, incorrect alignment targets, specification gaming by superintelligences, etc. may result in sum losses in universal utility n times greater than <number of humans> tortured in simulated hell for <unit of time>.
- In general, some amount of beauty has definitely existed and if we count it up among humans, it seems natural that it is in the computations performed by the 110 billion people who have existed. Then if we take for axiom that more cloned welfare = more good, we predict that a unit volume of computronium simulating these lifetimes over and over is a welfare-generating machine which massively outperforms the raw humans at generating welfare per unit spacetime.
- It seems plausible that an n supervising n+1 might have a slightly off generalized representation of what the platonic human (or, better, universal) CEV would be, and so n+1 generalizes this "off-color CEV." If this generalization process does not self-correct, then we have the above problem. Further, it is not necessarily the case that humans are the best basin for superintelligence to sprout from in general.