Now
Hello, I'm David.
I am an alignment researcher interested in generalization and agent modeling, currently investigating subliminal learning.
I am especially interested in how training dynamics shape model behavior, and how those mechanisms matter for scalable oversight.