Julian Minder Profile picture
PhD at EPFL with Robert West and Ryan Cotterell, MATS 7 Scholar with Neel Nanda
May 20 10 tweets 3 min read
New blog!
Synthetic Persona Pretraining (SPP): Alignment from Token Zero

Current alignment is shallow - values bolted on after pretraining can be routed around. To solve this, we wrote the desired persona directly into pretraining data. Early results, but we're very excited. 🧵 Image The Persona Selection Model posits that post-training picks from personas that pretraining already fixed, it doesn't build new ones. So if your pretraining corpus is a mess, no amount of post-training will save you.

(2/10)