PhD at EPFL with Robert West and Ryan Cotterell, MATS 7 Scholar with Neel Nanda
May 20 • 10 tweets • 3 min read
New blog!
Synthetic Persona Pretraining (SPP): Alignment from Token Zero
Current alignment is shallow - values bolted on after pretraining can be routed around. To solve this, we wrote the desired persona directly into pretraining data. Early results, but we're very excited. 🧵
The Persona Selection Model posits that post-training picks from personas that pretraining already fixed, it doesn't build new ones. So if your pretraining corpus is a mess, no amount of post-training will save you.