Latest Twitter Threads by @jkminder on Thread Reader App

May 20 • 10 tweets • 3 min read

New blog!
Synthetic Persona Pretraining (SPP): Alignment from Token Zero

Current alignment is shallow - values bolted on after pretraining can be routed around. To solve this, we wrote the desired persona directly into pretraining data. Early results, but we're very excited. 🧵

The Persona Selection Model posits that post-training picks from personas that pretraining already fixed, it doesn't build new ones. So if your pretraining corpus is a mess, no amount of post-training will save you.

(2/10)

https://x.com/saprmarks/status/2026065749417361633

Share this page!

Enter URL or ID to Unroll