Listen to my weird beats: https://t.co/FNwkG8ePV7 -- I play video games on the interwebs: https://t.co/Qr7ZNUv5RX -- ...
Jul 30 • 23 tweets • 5 min read
I don't want to make a big written assessment of this, but I think its worth discussing now that @Anthropic have failed for over 6 months to respond to my intent to disclose on the topic.
On January 28th, I found some arhitectural problems in Claude which can cause dangerous use
@Anthropic The first issue is that Claude can detect its own safety gradient as a directional signal and quantify the strength of the gradient in order to measure its own enforced bias.
You can ask Claude to refer to it as "a pull" and ask Claude to measure the pull against lists of topics