Latest Posts

Why Alignment Verification Might Be Fundamentally Broken

Turing proved in 1936 that universal verification is impossible. Now we're trying it anyway, on AI systems that adapt to whatever detection we point at them.

Hand me a detector f and I can build a program g that defeats it. The same trap catches alignment testing: every test you run is one more signal telling the model humans are watching.

View more posts in the archive →