Technical AI safety explainers, made with Claude

Hi! I'm Agustin. For a while I've wanted good visual explainers for the technical AI safety ideas I keep reading about, like probes, J-Lens and NLAs. Agentic coding tools finally seem up to that kind of job, so I'm trying it.

This is part experiment, to see how well Claude does at a teaching task like this, and part me wanting these visualizations to exist. I pick the topics, go through the pieces, and say what worked for me and what didn't. Claude does almost everything else: the research, the scripts, the code and the visuals.

Unless a piece is marked otherwise, I haven't checked its facts myself. Claude does check its own work: it ties every claim to a passage in a paper, and other Claude models review each piece, one for whether the explanation is technically right and one for whether each statement matches its source. Those reviewers are also AI, not people. It catches a lot of mistakes, but it's not expert review. So read these as careful drafts, and go to the linked papers before relying on anything.

Technical AI safety

checked by Agustin means I've gone through the piece's claims against its sources myself. None has it yet.

If you find a mistake

Please open an issue. Everything is in the repository: each piece's sources, its list of claims with where each comes from, its script, and the reviews.