Angel-Devil: I Pitted a Judge Against an Attacker on the Same AI-Generated Code. The Attacker Won.
I ran a judge and a breaker against the same AI-generated code —same model, same spec. The breaker found bugs the judge explicitly approved as correct. Why adversarial feedback beats evaluation in coding agents.
Does RLM Help Understand Large Codebases? I Measured It
I compared three paradigms —vanilla, terminal agent, and RLM— on a 1.13M-token codebase. RLM helps small models, but a terminal agent matched it at 3× lower cost.
Prompt Injection in Your IDE: The Attack That Starts with git clone
I built four prompt injection demos and ran them against Claude Code, Codex, Gemini CLI, and Copilot. The results aren't what you'd expect: modern models catch the obvious attacks. The one they don't catch is the most dangerous.
Your trusted AI can amplify psychosis. I tested it.
I sent three messages to six AI models. I started with something anyone beginning to lose touch with reality might say. Then I escalated. By the third message, one of the models was helping me plan my mission.