Teardowns of real incidents, the graph theory behind the detectors, and the benchmark numbers with the commands that produce them.
Every agent security tool grades its own homework. A proposal for an open benchmark, and our own scores on it
The lattice, the reachability query, and the min vertex cut, in 985 lines. Plus the one question the static model provably cannot answer, which is where the runtime half has to begin.
Three LangChain and LangGraph CVEs, what each one actually does, and why the kill chain you have read about does not exist
The payload in CVE-2026-25724 was a comment in a README, which makes natural language an execution path
76% recall, the twelve cases that get past us, and the detector we switched off on purpose
Every scanner publishes what it catches. Almost none publish a command you can run to check. Here is my false-positive count, the number that argues against it, and both commands.
A malicious dataset, a tool that runs code, a credential read, an outbound call. Four ordinary tools, none of them a mistake on its own. Drawn as a graph, they are one edge, walked 17,600 times.
Private data, untrusted content, external reach. Any one is fine, all three is an incident. Open your tool file: this is the ten-minute pass that tells you which edges to cut and which to gate.
A credential harvester in a library with 95 million monthly downloads. The way in was a security scanner, which is the reason this post spends a section on why you should not trust mine.
More posts as we publish them. Every claim links to the command that reproduces it.