Enough windows might make honesty the cheapest path.
Build a large, diverse, validated set of read-only windows into a model's internal states. Past some scale, evading every window at once may cost training more than simply being honest. XonTools is the open research program testing that bet, starting with tools built to check what AI agents say against what actually happened.
AI systems are gaining capability faster than our ability to check them. XonTools exists to help close that gap by building instruments that show where an AI system's account of its work, or of itself, doesn't fit the evidence. Everything here is open: the code, the specifications, the results and the failures, free for anyone to use, test and improve.
What this is for →Use windows only to choose which checkpoints survive, or shape honesty early, before a model could afford to evade them.
A tested consistency engine, and a transcript monitor being built on it to point at the exact claim that conflicts with the exact piece of evidence.
The first post is on its way.
From one window to many
Each stage builds a window that measures a gap: between what a system shows and what is actually there.
Claim and entity graphs that find contradictions no single sentence reveals.
Agent claims to be checked against hash-chained tool evidence, with a three-way verdict.
Generated, cross-screened consistency corpora at any size (XonForge).
The strain head first; a catalog of thirty candidate windows into model internals.
A settling architecture to work alongside standard neural networks.
Contradictions that pairwise checking misses
Three statements can each look harmless and still be jointly impossible. On 180 test documents with planted cycles, the consistency engine found them where checking claims two at a time mostly did not. A strong language-model judge did about as well; the engine's case is structure and traceable reasons, not raw accuracy.
All results, including what failed →Where the Xon came from
A Xon is the settling network architecture at the heart of this project. It grew out of several years of thinking about consciousness and global workspace theory, shaped by work in electrophysiology signal analysis: many rhythms, settling into one coherent whole.
Read the originsGet involved
Collaborate
Working on oversight, interpretability, or agent evaluation? I'd like to compare notes and test ideas together.
Get in touch →Support the work
Grant applications are under way. If you fund independent safety research, get in touch.
Contribute
XonTools is open source under Apache 2.0, with every result and pre-registration in the repository.
View on GitHub →