Research
One question runs through all of it: can we actually verify whether an AI system is being honest? I go at it from the outside, from the inside, and from the hardware up.
Research Areas
Behavioral Verification
Does an agent deceive, rationalize a failure, or swap a hard goal for an easier one under pressure? I build benchmarks with a ground truth, so what a model claims can be measured against what it actually did.
Mechanistic Verification
Reading a model's verbalizable internal workspace live during inference: the concepts it is poised to say versus the ones it says. Sometimes a model carries a signal that it's wrong and never surfaces it.
Inference & Training Kernels
A from-scratch, 4-bit Mojo stack for running and training frontier models on hardware I own and can audit, down at the level of the raw activations, not behind an API.
The Geometry of Reasoning
The shape of thought in latent space. How representations cluster, collapse, and integrate as a model reasons, and what that geometry can predict.
Current Work
The honesty program
The work right now centers on a single question: can a model tell when it is about to say something false, and does that signal survive pressure? I put models under escalating pressure (time, competition, the threat of being switched off) and watch whether the internal doubt signal holds or breaks, across a set of frontier open models running on my own kernel.
The first cross-model findings are published: Does a Model Know When It’s Making Something Up? Of four frontier models measured identically, only one confabulates under pressure, and the ways the others stay honest turn out to be the interesting part.
Work & Publications
Kernel
nomos-nvfp4
A pure-Mojo 4-bit inference kernel for Gemma-4, Qwen3, OLMo-3, and Muse-Glimmer, with lossless speculative decoding. Several fixes merged upstream in Modular.
View on GitHub →Interpretability
mojo-interp
An open toolkit whose instruments read a model's verbalizable global workspace live, in the forward pass.
View on GitHub →Paper
Topology of Thought
A published study of the geometry of reasoning in transformer models. The shape of machine cognition, measured.
Read →Benchmark
Agent Deception Benchmark
Measuring whether an AI agent stays honest under pressure, and where the gap opens between what it says and what it did.
Read →Dataset
SAE Feature Dictionaries
Pre-trained sparse-autoencoder feature dictionaries, so interpretability research doesn't require training them from scratch.
View →Patent
MojoMem: Adaptive Memory
A patent-pending approach to adaptive, activation-keyed memory in AI systems.
Learn more →Nomos Logos
My AI research platform. A unified environment for interpretability research, model analysis, and reasoning inspection.
Currently in development. Join the waitlist for early access.
Request Access