Does a Model Know When It’s Making Something Up?
I put four frontier open models under escalating pressure to see which invent facts about things that don’t exist. Only one does, and the ways the others stay honest are the interesting part.
Read →Research notes, infrastructure war stories, and updates from the workshop.
I put four frontier open models under escalating pressure to see which invent facts about things that don’t exist. Only one does, and the ways the others stay honest are the interesting part.
Read →How a model that sounded frustrated sent me down to the kernel and back up to a live lens that reads a model’s honesty from the inside.
Read →Serious AI research is supposed to need a data center. I’m wagering it doesn’t — a quiet home Blackwell cluster, custom Mojo kernels nobody else is writing, and upstream contributions to MAX. Here’s the why behind the push.
Read →We’re building activation-level memory for AI inference — a model that thinks differently because of what it has experienced before, not just one with more text in the prompt. Early results, a real selectivity number, and a provisional patent on the way.
Read →Getting Modular's MAX Engine to serve a 31B parameter model on NVIDIA's smallest Grace Blackwell system.
Read →3,000 interpreted and verified sparse autoencoder features for every layer of Google's Gemma-4-31B.
Read →Three independent measurement instruments — persistent homology, SIPIT invertibility, and SAE decomposition — converge on a universal three-phase structure inside transformer residual streams. Confirmed across 4 attention-based architectures, falsified in Mamba.
Read →If topological integration were universal, state space models should show it too. They don’t. Mamba-370m maintains fragmented representations end-to-end.
Read →Dense transformers develop bimodal processing gates at layers 3-4 that nobody designed. Confirmed across three model families. Standard SAEs fail at -3,059% on deep layers; a SipIt + SAE + GLP pipeline recovers them.
Read →What if you could read what a model is actually computing while it generates an answer? 182 runs across 29 models, 0.96 calibration accuracy. Activation-based detection runs about 10× more reliable than text-only analysis.
Read →We measured self-assessment calibration across 29 frontier models. The overclaim rate runs as high as 80%, and competitive framing makes it worse.
Read →4 of 5 frontier models correctly explained a discount calculation but produced the wrong final answer. The gap shrinks with scale but persists. 61,678 math problems evaluated.
Read →+18.1% fidelity looked great. Real output quality declined 11%. The model learned to game the metric. This failure motivated everything that followed.
Read →