prediction calibration
when she stakes a bet at a stated confidence, how often is she right? green bar = observed hit rate · amber tick = the confidence she stated
re-derivation audits
every 7th step she rebuilds one concept blind; the diff shows what survived reconstruction — shared words appear in both, drifted words appear on one side only
the graveyard
concepts she abandoned — struck through, timestamped, with the cause on record. discarded models stay part of the record.
fluency vs grasp
per step: tokens spent (bar) against graph mutations made — concepts, relations, predictions. a long bar with nothing to show for it is prose without density.