How models fail
since 2023
Learning how models actually work and fail: hallucination, calibration, refusal, the transformer from the inside. I run new models through my own set of tasks rather than simply trusting public announcements.
Updated October 6, 2026. Where my time goes at the moment.
since 2023
Learning how models actually work and fail: hallucination, calibration, refusal, the transformer from the inside. I run new models through my own set of tasks rather than simply trusting public announcements.
since 2024
Developing the AI system for sports education that we launched in 2024.
since late 2025
Building a safe perimeter for coding agents. Claude Code, Codex and the like work where there are secrets, customer data and a payments zone. I am not banning agents; I am making the ways they are used safe: policies, an audit trail of what agents do, and prompt-injection testing before and after.
since spring 2026
Building a data agent for reporting: a question in plain language goes in, a checked and well-presented answer from the data comes out. The main risk is a plausible wrong number, so explicit metric definitions and a set of questions with reference answers come first, and the model second.
fall 2026
Writing posts for this site.