AI Evals · Benchmark · Agent Experiments

I study how AI systems fail — and how we can evaluate them better.

Notes on AI evaluation, benchmark design, agent experiments, and research—alongside observations from work and life that deserve a second look.

Recent articles

View all articles →