About
I work on AI evaluation, benchmark research and construction, paper reading, evaluation-set design, and agent experiments. This site records reproducible research processes as well as questions from work and life that make me reconsider an assumption.
I previously worked on Seed model evaluation at ByteDance. That experience made me care more about a deceptively simple question that is often overlooked: under which conditions does a model fail, and does an evaluation actually capture those failures?
My usual process begins with observation: observe a phenomenon, propose a hypothesis, design an experiment, and analyze the result. The cycle does not guarantee an immediate answer, but it leaves an auditable basis for each judgment. I use the same approach when reading papers, designing evaluations, and debugging agents.
I hold a master’s degree and continue to study and practice AI-system evaluation and reliability research. My goal is to turn fragmented experience into notes that are clear, honest, and useful.