Blog

Notes on things I have worked on, most recent first.

Depth, dependency, and the cost of binary grading

Verifier-reward RL for theorem proving: a dependency graph reconstructed by necessity, why all-or-nothing grading distorts a compositional difficulty curve, and what it costs to train a 7B prover through the full depth curriculum.

Unlearning is a sequence

What happens to privacy when a model provider switches unlearning mechanism mid-stream. The cross-stage leak, how we measured it, and the alignment objective that removes it.