How well models write high-quality formal code
Monday, August 31, 2026
Today's models can formalize proofs from large mathematical projects. But as AI-generated formal code becomes dominant, a more important question arises: can models actually write good formal code?
We are building an evaluation that measures whether models can truly meet the standards of high-quality, maintainable formal code. The evaluation is being built with contributors to mathlib. Mathlib is Lean's community-maintained library of formalized mathematics.
We find that even today's most capable models struggle against this new standard.