How well models write high-quality formal code

Monday, August 31, 2026

Today's models can formalize proofs from large mathematical projects. But as AI-generated formal code becomes dominant, a more important question arises: can models actually write good formal code?

We are building an evaluation that measures whether models can truly meet the standards of high-quality, maintainable formal code. The evaluation is being built with contributors to mathlib. Mathlib is Lean's community-maintained library of formalized mathematics.

We find that even today's most capable models struggle against this new standard.

Request early access.