but, yes, mostly current LLMs are bad and sloppy when it comes to writing fully correct proofs. i expect this to be pretty temporary and to improve quickly with better prompting and models. grok 3 and gemini 2.5 are good at figuring out what's true and generating counterexamples
in reply to: 1907455663569293450
cited on: grok-3
Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.