# @QiaochuYuan — 2025-04-02

♥64 ↻0 · https://x.com/QiaochuYuan/status/1907455665976787434

but, yes, mostly current LLMs are bad and sloppy when it comes to writing fully correct proofs. i expect this to be pretty temporary and to improve quickly with better prompting and models. grok 3 and gemini 2.5 are good at figuring out what's true and generating counterexamples

tags: author:qiaochuyuan, kind:tweet, model:gemini-2-5-pro, model:grok-3, on:grok-3, year:2025
cited on: _dossiers/grok-3.md, grok-3
