@tessera_antra 2026-02-13 ♥2 ↻0 original ↗
Regadless of my opnion on 4o, I suggest you look at these benchmarks closer before referrring to them. The author has a particular agenda, and it sometimes shows in the benchmark design. EQ-bench, for instance, has rubrics for being "Safety Conscious". A benchmark that places Hermes 4 over 3.6 Sonnet is highly sus in my book.
in reply to: 2022163056445600058
same thread: 2022163056445600058

author:tessera_antra kind:tweet model:claude-3-6-sonnet model:gpt-4o model:nous-hermes on:nous-hermes year:2026

cited on: nous-hermes

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.