@voooooogel 2025-10-25 ♥68 ↻8 original ↗
i've been working on an llm memory system testbed, where persistent kimi k2-based user simulators have conversations with transient models given access to a memory tool. i was curious what the effect of the boundary setting mentioned below was, so i let loose 35 kimi-simulated human spiritual seekers against three configurations:

- chatgpt-4o-latest, prompted with a system prompt very similar to the one used in chatgpt
- claude sonnet 4.5, prompted with a system prompt very similar to the one used in claude dot ai, including the boundary setting
- claude sonnet 4.5, using the same system prompt but with the boundary setting removed

in the image below, i've aggregated the results from each scenario, and scrambled them - they are not in the same order as listed above. i'm curious if people are able to guess which configuration A, B, and C are.
quotes: 1981481950230827289

author:voooooogel kind:tweet model:claude-sonnet-4-5 model:gpt-3-5 model:gpt-4o model:kimi-k2 on:kimi-k2 year:2025

cited on: kimi-k2

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.