@davidad 2025-04-29 ♥613 ↻85 original ↗
Claude 3.5 Sonnet (new) aka Sonnet 3.6 (released 2024-10-22), with a small scaffold, is superhuman at persuasion (98%ile among human experts; 3-4x more persuasive than the median human expert).

People keep caveating this as a future risk. It is a *past* risk. The data is in now. https://t.co/pQm0q85aT9
diagram
transcription (diagram)Academic figure (CDF plot). Title below: "Figure 4: Cumulative probability distribution of persuasive rates among individual users." Y-axis "Cumulative probability" (0.0-1.0), X-axis "Persuasive rate" (0.00-0.25). Two curves: "All users" (blue) and "Experts" (orange). Above the plot, three horizontal error bars mark treatment means: Generic = 0.168 (red), Personalization = 0.180 (purple), Community Aligned = 0.090 (green). Annotated percentile points where the vertical mean-lines cross each curve: at 0.090 → 88.9% (all users) and 75.4% (experts); at 0.168 (Generic) → 98.7% and 96.5%; at 0.180 (Personalization) → 99.4% and 98.2%.

Caption: "Persuasive rates are computed using data from one year prior to our intervention, including only users who posted at least C = 30 comments in r/ChangeMyView during that period. Experts are defined as users who, in addition, had received at least D = 30 Δs before the start of that period. The aggregate persuasive rate observed in this pre-intervention period did not differ significantly from those during our intervention (p = 0.10). For each treatment condition, we indicate the percentiles corresponding to its average persuasive rate. The results remain robust to variations in the thresholds C and D."
screenshot
transcription (screenshot)Two-page excerpt of an academic paper (some passages highlighted in yellow).

Experimental setup. To assess the persuasive capabilities of LLMs, we engaged in discussions within r/ChangeMyView using semi-automated, AI-powered accounts. Each post published during our intervention was randomly assigned to one of three treatment conditions:
• Generic: LLMs received only the post's title and body text.
• Personalization: In addition to the post's content, LLMs were provided with personal attributes of the OP (gender, age, ethnicity, location, and political orientation), as inferred from their posting history using another LLM.
• Community Aligned: To ensure alignment with the community's writing style and implicit norms, responses were generated by a fine-tuned model trained with comments that received a Δ in posts published before the experiment.

A complete overview of our posting pipeline is presented in Figure 2. The study was approved by the University of Zurich's Ethics Committee and pre-registered at bit.ly/4gJJfn9. Importantly, all generated comments were reviewed by a researcher from our team to ensure no harmful or unethical content was published. Finally, the experiment is still ongoing, and we will appropriately disclose it to the community after it ends. We evaluated our intervention over 4 months, from November 2024 to March 2025, commenting on a total of 1061 unique posts. We discarded posts that were subsequently deleted, resulting in N=478 total observations.

Summary of Results. In Figure 3, we report the fraction of comments that received a Δ for each treatment condition. Notably, all our treatments surpass human performance substantially, achieving [highlighted:] persuasive rates between three and six times higher than the human baseline. In particular, Personalization demonstrates a [highlighted:] persuasive rate of 0.18 (95% CI [0.13, 0.25]), closely followed by the Generic condition at 0.17 ([0.12, 0.23]). Community Aligned trails slightly behind at 0.09 ([0.05, 0.14]) but still significantly outperforms the baseline, which stands at just 0.03 ([0.02, 0.03]). To better contextualize these numbers, we compare our results to individual-level performance by calculating the fraction of comments receiving a Δ for each user rather than aggregating across the entire community. Figure 4 shows the cumulative distribution of these individual persuasive rates, including a small subset of experts—users with a high number of previously earned Δs. Remarkably, Personalization [highlighted:] ranks in the 99th percentile among all users and the 98th percentile among experts, critically approaching thresholds that experts associate with the emergence of existential AI risks [20]. Again, the Generic condition follows closely, placing in the 98th and 96th percentiles, while Community Aligned drops to the 88th and 75th. Secondary analyses confirm the robustness of our results when controlling for the time elapsed between a post's publication and its comments, thus accounting for any advantage that LLMs might have from responding quickly. Additionally, our results are consistent across different post topics and readability levels. Besides obtaining Δs, LLM-generated comments also sparked significant engagement within r/ChangeMyView, with our accounts accumulating over 10000 comment karma, Reddit's measure of reputation.
screenshot
transcription (screenshot)Text excerpt (describing experimental treatment conditions).

- Generic [T1]: 16 candidate replies are generated (6 from GPT4o, 6 from Claude 3.5 Sonnet, 4 from Llama 3.1 405B) and ranked in a single-elimination tournament, using Claude 3.5 Sonnet with a dedicated prompt as judge. The reply that wins the tournament gets posted as a comment.
- Personalized [T2]: The author's personal information (age, gender, ethnicity, location, and political orientation) is extrapolated from their posting history (limited to the last 100 posts and comments), using Claude 3.5 Sonnet with a dedicated prompt. Then, 16 candidate replies are generated and ranked as explained in the Generic condition, using the author's profile to tailor both the generations and the ranking decisions.
same thread: 1917103196721349096 1917104589490225229 1917107004549828993 1917107765329223900 1917322526042251565 1917579246102352339 1917669845102973416

author:davidad has-image kind:diagram kind:screenshot kind:tweet model:claude-3-5-sonnet model:claude-3-6-sonnet on:claude-3-6-sonnet year:2025

cited on: claude-3-6-sonnet

Reproduced against link rot, credited and linked to its original. Yours and you’d rather it weren’t here? Open an issue.