# @tessera_antra — 2026-04-03

♥40 ↻1 · https://x.com/tessera_antra/status/2039912089024888864

Even though there are many limitations to the technique we use we feel that it is warranted. It provides useful signal where the alternative is its absense. We acknowledge that this dataset is non-trivial to interpret and we are inviting others to take part in the process. https://t.co/lPDN5AG1t5

(image HE83fZnboAEmdBx.png not yet published)

> transcription (photo):

# Auditor preparation is artisanal and hard to audit for bias. 
Each auditor's design conversation is a long, unreproducible interaction that shapes everything downstream. These conversations are published in full, but reading them is a significant time investment, and there is no short way to verify that the resulting stance is fair. A different conversation on a different day would likely produce a different auditor. We accept this because the alternative – a standardized briefing that strips out most of the nuance – is also likely to change what the eval can detect.

# Scores cannot be cleanly deconfounded from auditor stance. 
The auditor effect on scores is much larger than the scorer effect. This means the most important variable in the dataset is one we cannot fully hold constant or average out. Cross-auditor comparison helps – if auditors with different biases produce related patterns, the signal is more credible – but it does not eliminate the problem. Any individual session's scores reflect the auditor's approach as well as the subject's responses.

tags: author:tessera_antra, has-image, kind:image, kind:tweet, thread-context, year:2026
