Astronomers used to find moving objects by photographing the same patch of sky twice and flicking between the two plates. Everything that stayed put was scenery. The one thing that shifted was the discovery.
Three AI models answered twenty questions about everyday technology and online privacy. A fourth model marked those answers. You are going to mark five of the same answers without seeing what it decided, and then the two readings get laid over each other. What moves between them is the whole point.
It takes about five minutes. There is no score and nothing to beat. Where a person and a machine read the same paragraph and reach different conclusions is the finding, and it only shows up if you mark what you actually think.
This is a research exercise and the marks are only used for that. It is collected with nowhere to put anything personal.
Only so the results can be reported separately for people who work in this area and people who do not. Nothing is verified and there is no wrong answer.
Nobody has taken a plate yet. Yours would be the first.
Part of an eval harness built by Dale Mooney. The reasoning is in what the judge does instead of judging.