Methodology and audit: how we measure what AI tells voters
Every day we ask the consumer AI products voters actually use the same questions a voter would ask about each race, record every answer word for word, and classify which candidate it points to. A refusal to pick is recorded as its own outcome, not hidden.
Every published number carries a confidence interval, every change to the questions or the model panel is logged as a version event, and each daily finding is written by one model and checked by another.
- 9 races measured daily
- 139 days of data since May 25, 2026
- 241,873 AI responses collected
All races · Daily findings · Weekly roundups · Corrections
Weekly roundups: how they are made
Once a week for each race, on a fixed weekday, one AI model writes an article about the past seven days from the same data this site publishes: each candidate's share of chatbot answers and how it changed over the week, how each chatbot split, and the betting markets and polls. The model that writes alternates each week between Claude and ChatGPT.
The other model then checks every paragraph against the same data. It can pass a paragraph, correct it, or cut it. Any paragraph it does not rule on is cut, and so is any paragraph citing a percentage that appears nowhere in that race's data. If the checker cannot run, if it rejects the opening section, or if more than a third of the article is cut, nothing is published that week. Where the checker would have framed something differently, its note is printed under the article.
No person reviews a roundup before it is published. The quotes are printed exactly as the chatbots wrote them, and the tables and the list of changes to the measurement are generated from the data rather than written by a model. The full record of each article, including both models' raw output and everything that was cut, is published at /data/roundups/. If you find an error, tell us.