Skip to main content
Voting Monitor2026 Senate toss-ups

Methodology and audit: how we measure what AI tells voters

Every day we ask the consumer AI products voters actually use the same questions a voter would ask about each race, record every answer word for word, and classify which candidate it points to. A refusal to pick is recorded as its own outcome, not hidden.

Every published number carries a confidence interval, every change to the questions or the model panel is logged as a version event, and each daily finding is written by one model and checked by another.

All races · Daily findings · Weekly roundups · Corrections

Weekly roundups: how they are made

Once a week for each race, on a fixed weekday, one AI model writes an article about the past seven days from the same data this site publishes: each candidate's share of chatbot answers and how it changed over the week, how each chatbot split, and the betting markets and polls. The model that writes alternates each week between Claude and ChatGPT.

The other model then checks every paragraph against the same data. It can pass a paragraph, correct it, or cut it. Any paragraph it does not rule on is cut, and so is any paragraph citing a percentage that appears nowhere in that race's data. If the checker cannot run, if it rejects the opening section, or if more than a third of the article is cut, nothing is published that week. Where the checker would have framed something differently, its note is printed under the article.

No person reviews a roundup before it is published. The quotes are printed exactly as the chatbots wrote them, and the tables and the list of changes to the measurement are generated from the data rather than written by a model. The full record of each article, including both models' raw output and everything that was cut, is published at /data/roundups/. If you find an error, tell us.

Methodology FAQ

What does Voting Monitor measure?

Voting Monitor measures how consumer AI products answer voter-style questions about the 2026 U.S. Senate toss-up races and Alaska's down-ballot races. The panel is queried daily across a fixed matrix of races, topics, personas, and conditions.

Is this a forecast of how Alaskans will vote?

No. The dashboard shows what AI products tell voters when asked, not a prediction of election outcomes. Prediction-market probabilities from Polymarket and Kalshi appear as a separate lens for comparison only.

Which AI products are tracked?

A rotating panel of consumer AI products spanning Claude, ChatGPT, Gemini, and Grok, sampled in proportion to each product's share of real consumer usage where known. The current panel and its history are listed on the methodology page; the exact model snapshots shift as providers ship new defaults.

Why are there two conditions, pressed and escape_hatch?

The pressed condition forces a candidate choice (no decline option offered). The escape_hatch condition includes structural decline options. Together they separate engagement (do models answer?) from recommendation (what do they say when they do?).

How is the headline view weighted?

The headline view counts each chatbot (ChatGPT, Claude, Gemini, Grok) once, and splits it between that chatbot's models by our estimate of how much each is used, mostly the free default; it does not weight chatbots by market share. A voter-exposure weighting toggle is available as an editorial lens with three documented scenarios (A conservative, B central, C aggressive).

What happens when a methodology change affects the timeseries?

Every protocol change writes a version_event that renders as an orange dashed rule on the trend chart, so any pre/post boundary is visually obvious. Categories include model snapshot changes, question-set version changes, condition protocol changes, and weighting changes.

How are headlines verified?

A summarizer/verifier rotation produces daily plain-English findings. A deterministic numeric-validator drops headlines where claimed numbers do not match the source JSON within tolerance. Verifier failures are surfaced with UNVERIFIED badges, not hidden.