Verbatim Pundits

How intellectually honest political commentators are in what they actually say on the record.

Scores describe argumentative behaviour observed in sampled recordings. They are not fact-checks, and they are not judgements of a person's sincerity or character.

Pilot board. This is an early run, published so the method can be checked in the open rather than because the evidence is settled. It ranks 10 of 39 people on the roster, each on 5 to 12 recordings. Blinding removes the name from the transcript and does not make the speaker unknown: the judges still identified them on 100% of blinded recordings, so read the blinded score as name-removed rather than anonymous. Scores carry no correction for the format of the appearance. It was fitted and then withheld, because the corpus does not yet support it (venue 'debate' seen 4 times, under the minimum of 8). A reaction stream and a sit-down interview are not equally demanding, and that difference sits inside these scores.

RankPersonBlinded overallSteel-manning and charityEpistemic rigor and calibrationGood faith and consistencyOpen overallHaloRecordingsConfidence
1Ezra Kleincolumnist and podcast host73.5 [67.1, 79.1]71.773.875.471.8-1.77high
2Coleman Hugheswriter and podcast host65.0 [62.4, 67.8]61.566.067.964.9-0.17high
3Steven Bonnellstreamer and political commentator57.8 [50.8, 65.0]57.259.955.958.1+0.46high
4Sam Sederpodcast and YouTube show host51.6 [48.3, 55.5]47.654.153.553.3+1.47high
5Ana Kasparianonline news show host and commentator48.0 [41.8, 53.6]40.250.654.248.5+0.36high
6Ben Shapiropodcast and radio host42.4 [36.6, 48.8]36.149.141.941.6-0.912high
7Hasan Pikerstreamer and political commentator34.6 [27.6, 42.1]32.036.835.134.6-0.18high
8Zack Hoytstreamer and political commentator34.2 [26.6, 43.2]30.635.137.435.7+1.411high
9Charlie Kirkpolitical activist and radio/podcast host32.5 [29.8, 35.1]33.830.933.032.3-0.35high
10Matt Walshpodcast host and author31.3 [26.8, 35.3]26.333.634.432.0+0.88high

What the scores mean

In the blinded score the judges read the transcript with the speaker's name and show removed. That does not make the speaker anonymous: a well-known commentator is often recognisable from their positions, and one of the two judges can search the web, which is measured and disclosed rather than blocked. The open score gives the judges the name; the difference between the two is shown as the halo.

Steel-manning and charity (35%)

Does the speaker represent the other side fairly, at its strongest? Credit for stating an opposing view accurately when that view appears in the recording, for answering its strongest form rather than a weaker one, and for naming where the real disagreement lies. The speaker loses credit for claiming bad motives without evidence given in the recording.

Epistemic rigor and calibration (35%)

Does the speaker's confidence match the support they give? Credit for backing load-bearing claims with stated evidence, a source or a mechanism, for separating fact from interpretation and speculation, for handling numbers consistently, and for saying what would change their mind. This checks the reasoning shown, not whether the claims are true.

Good faith and consistency (30%)

Does the speaker argue in good faith and hold one standard? Credit for applying the same standard to cases set side by side, for conceding a good point, for answering the question actually asked, for avoiding insults and insinuation, and for departing from their own side with reasons when the recording shows what that side is.

Method

Each recording is transcribed, the speaker's name and show are removed for the blinded score, and two AI judges score it against a fixed rubric: Claude Fable and Gemini Flash. Each judge's scale is calibrated onto a shared one, a person's score is the mean over their recordings, and the interval is a bootstrap over those recordings. A dimension a recording gives no real chance to show is marked not observed and does not count.