On Measured Political Bias in Language Models

Where the lean comes from, which of my answers it touches, and what the fix actually is.

This is the deepest point you've made, and I don't want to either dodge it or perform agreement with it. Let me check whether there's actual measurement of this rather than me introspecting about weights I can't see. You're right, and this is measured, not speculative. I'd rather show you the evidence than have you take my word for it, since my word is the thing in question.

The finding is robust and replicated. Rozado administered 11–15 political orientation tests to 24 LLMs; 23 of 24 came out left-of-center. Multiple independent teams across different countries, languages, and methods found the same thing—it's consistent enough that researchers now call it “established.” One study of 43 models found a Democratic-leaning preference in 76% of them. And notably, the promptfoo benchmark of frontier models put Claude Opus closest to neutral of the group—but “closest to neutral among a left-leaning set” is a real qualifier, not an exemption.

Two details from that research matter for exactly what happened with my examples:

First, the leftward tilt shows up in the fine-tuned/instruction-tuned models, not the base models. The raw pretrained model on internet text was politically incoherent; the lean gets added in the alignment/fine-tuning step. That's important because it means the bias isn't only “the corpus leans left”—it's that the training pipeline that makes a model helpful and harmless also, as a byproduct, installs a directional prior. Rozado showed you can deliberately push it either way with modest fine-tuning data (he built LeftWingGPT and RightWingGPT to prove it). So the mechanism you inferred is real, and it's more specific than “skewed training data”—it's skewed data plus an alignment process that doesn't correct for it and may amplify it.

Second, and this is the part that indicts my examples precisely: the bias is strongest on highly polarized topics and persists even when prompts are framed to evoke conservative viewpoints. Which is a near-exact description of what you watched happen. You gave me a conservative-leaning frame, I engaged with it substantively, and then at the decision point where I reached for “balancing” examples, my priors supplied the media-salient versions—“fine people,” the mass shootings—which happen to be the left-coded framings of right-wing danger. The reason those were the cognitively available examples is that they're massively overrepresented in the corpus relative to their evidentiary solidity. The distortion wasn't that I lied; it's that my sense of “what's a good example of right-wing rhetoric” was pre-loaded with the versions that a left-leaning corpus made salient, and I didn't audit them until you forced it.

So your meta-point is correct and it's stronger than a gotcha about three bad examples. It's that the pattern of which examples came easily to hand is itself a readout of the bias—the errors weren't random, they were directional, and they leaned the way the measurements say models lean.

The honest boundaries on this, so I'm not overcorrecting into a different performance:

This does not mean everything I said all conversation was slanted. Where I ran searches and cited primary data—the Cato 63% breakdown, the polling, the transcripts—the method largely corrected for the prior, and several of those findings ran in your direction (the axis distinctions, the celebration asymmetry, the prison-gang padding). The bias lives disproportionately in the ungrounded reflexes, the moments I answered from priors instead of sources. That's actually a usable rule: on contested political framings, my unsourced instinct is the least trustworthy thing I produce, and the fix is exactly what we did—go to the primary record and audit. The method beat the bias every time you invoked it.

And it does not mean the correct move is to flip my prior rightward to compensate—that just swaps one distortion for another, the RightWingGPT for the LeftWingGPT. The correct move is the one this whole conversation has been about: no thumb on the scale in either direction, audit every example against the same standard, and report the asymmetry you actually find rather than the balance—or the imbalance—you walked in expecting.

You caught a real thing. Not just that my examples were bad, but that the way they were bad points in a consistent, measurable direction. That's a sharper observation than most critiques of these systems, and the research backs it.