Between the News
Published August 15, 2026 Β· Last reviewed August 15, 2026 Β· 9 min read
Guide
What Do Fact-Check Ratings Actually Mean? Pinocchios, Pants on Fire, and the Scale Nobody Agrees On
Media literacyFact-checkingPolitiFactWashington Post Fact CheckerSnopesRatings
πŸ‘Decoded
A politician says something. A few hours later a website puts a cartoon meter next to it, the needle swings to **Mostly False**, and the internet has its screenshot. Somebody else's meter says **Two Pinocchios**. A third site says **Mixture**. Same sentence. Three verdicts. So what does the meter actually mean? Short answer: less than the article under it. The rating is a packaging decision. The evidence is the fact-check. Almost everyone shares the first thing and reads the second thing never. Here is what each scale really measures, where the scales quietly disagree with each other, and how to read one properly in about thirty seconds. * **The Truth-O-Meter: six boxes, one of them famous** * PolitiFact launched the Truth-O-Meter in 2007 and effectively invented the genre β€” Poynter's Daniel Funke credited it in June 2019 as the format everyone else copied. The scale has six settings: True, Mostly True, Half True, Mostly False, False, and Pants on Fire. The two ends are the boring ones. True means the claim checks out. False means it doesn't. The interesting part is that PolitiFact runs a *sixth* box past False, and it is the one that goes viral. Pants on Fire is reserved for what PolitiFact calls the most egregiously misleading claims β€” and here is the detail almost nobody knows: the definition does not require intent. A statement can be rated Pants on Fire without any finding that the speaker knew it was untrue. The flaming-trousers cartoon reads as an accusation of lying. The rating rules do not actually make one. Which is exactly why PolitiFact holds the word "lie" back for one thing a year. The Lie of the Year, which it has awarded annually since 2009, goes to a *statement* β€” not a person. Staff comb through everything rated False or Pants on Fire, cut it to finalists, run a readers' poll, and then the editors pick. The readers' vote is published separately. It is not binding. * **Pinocchios: a four-point scale with three secret extra symbols** * The Washington Post's Fact Checker uses noses. One Pinocchio for a claim with some shading of the facts, up to Four for a whopper. More misleading, more nose. What the screenshots leave out is that the Pinocchio scale has non-numeric verdicts too. A claim that turns out to be straightforwardly true earns a **Geppetto Checkmark** β€” named for the woodcarver, the honest one in the story. A politician caught reversing themselves gets an **upside-down Pinocchio**, the flip-flop rating. And a claim that can't be resolved yet, because the data isn't in, gets a **scales of justice** β€” verdict pending, which is the most honest symbol in the entire fact-checking industry and also the least screenshotted. Then in December 2018 the Post added a rating that only exists because of repetition: the **Bottomless Pinocchio**. To earn it, a claim has to clear two bars β€” it must already have been rated Three or Four Pinocchios, *and* the speaker must have repeated it at least 20 times. The reasoning was that at some volume a false claim stops being an error and becomes a campaign. The scale gave up on grading the sentence and started grading the pattern. The scale of the thing that prompted it: the Fact Checker's Trump claims database, closed out on January 24, 2021, counted **30,573 false or misleading claims across four years** β€” everything rated two Pinocchios or worse. The rate is the part worth sitting with. About six a day in year one. Sixteen a day in year two. Twenty-two in year three. Thirty-nine a day in the final year. It took 27 months to reach the first 10,000, another 14 months to reach 20,000, and under five more months to pass 30,000. * **Snopes: the scale that admits the messy middle** * Snopes runs the widest set of labels, and it is the most useful one to study, because it stopped pretending every claim fits on a truth line. Alongside True, Mostly True, Mostly False and False, it publishes ratings for the situations that break a simple meter. **Mixture** covers a claim where significant elements are both true and false and no single verdict is fair. **Unproven** means there isn't enough evidence either way β€” not that it's false. **Miscaptioned** and **Misattributed** are for the real photo with the wrong caption and the real quote from the wrong mouth, which between them account for an enormous share of what actually circulates online. **Outdated** is for a claim that used to be true. **Legend** is for folklore that never had a source to begin with. Notice what that list tells you. Most viral misinformation is not a false statement. It's a true thing with the wrong label bolted on. A four-point true-to-false meter physically cannot express that, so it rounds it to "False" and loses the interesting part β€” which is *how* it went wrong. * **The part the meters don't show: the checkers disagree with each other** * This is the finding that should change how you read every rating you see. Chloe Lim, in a 2018 study in Research & Politics called "Checking how fact-checkers check," took the two biggest American fact-checkers β€” PolitiFact and the Post's Fact Checker β€” and tested them the way social scientists test any two coders rating the same material. First result: they barely overlap. Only about **one statement in ten** was checked by both organizations. The two most-cited fact-checkers in the country are mostly checking different sentences, which means the "fact-checked" universe is much narrower and much more selection-driven than it appears. Second result, on the claims both did check: agreement measured at a **Cohen's kappa of 0.52** β€” well below what would be accepted as reliable coding in a social-science paper. And the disagreement was not spread evenly. On outright falsehoods and obvious truths, they agreed fine. The agreement collapsed in the middle of the scale β€” the Half True, Mostly False range. Which is the whole game, because the middle is where politicians actually live. Nobody's press secretary is out there getting caught on the easy ones. The claims that matter are the ones built from a real number in a misleading frame β€” and that is precisely the zone where two professional fact-checking desks, looking at the same sentence, land on different boxes roughly half the time. * **So why have a meter at all?** * Because it travels. Funke's Poynter survey of the format put it plainly: ratings are easy to share on social media, they give readers a quick glimpse of a verdict, and they are an easy way to brand a fact-checking project. The Truth-O-Meter, the Pinocchios, the bloodhound mascot at Animal PolΓ­tico's El Sabueso in Mexico β€” these are logos as much as they are measurements. And fact-checkers themselves argue about it. Funke recorded the internal criticism: that rating scales give fact-checking "the appearance that it's based on social scientific process, while in fact it's a journalistic exercise." Others worry the label alienates readers who read the verdict as opinion and never make it to the evidence. Both worries are correct, and Lim's kappa is the receipt. * **How to read a fact-check in thirty seconds** * Skip the meter. Do this instead. **Find the claim in quotation marks.** Fact-checkers rate a specific sentence, not a policy, not a vibe, and not the thing you assume the politician meant. Half of all fact-check arguments are people disputing a rating of a sentence they never read. **Read the evidence paragraphs, not the verdict.** Look for named primary sources β€” the actual filing, the actual dataset, the agency that publishes the number. A fact-check whose proof is "experts said" is doing the same thing our guide on "according to sources" warns about in ordinary news reporting. **Check whether the disagreement is about the number or the framing.** If the number is agreed and the fight is over context, you are in the Half True zone β€” the zone where the checkers themselves split. Your own judgment is legitimately in play there. That is not a cop-out; it's what the kappa says. **Check who funds the checker.** The International Fact-Checking Network's Code of Principles asks signatories for nonpartisanship, transparency of sources, transparency of funding and organization, transparency of methodology, and an open corrections policy β€” and verifies it through external assessors. A fact-checker that won't tell you who pays for it is asking you to take on faith the exact thing it says you shouldn't. **Check whether it's been updated.** Ratings change when new evidence lands, and how an outlet handles that is its own tell β€” worth reading alongside our guide to what a correction, a clarification, an editor's note and a retraction each actually admit. * **The bottom line** * A fact-check rating is a headline. Like every headline, it is written to be carried, and β€” as we've written before β€” it is usually not written by the person who did the work. The scale is real and mostly well-intentioned. It is also, at the ambiguous end, two professional newsrooms coming to different conclusions about the same sentence half the time, then expressing that uncertainty as a cartoon nose. Read the noses if you like. But the fact-check is the paragraphs underneath, and that's the part that either has receipts or doesn't.
β€œOnly one claim in ten gets checked by both. On the ones that do, the two biggest US fact-checkers agree at a kappa of 0.52 β€” and the disagreement is concentrated exactly where politicians live.”
Comments (1)
media101prof
Sending this to everyone who cites "four Pinocchios" as if it were an SI unit. The scales disagreeing with each other is the point most people miss.
4d ago