Methodology
What Press Them measures, how it measures it, and what the numbers do not mean.
One scale for both countries
Every outlet, US or UK, sits on a single left-to-right scale anchored on the US center. "Center" always means the US center. We do this on purpose: the UK's mainstream press sits to the left of America's, and a shared scale lets you see that lean instead of hiding it inside two separate national scales.
Each outlet also has a domestic position (where it sits within its own country's press). We keep it for balancing the roster, so every wing has counterparts on both sides. Readers only see the shared position.
Temperature: how warmly a story is told
Favorability toward the story's main subject runs from −10 (hostile) to +10 (glowing). We show it as temperature: hostile coverage is cold, neutral is warm, flattering is hot.
Ice cold-8.0
Cool-4.0
Warm0.0
Hot+4.0
Scorching+8.0
The story dial shows the median across outlets. Heat strips show one tile per outlet, ordered left to right, so you can see whether heat clusters on one wing.
Discovery
Each edition starts from a fixed roster of US and UK outlets, deliberately balanced so that every point on the spectrum has counterparts on the other side. For each outlet we try, in order: its own RSS or Atom feed, an OpenRSS mirror, its news sitemap, a shallow crawl of its front and section pages, a Google News feed for the outlet's domain, and GDELT. We respect robots directives and never attempt to bypass paywalls or bot protection.
Where an outlet is behind a paywall or blocks automated access, we record that status and work from the headline, standfirst and any publicly visible summary only. Those articles are labeled so you can see the difference.
Clustering and matching
Articles are embedded, grouped by vector similarity, and borderline cases are adjudicated by a model. Each story gets a centroid; US and UK stories are matched into transatlantic pairs using those centroids, with a model check on close calls. Where no counterpart exists, we say so rather than forcing a pair.
Blind multi-model scoring
Before scoring, article text is redacted: outlet names, bylines, mastheads and other identifying boilerplate are removed, and the order of articles is shuffled. Each active scorer — currently models from more than one vendor — then rates the same redacted text against the same versioned prompt and the same worked calibration anchors. Prompt text is hashed and stored with every score, so a result can be traced back to the exact instructions that produced it.
The published favorability and lens numbers are the ensemble median. We also record each model's own score, the spread between them, and each model's average drift from the median over time. That last number is the models' bias, not the outlets' — it is published on the model spin tracker for the same reason we publish everything else.
Coverage gaps and omitted facts
For each story we build a matrix of country by political bucket, weighted by each outlet's audience reach, and flag where a story is carried heavily on one side and barely at all on the other. Separately, we extract atomic factual claims from the coverage, merge duplicates, rate how material each claim is, and record which outlets included it.
An empty cell or an omission means the item was not found on the surfaces we monitor at the time we checked. It is not proof that an outlet never covered it — feeds lag, paywalls hide text, and sections we do not crawl exist.
Limits
Lens placements for outlets are editorial judgments, published openly and adjustable. Scores are model output, not fact. Every score is reproducible from the stored prompt version, model id and inputs, and the whole roster is visible on the outlets page so you can judge the frame we are judging through.