“Let’s Bring Back Hope”: The Lexical Appendix
The Vocabulary of the Reply
What the public actually said back when Andy Burnham asked the country to bring back hope — read as words rather than scores. Every cloud below is drawn from the scored corpus of the first 48 hours, cut by platform, by stance toward the appeal, and by DDI band.
The corpus as a whole
All platforms, all stancesOne word dominates the reply, and it is Burnham's own. Hope is the most frequent content word in the corpus — but it is overwhelmingly returned to sender: quoted, bargained over, or thrown back. The words standing closest to it are not affective but procedural — election, country, people, Labour.
All comments
Every scored comment, four platformsn = 8,086 commentsTwo-word phrases
Recurring pairs across the whole corpusn = 8,086 commentsBy platform
Four arenas, four vocabulariesEach panel shows two clouds. The left is raw frequency — what that audience talked about. The right is distinctiveness: terms weighted by how much more they occur here than in the rest of the corpus, which strips out shared vocabulary and leaves each arena's signature. The frequency clouds look broadly alike; the distinctiveness clouds do not. Tap any cloud to open it full size.
X
The most hostile arenan = 1,412 commentsTikTok
Direct address, retail politicsn = 1,048 commentsBy stance toward the appeal
Embrace · Conditional · Reject · Off-topicThe stance cuts show four distinct grammars. Embrace speaks in the vocabulary of ceremony (congratulations, luck, wishing). Conditional speaks in the vocabulary of terms and proof (actions, earn, louder, otherwise). Reject speaks in the vocabulary of character and fraud. Off-topic barely engages the appeal at all — it uses the thread as a distribution channel.
Embrace
Accepts the hope framingn = 1,201 commentsConditional
Hope offered on termsn = 1,560 commentsReject
Refuses the framing outrightn = 4,561 commentsBy DDI band
Where composite score meets vocabularyThese panels use signature terms rather than raw counts, because the raw counts across bands are near-identical — the same subjects appear at every score. What separates the bands is not what people discuss but how. Risk is carried by dehumanising and criminalising terms; Concerning by procedural and remedial ones; Mixed and Healthy, a residue of 138 comments, by specific personal circumstance.
Risk
Composite 0–24n = 6,149 commentsConcerning
Composite 25–49n = 1,795 commentsMixed and Healthy
Composite 50 and aboven = 138 commentsHow these clouds were made
Method| Step | Treatment |
|---|---|
| Source | The scored corpus, read with quoting disabled — stray quotation marks inside comment text otherwise merge rows and silently drop about forty comments. |
| Cleaning | URLs removed; @ and # handles removed; the leading capitalised name tag that Facebook prepends to threaded replies removed. The handle andyburnham is folded into burnham. |
| Stopwords | Standard English function words plus platform furniture. Negation particles are removed as single tokens but retained inside phrases. Hope, stop, need, back and let are deliberately kept — they are the substance here, not noise. |
| Frequency clouds | Up to 95 terms on desktop, sized by count with relative scaling so a single dominant term cannot flatten the rest. On a phone a reduced cut of the same data is shown, around 25 terms, so the type stays legible; the full cloud opens on tap. |
| Distinctiveness | Log-odds of a term in the cut against its rate in the rest of the corpus, damped by log frequency, minimum eight occurrences. Only over-represented terms are shown. |
| Caution | Word clouds show salience, not meaning. A term's prominence says nothing about the stance it carries — hope is the clearest case, appearing most often in comments that reject it. Read each cloud against its band and stance cuts, never alone. |
Lexical appendix to the Democracy Discourse Index · United Kingdom study
July 2026 · gcrd.org.uk
Read the the full report at GCRD Insights

