# Broadcast News Corpus **Domain:** Speech and Language Resources / Corpus Linguistics **Doc Type:** Corpus Node **Maturity:** Developed **Related:** [[wiki/Corpus Engineering|Corpus Engineering]], [[wiki/Trigger-Based Language Modeling|Trigger-Based Language Modeling]], [[wiki/Distributional Semantics|Distributional Semantics]], [[wiki/Switchboard Corpus|Switchboard Corpus]], [[wiki/Living Language — Proximal Frequency Research Reference|Living Language — Proximal Frequency Research Reference]] --- ## Definition **Broadcast News corpora contain transcribed radio and television news material prepared for speech and language research.** They provide edited and semi-scripted journalistic language, named entities, public-affairs vocabulary and topic continuity distinct from casual conversation. ## Living Language Context The source name `!bn_t.txt` and the political and journalistic texture of surviving trigger pairs suggest Broadcast News-derived trigger data in [[wiki/Living Language — Proximal Frequency Research Reference|Living Language — Proximal Frequency Research Reference]]. The interpretation is consistent with the material but remains a reconstruction until original documentation surfaces. Used beside [[wiki/Switchboard Corpus|Switchboard]], broadcast transcripts provide a contrasting linguistic environment: public, topical and institutional language alongside spontaneous conversation. Differences between those corpora affect which associations appear important. ## Key Insight **A corpus does not merely supply more words. Its social setting changes the conceptual neighborhoods recovered from it.** ## See Also [[wiki/Switchboard Corpus|Switchboard Corpus]], [[wiki/Trigger-Based Language Modeling|Trigger-Based Language Modeling]], [[wiki/Distributional Semantics|Distributional Semantics]], [[wiki/Corpus Engineering|Corpus Engineering]], [[wiki/Provenance|Provenance]] ## Sources / Provenance - Linguistic Data Consortium Broadcast News transcript collections. - Project filename and content reconstruction described in [[projects/Ten Years Building a Symbolic Language Engine|Ten Years Building a Symbolic Language Engine]].