Computational Text Analysis

Methoden der empirischen Kommunikations- und Medienforschung

Marko Bachl

Freie Universität Berlin

Fragen zur Hausaufgabe?

Compuational Text Analysis

Basiert auf Bachl & Scharkow (2024)

Agenda

  1. Was ist Computational Text Analysis?
  2. Historische Entwicklung
  3. Text as data
  4. Typologie: Unsupervised vs. Supervised
  5. Fazit

Was ist Computational Text Analysis?

Was ist Computational Text Analysis?

  • Definition: Methoden zur Messung von Texteigenschaften oder -inhalten mithilfe von Computeralgorithmen.
  • Häufig synonym oder ähnlich verwendete Begriffe: Text-as-data, Automated Content Analysis, Natural Language Processing (NLP), Text Mining
  • Kleinster gemeinsamer Nenner:
    • Nach der Konfiguration führt das System die Messung unabhängig und ohne manuellen Eingriff durch.
    • Dadurch skalieren die Methoden gut. Sie sind für die Analyse großer Textmengen geeignet.


  • Überblick über Grundbegriffe, Konzepte und einige Verfahren aus der Vogelperspektive mit einigen Beispielen in R. Weder vollständig noch nuanciert.
  • Hinweise und Links zum Weitermachen in den Folien. Einzelne Verfahren detaillierter im nächsten Semester.

Historische Entwicklung

Die Anfänge (1950er–1960er)

  • Computer als neues Werkzeug für Sozialwissenschaften
  • Textstatistiken (z.B. Wortzählungen)
  • General Inquirer: Prototyp für Textanalyse-Software, u.a. mit Preprocessing-Routinen und wörterbuchbasierter Klassifikation
  • Aber: Hohe Kosten (Rechenzeit teurer als manuelle Arbeit), aufwendige Digitalisierung (Lochkarten)
  • Nachlassendes Interesse nach anfänglichem Enthusiasmus

Erste Welle der Digitalisierung (1970er–1980er)

  • Zugang zu Medieninhalten als Nebenprodukt der Digitalisierung der Medienproduktion
  • Erste US-amerikanische Zeitungen in LexisNexis
  • Mehr Anwendungen in Kommunikations- und Politikwissenschaft
  • Computer noch immer zu wenig leistungsfähig, um Texte zu “verstehen”

Internet und PCs (1990er–2000er)

  • Internet und allgemeine Digitalisierung sorgen für immer mehr computerlesbare Medien- und Kommunikationsinhalte
  • PCs allgemein verbreitet rechenstärker, benutzungsfreundlicher
  • Statistische Analyse großer Textmengen möglich: “text as data”, maschinelles Lernen auf Basis von “Bag of Words” (siehe nächster Abschnitt)
  • Zunehmende Verbreitung in den Sozialwissenschaften (z.B. Grimmer & Stewart, 2013; Hase et al., 2023)

(Large) Language Models (2010er–heute)

  • Deep Learning:
    • Embeddings (Kontextwissen)
    • Transfer Learning (Übernahme des Kontextwissens beim Lernen neuer Aufgaben)
    • Zero-shot classification (Neue Aufgaben ohne Training)
  • Verfügbarkeit und einfache Nutzbarkeit solcher Modelle

Fragen?

Text as data

Text as data

Texte müssen in irgendeiner Form in Zahlen umgewandelt werden, damit Computer etwas berechnen können


  1. Bag of Words (BoW; “Sack mit Wörtern”)
  2. (Semantische Netzwerke; Konzept sehr spannend, hat sich nie durchgesetzt (z.B. Atteveldt, 2008))
  3. Embeddings (“Einbettungen” [sagt aber niemand])

Bag of Words (BoW)

  • Text als ungeordnete Sammlung (“Sack”) von Wörtern.
  • Grammatik und Wortreihenfolge werden ignoriert; nur die Häufigkeit zählt.
  • Preprocessing (Vorbereitung) der Texte notwendig, u.a.:
    • Tokenisierung
    • Stop-Wörter entfernen
    • Stemming / Lemmatisierung
  • Document-Term-Matrix (DTM), auch Document-Feature-Matrix (DFM)
  • Entscheidungen beim Preprocessing beeinflussen Ergebnisse (Denny & Spirling, 2018; Pipal et al., 2023)
  • Ungeeignet für kurze und nicht standardsprachliche Texte; Social Media!

Bag of Words (BoW): Beispiel

Populäre R Pakete für BoW-basierte Computational Text Analysis

Bag of Words (BoW): Beispiel

Inaugural Addresses der US-Präsidenten

data_corpus_inaugural <- data_corpus_inaugural |> rev.default() # Umdrehen, damit aktuellste Reden zuerst kommen
data_corpus_inaugural |>
  _[1] |>
  as.character() |>
  str_sub(end = 3000) |>
  str_wrap(width = 80) |>
  cat()

Bag of Words (BoW): Beispiel

Thank you. Thank you very much, everybody. Wow. Thank you very, very much.
Vice President Vance, Speaker Johnson, Senator Thune, Chief Justice Roberts,
justices of the Supreme Court of the United States, President Clinton, President
Bush, President Obama, President Biden, Vice President Harris, and my fellow
citizens, the golden age of America begins right now. From this day forward, our
country will flourish and be respected again all over the world. We will be the
envy of every nation, and we will not allow ourselves to be taken advantage of
any longer. During every single day of the Trump administration, I will, very
simply, put America first. Our sovereignty will be reclaimed. Our safety will
be restored. The scales of justice will be rebalanced. The vicious, violent,
and unfair weaponization of the Justice Department and our government will end.
And our top priority will be to create a nation that is proud, prosperous, and
free. America will soon be greater, stronger, and far more exceptional than ever
before. I return to the presidency confident and optimistic that we are at the
start of a thrilling new era of national success. A tide of change is sweeping
the country, sunlight is pouring over the entire world, and America has the
chance to seize this opportunity like never before. But first, we must be honest
about the challenges we face. While they are plentiful, they will be annihilated
by this great momentum that the world is now witnessing in the United States
of America. As we gather today, our government confronts a crisis of trust. For
many years, a radical and corrupt establishment has extracted power and wealth
from our citizens while the pillars of our society lay broken and seemingly in
complete disrepair. We now have a government that cannot manage even a simple
crisis at home while, at the same time, stumbling into a continuing catalogue
of catastrophic events abroad. It fails to protect our magnificent, law-abiding
American citizens but provides sanctuary and protection for dangerous criminals,
many from prisons and mental institutions, that have illegally entered our
country from all over the world. We have a government that has given unlimited
funding to the defense of foreign borders but refuses to defend American borders
or, more importantly, its own people. Our country can no longer deliver basic
services in times of emergency, as recently shown by the wonderful people of
North Carolina — who have been treated so badly — (applause) — and other states
who are still suffering from a hurricane that took place many months ago or,
more recently, Los Angeles, where we are watching fires still tragically burn
from weeks ago without even a token of defense. They’re raging through the
houses and communities, even affecting some of the wealthiest and most powerful
individuals in our country — some of whom are sitting here right now. They don’t
have a home any longer. That’s interesting. But we can’t let

Bag of Words (BoW): Beispiel

Tokenisierung

data_corpus_inaugural |>
  _[1] |>
  tokens(
    what = "word",
    remove_punct = TRUE,
    remove_symbols = FALSE,
    remove_numbers = TRUE,
    remove_url = FALSE,
    remove_separators = TRUE,
    split_hyphens = FALSE,
    split_tags = FALSE
  ) |>
  as.character() |>
  head(100)

Bag of Words (BoW): Beispiel

  [1] "Thank"          "you"            "Thank"          "you"           
  [5] "very"           "much"           "everybody"      "Wow"           
  [9] "Thank"          "you"            "very"           "very"          
 [13] "much"           "Vice"           "President"      "Vance"         
 [17] "Speaker"        "Johnson"        "Senator"        "Thune"         
 [21] "Chief"          "Justice"        "Roberts"        "justices"      
 [25] "of"             "the"            "Supreme"        "Court"         
 [29] "of"             "the"            "United"         "States"        
 [33] "President"      "Clinton"        "President"      "Bush"          
 [37] "President"      "Obama"          "President"      "Biden"         
 [41] "Vice"           "President"      "Harris"         "and"           
 [45] "my"             "fellow"         "citizens"       "the"           
 [49] "golden"         "age"            "of"             "America"       
 [53] "begins"         "right"          "now"            "From"          
 [57] "this"           "day"            "forward"        "our"           
 [61] "country"        "will"           "flourish"       "and"           
 [65] "be"             "respected"      "again"          "all"           
 [69] "over"           "the"            "world"          "We"            
 [73] "will"           "be"             "the"            "envy"          
 [77] "of"             "every"          "nation"         "and"           
 [81] "we"             "will"           "not"            "allow"         
 [85] "ourselves"      "to"             "be"             "taken"         
 [89] "advantage"      "of"             "any"            "longer"        
 [93] "During"         "every"          "single"         "day"           
 [97] "of"             "the"            "Trump"          "administration"

Bag of Words (BoW): Beispiel

Kleinschreibung

data_corpus_inaugural |>
  _[1] |>
  tokens(
    what = "word",
    remove_punct = TRUE,
    remove_symbols = FALSE,
    remove_numbers = TRUE,
    remove_url = FALSE,
    remove_separators = TRUE,
    split_hyphens = FALSE,
    split_tags = FALSE
  ) |>
  tokens_tolower() |>
  as.character() |>
  head(100)

Bag of Words (BoW): Beispiel

  [1] "thank"          "you"            "thank"          "you"           
  [5] "very"           "much"           "everybody"      "wow"           
  [9] "thank"          "you"            "very"           "very"          
 [13] "much"           "vice"           "president"      "vance"         
 [17] "speaker"        "johnson"        "senator"        "thune"         
 [21] "chief"          "justice"        "roberts"        "justices"      
 [25] "of"             "the"            "supreme"        "court"         
 [29] "of"             "the"            "united"         "states"        
 [33] "president"      "clinton"        "president"      "bush"          
 [37] "president"      "obama"          "president"      "biden"         
 [41] "vice"           "president"      "harris"         "and"           
 [45] "my"             "fellow"         "citizens"       "the"           
 [49] "golden"         "age"            "of"             "america"       
 [53] "begins"         "right"          "now"            "from"          
 [57] "this"           "day"            "forward"        "our"           
 [61] "country"        "will"           "flourish"       "and"           
 [65] "be"             "respected"      "again"          "all"           
 [69] "over"           "the"            "world"          "we"            
 [73] "will"           "be"             "the"            "envy"          
 [77] "of"             "every"          "nation"         "and"           
 [81] "we"             "will"           "not"            "allow"         
 [85] "ourselves"      "to"             "be"             "taken"         
 [89] "advantage"      "of"             "any"            "longer"        
 [93] "during"         "every"          "single"         "day"           
 [97] "of"             "the"            "trump"          "administration"

Bag of Words (BoW): Beispiel

Häufige Wörter (“stopwords”) entfernen

stopwords()
  [1] "i"          "me"         "my"         "myself"     "we"        
  [6] "our"        "ours"       "ourselves"  "you"        "your"      
 [11] "yours"      "yourself"   "yourselves" "he"         "him"       
 [16] "his"        "himself"    "she"        "her"        "hers"      
 [21] "herself"    "it"         "its"        "itself"     "they"      
 [26] "them"       "their"      "theirs"     "themselves" "what"      
 [31] "which"      "who"        "whom"       "this"       "that"      
 [36] "these"      "those"      "am"         "is"         "are"       
 [41] "was"        "were"       "be"         "been"       "being"     
 [46] "have"       "has"        "had"        "having"     "do"        
 [51] "does"       "did"        "doing"      "would"      "should"    
 [56] "could"      "ought"      "i'm"        "you're"     "he's"      
 [61] "she's"      "it's"       "we're"      "they're"    "i've"      
 [66] "you've"     "we've"      "they've"    "i'd"        "you'd"     
 [71] "he'd"       "she'd"      "we'd"       "they'd"     "i'll"      
 [76] "you'll"     "he'll"      "she'll"     "we'll"      "they'll"   
 [81] "isn't"      "aren't"     "wasn't"     "weren't"    "hasn't"    
 [86] "haven't"    "hadn't"     "doesn't"    "don't"      "didn't"    
 [91] "won't"      "wouldn't"   "shan't"     "shouldn't"  "can't"     
 [96] "cannot"     "couldn't"   "mustn't"    "let's"      "that's"    
[101] "who's"      "what's"     "here's"     "there's"    "when's"    
[106] "where's"    "why's"      "how's"      "a"          "an"        
[111] "the"        "and"        "but"        "if"         "or"        
[116] "because"    "as"         "until"      "while"      "of"        
[121] "at"         "by"         "for"        "with"       "about"     
[126] "against"    "between"    "into"       "through"    "during"    
[131] "before"     "after"      "above"      "below"      "to"        
[136] "from"       "up"         "down"       "in"         "out"       
[141] "on"         "off"        "over"       "under"      "again"     
[146] "further"    "then"       "once"       "here"       "there"     
[151] "when"       "where"      "why"        "how"        "all"       
[156] "any"        "both"       "each"       "few"        "more"      
[161] "most"       "other"      "some"       "such"       "no"        
[166] "nor"        "not"        "only"       "own"        "same"      
[171] "so"         "than"       "too"        "very"       "will"      

Bag of Words (BoW): Beispiel

Häufige Wörter (“stopwords”) entfernen

data_corpus_inaugural |>
  _[1] |>
  tokens(
    what = "word",
    remove_punct = TRUE,
    remove_symbols = FALSE,
    remove_numbers = TRUE,
    remove_url = FALSE,
    remove_separators = TRUE,
    split_hyphens = FALSE,
    split_tags = FALSE
  ) |>
  tokens_tolower() |>
  tokens_remove(stopwords()) |>
  as.character() |>
  head(100)

Bag of Words (BoW): Beispiel

  [1] "thank"          "thank"          "much"           "everybody"     
  [5] "wow"            "thank"          "much"           "vice"          
  [9] "president"      "vance"          "speaker"        "johnson"       
 [13] "senator"        "thune"          "chief"          "justice"       
 [17] "roberts"        "justices"       "supreme"        "court"         
 [21] "united"         "states"         "president"      "clinton"       
 [25] "president"      "bush"           "president"      "obama"         
 [29] "president"      "biden"          "vice"           "president"     
 [33] "harris"         "fellow"         "citizens"       "golden"        
 [37] "age"            "america"        "begins"         "right"         
 [41] "now"            "day"            "forward"        "country"       
 [45] "flourish"       "respected"      "world"          "envy"          
 [49] "every"          "nation"         "allow"          "taken"         
 [53] "advantage"      "longer"         "every"          "single"        
 [57] "day"            "trump"          "administration" "simply"        
 [61] "put"            "america"        "first"          "sovereignty"   
 [65] "reclaimed"      "safety"         "restored"       "scales"        
 [69] "justice"        "rebalanced"     "vicious"        "violent"       
 [73] "unfair"         "weaponization"  "justice"        "department"    
 [77] "government"     "end"            "top"            "priority"      
 [81] "create"         "nation"         "proud"          "prosperous"    
 [85] "free"           "america"        "soon"           "greater"       
 [89] "stronger"       "far"            "exceptional"    "ever"          
 [93] "return"         "presidency"     "confident"      "optimistic"    
 [97] "start"          "thrilling"      "new"            "era"           

Bag of Words (BoW): Beispiel

Auf Wortstämme reduzieren (“stemming” oder “lemmatization”)

data_corpus_inaugural |>
  _[1] |>
  tokens(
    what = "word",
    remove_punct = TRUE,
    remove_symbols = FALSE,
    remove_numbers = TRUE,
    remove_url = FALSE,
    remove_separators = TRUE,
    split_hyphens = FALSE,
    split_tags = FALSE
  ) |>
  tokens_tolower() |>
  tokens_remove(stopwords()) |>
  tokens_wordstem() |>
  as.character() |>
  head(100)

Bag of Words (BoW): Beispiel

  [1] "thank"       "thank"       "much"        "everybodi"   "wow"        
  [6] "thank"       "much"        "vice"        "presid"      "vanc"       
 [11] "speaker"     "johnson"     "senat"       "thune"       "chief"      
 [16] "justic"      "robert"      "justic"      "suprem"      "court"      
 [21] "unit"        "state"       "presid"      "clinton"     "presid"     
 [26] "bush"        "presid"      "obama"       "presid"      "biden"      
 [31] "vice"        "presid"      "harri"       "fellow"      "citizen"    
 [36] "golden"      "age"         "america"     "begin"       "right"      
 [41] "now"         "day"         "forward"     "countri"     "flourish"   
 [46] "respect"     "world"       "envi"        "everi"       "nation"     
 [51] "allow"       "taken"       "advantag"    "longer"      "everi"      
 [56] "singl"       "day"         "trump"       "administr"   "simpli"     
 [61] "put"         "america"     "first"       "sovereignti" "reclaim"    
 [66] "safeti"      "restor"      "scale"       "justic"      "rebalanc"   
 [71] "vicious"     "violent"     "unfair"      "weapon"      "justic"     
 [76] "depart"      "govern"      "end"         "top"         "prioriti"   
 [81] "creat"       "nation"      "proud"       "prosper"     "free"       
 [86] "america"     "soon"        "greater"     "stronger"    "far"        
 [91] "except"      "ever"        "return"      "presid"      "confid"     
 [96] "optimist"    "start"       "thrill"      "new"         "era"        

Bag of Words (BoW): Beispiel

Document-Term-Matrix (Document-Feature-Matrix in {quanteda})

data_corpus_inaugural |>
  tokens(
    what = "word",
    remove_punct = TRUE,
    remove_symbols = FALSE,
    remove_numbers = TRUE,
    remove_url = FALSE,
    remove_separators = TRUE,
    split_hyphens = FALSE,
    split_tags = FALSE
  ) |>
  tokens_tolower() |>
  tokens_remove(stopwords()) |>
  tokens_wordstem() |>
  dfm() |>
  print(max_ndoc = 20, max_nfeat = 9)

Bag of Words (BoW): Beispiel

Document-feature matrix of: 60 documents, 5,478 features (89.35% sparse) and 4 docvars.
                 features
docs              thank much everybodi wow vice presid vanc speaker johnson
  2025-Trump         23    6         1   1    2     10    1       1       1
  2021-Biden          3    8         0   0    3      7    0       1       0
  2017-Trump          3    1         0   0    0      5    0       0       0
  2013-Obama          1    0         0   0    1      2    0       0       0
  2009-Obama          2    1         0   0    0      1    0       0       0
  2005-Bush           0    0         0   0    1      4    0       0       0
  2001-Bush           2    2         0   0    1      3    0       0       0
  1997-Clinton        0    1         0   0    0      1    0       0       0
  1993-Clinton        2    2         0   0    0      3    0       0       0
  1989-Bush           5    3         0   0    1      7    0       3       0
  1985-Reagan         1    4         0   0    1      3    0       1       0
  1981-Reagan         2    4         0   0    2      6    0       1       0
  1977-Carter         1    0         0   0    0      3    0       0       0
  1973-Nixon          0    1         0   0    1      1    0       1       0
  1969-Nixon          1    0         0   0    2      3    0       0       1
  1965-Johnson        0    1         0   0    0      0    0       0       0
  1961-Kennedy        0    1         0   0    2      4    0       1       1
  1957-Eisenhower     0    2         0   0    1      1    0       1       0
  1953-Eisenhower     0    0         0   0    0      0    0       0       0
  1949-Truman         0    0         0   0    1      1    0       0       0
[ reached max_ndoc ... 40 more documents, reached max_nfeat ... 5,469 more features ]

Fragen?

Embeddings

  • Wörter werden als dichte Vektoren in einem hochdimensionalen Raum repräsentiert.
  • Distributional Hypothesis: Wörter, die in großen Textmengen in ähnlichen Kontexten stehen, haben ähnliche Bedeutungen.
  • Training (einfacher Word) Embeddings: Vorhersage eines Wortes durch die das Wort umgebenden Wörter.
  • Typen:
    • Statisch: z.B. word2vec, GloVe (ein Vektor pro Wort).
    • Kontextualisiert: z.B. BERT, GPT (Vektor passt sich dem Kontext an).
  • Kann für einzelne Wörter oder längere Texte (z.B. Sätze) bestimmt werden.
  • Machen Preprocessing weitgehend unnötig. Sind für kürzere und nicht standardsprachliche Texte geeignet.
  • Vortrainierte Embeddings (mit Biases der Trainingsdaten) oder eigene Embeddings (erfodern sehr große Textmengen)

Embeddings: Beispiel

R Pakete für Embedding-basierte Computational Text Analysis

Embeddings: Beispiel

100 häufigste Wörter

top100 <- data_corpus_inaugural |>
  tokens(
    what = "word",
    remove_punct = TRUE,
    remove_symbols = FALSE,
    remove_numbers = TRUE,
    remove_url = FALSE,
    remove_separators = TRUE,
    split_hyphens = FALSE,
    split_tags = FALSE
  ) |>
  tokens_remove(stopwords()) |>
  dfm() |>
  topfeatures(n = 100) |>
  names()

Embeddings: Beispiel

Word Embeddings

top100_emb <- top100 |>
  embed_text(model = "qwen3-embedding:0.6b") |>
  mutate(feature = top100) |>
  relocate(feature)
top100_emb |>
  mutate(across(where(is.numeric), \(x) round(x, digits = 3))) |>
  print(n = 50)

Embeddings: Beispiel

# A tibble: 100 × 1,025
   feature  dim_1  dim_2  dim_3  dim_4  dim_5  dim_6  dim_7  dim_8  dim_9 dim_10
   <chr>    <dbl>  <dbl>  <dbl>  <dbl>  <dbl>  <dbl>  <dbl>  <dbl>  <dbl>  <dbl>
 1 people  -0.045 -0.022 -0.013 -0.007  0.036 -0.068  0.047 -0.034 -0.027  0.024
 2 govern… -0.024 -0.062 -0.011 -0.022  0.053 -0.112 -0.009  0.034 -0.021  0.049
 3 us       0.009 -0.037 -0.014 -0.066  0.053  0.004  0.001 -0.003 -0.037  0.006
 4 can     -0.006 -0.065 -0.014 -0.062  0.025 -0.009  0.036  0.015 -0.007  0.045
 5 must    -0.029 -0.052 -0.013 -0.058  0.039 -0.013  0.051  0.003 -0.02   0.058
 6 upon     0.009 -0.062 -0.014 -0.055  0.038 -0.004  0.026  0.03  -0.035  0.014
 7 great   -0.022 -0.029 -0.014  0.017  0.055  0.009 -0.029  0.049 -0.017  0.025
 8 states   0.006 -0.067 -0.014 -0.067  0.039  0.041  0.039 -0.02  -0.005  0.04 
 9 may     -0.057 -0.009 -0.014 -0.045  0.043 -0.011  0.035  0.034 -0.001  0.013
10 world   -0.015  0.017 -0.012 -0.024  0.057 -0.058  0.009  0.024 -0.067 -0.014
11 nation   0.009 -0.033 -0.013 -0.022  0.065 -0.05  -0.022  0.027 -0.03   0.014
12 country -0.017  0     -0.012 -0.008  0.084 -0.069 -0.013 -0.005 -0.007  0.032
13 shall   -0.002 -0.074 -0.014 -0.043  0.003 -0.028  0.021  0.033 -0.058  0    
14 every   -0.038 -0.05  -0.015 -0.046  0.041 -0.035  0.043  0.035 -0.028 -0.012
15 one      0.007 -0.068 -0.017 -0.057  0.028  0.035  0.006  0.001 -0.039  0.019
16 peace   -0.007 -0.003 -0.013  0.016  0.026 -0.053 -0.019  0.035 -0.018  0.034
17 new     -0.017 -0.054 -0.016 -0.056  0.035 -0.031  0.01   0.083 -0.058 -0.001
18 power    0.015 -0.036 -0.013 -0.004  0.05   0.031  0.02  -0.018 -0.021  0.062
19 now     -0.012 -0.065 -0.015 -0.059  0.05  -0.018  0.029  0.041 -0.046  0.021
20 public  -0.024 -0.071 -0.012 -0.035  0.023 -0.083 -0.005  0.01  -0.031  0.058
21 time    -0.048 -0.013 -0.014 -0.064  0.067  0.023  0.009  0.028 -0.044 -0.007
22 america -0.011  0     -0.012 -0.038  0.039 -0.023 -0.012 -0.004 -0.016  0.018
23 citize… -0.008 -0.043 -0.013 -0.007  0.055 -0.126  0.066  0.003 -0.019  0.017
24 united   0.009 -0.025 -0.014 -0.039  0.053 -0.029  0.004  0.008 -0.055  0.022
25 consti…  0.02  -0.091 -0.013 -0.044  0.034 -0.021 -0.01   0.012 -0.042  0.018
26 nations  0.006 -0.036 -0.013 -0.06   0.041 -0.07   0.018  0.02  -0.039  0.006
27 union    0.009 -0.063 -0.016 -0.062  0.054  0.015  0.008  0.022 -0.026  0.022
28 freedom  0.005 -0.043 -0.013 -0.029  0.053 -0.106 -0.015  0.056 -0.024  0.02 
29 free    -0.02  -0.042 -0.015 -0.086  0.062 -0.052 -0.021  0.072 -0.04   0.042
30 americ… -0.01  -0.009 -0.015 -0.019  0.051 -0.04  -0.033  0.015  0.007  0.016
31 war      0.029 -0.05  -0.012  0.003  0.053 -0.056 -0.03   0.019 -0.047  0.007
32 nation…  0.003 -0.052 -0.015 -0.009  0.056 -0.062 -0.026  0.038 -0.025  0.02 
33 let     -0.004 -0.046 -0.013 -0.069 -0.015  0.012 -0.029  0.058 -0.05   0.031
34 made    -0.034 -0.044 -0.014 -0.07   0.018 -0.012  0.048  0.022 -0.031  0.064
35 years   -0.025  0.004 -0.015 -0.009  0.078 -0.041  0.057  0.022 -0.051 -0.02 
36 make    -0.032 -0.041 -0.014 -0.08   0.031 -0.046  0.039  0.037 -0.055  0.054
37 good    -0.021 -0.045 -0.015 -0.01   0.038 -0.006 -0.039  0.065 -0.008  0.066
38 justice  0.017 -0.047 -0.012  0      0.034 -0.069 -0.009  0.062 -0.015 -0.021
39 spirit   0.011  0.035 -0.011  0.002  0.049 -0.025  0.015  0.013 -0.021  0.019
40 never   -0.018 -0.055 -0.015 -0.051  0.029 -0.012 -0.01   0.05  -0.047  0.019
41 without -0.026 -0.064 -0.016 -0.077  0.045 -0.019  0.001  0.034 -0.047  0.027
42 life    -0.045  0.007 -0.013  0.004  0.047 -0.028  0.02   0.033 -0.028  0.021
43 men     -0.023 -0.049 -0.015 -0.014  0.009 -0.057  0.038 -0.015 -0.014  0.013
44 rights  -0.009 -0.021 -0.015 -0.075  0.037 -0.098 -0.029  0.047 -0.031  0.026
45 law     -0.014 -0.047 -0.011 -0.016  0.024 -0.078 -0.019  0.074 -0.009 -0.011
46 just     0.031 -0.042 -0.012 -0.038  0.025  0.017  0.02   0.043 -0.04   0.036
47 congre… -0.019 -0.068 -0.011 -0.008  0.035 -0.051 -0.009  0.003 -0.044  0.034
48 laws    -0.023 -0.037 -0.013 -0.024  0.024 -0.076 -0.02   0.067 -0.013 -0.015
49 right   -0.015 -0.032 -0.014 -0.015  0.041 -0.015 -0.038  0.056 -0.026  0.027
50 work    -0.015 -0.048 -0.012 -0.046  0.037 -0.03   0.051  0.03  -0.008  0.049
# ℹ 50 more rows
# ℹ 1,014 more variables: dim_11 <dbl>, dim_12 <dbl>, dim_13 <dbl>,
#   dim_14 <dbl>, dim_15 <dbl>, dim_16 <dbl>, dim_17 <dbl>, dim_18 <dbl>,
#   dim_19 <dbl>, dim_20 <dbl>, dim_21 <dbl>, dim_22 <dbl>, dim_23 <dbl>,
#   dim_24 <dbl>, dim_25 <dbl>, dim_26 <dbl>, dim_27 <dbl>, dim_28 <dbl>,
#   dim_29 <dbl>, dim_30 <dbl>, dim_31 <dbl>, dim_32 <dbl>, dim_33 <dbl>,
#   dim_34 <dbl>, dim_35 <dbl>, dim_36 <dbl>, dim_37 <dbl>, dim_38 <dbl>, …

Embeddings: Beispiel

Word Embeddings: Ähnlichkeit

Cosine similarity:

  • Wertebereich [\(-1;+1\)], entgegengesetzt bis identisch
  • Hier: Distanz im 1024-dimensionalen Vektorraum
top100_sim_matrix <- simil(as.matrix(select(top100_emb, -feature)), method = "cosine")
dimnames(top100_sim_matrix) <- list(top100, top100)
top100_sim_matrix |>
  as.matrix() |>
  data.frame() |>
  rownames_to_column(var = "feature") |>
  as_tibble() |>
  gather("feature2", "similarity", -feature) |>
  filter(feature > feature2) |>
  arrange(desc(similarity)) |>
  print(n = 30)

Embeddings: Beispiel

# A tibble: 4,950 × 3
   feature      feature2 similarity
   <chr>        <chr>         <dbl>
 1 laws         law           0.970
 2 states       state         0.964
 3 national     nation        0.951
 4 nation       country       0.933
 5 men          man           0.924
 6 duty         duties        0.906
 7 people       men           0.906
 8 national     country       0.897
 9 americans    america       0.895
10 political    policy        0.892
11 nation       america       0.892
12 law          justice       0.891
13 country      america       0.890
14 one          first         0.884
15 without      within        0.877
16 nations      nation        0.877
17 without      now           0.877
18 without      upon          0.872
19 good         best          0.871
20 constitution congress      0.870
21 americans    american      0.869
22 now          new           0.867
23 today        day           0.867
24 much         many          0.865
25 laws         justice       0.865
26 make         made          0.864
27 now          never         0.863
28 government   congress      0.861
29 still        now           0.861
30 never        ever          0.860
# ℹ 4,920 more rows

Embeddings: Beispiel

Sentence Embeddings (300 zufällig ausgewählte Sätze)

set.seed(1)
sentence_sample <- data_corpus_inaugural |>
  tokens(what = "sentence") |>
  as.character() |>
  sample(300) |>
  unique()
sentence_emb <- sentence_sample |>
  embed_text(model = "qwen3-embedding:0.6b") |>
  mutate(sentence = sentence_sample) |>
  relocate(sentence)
print(sentence_emb, n = 50)

Embeddings: Beispiel

# A tibble: 300 × 1,025
   sentence          dim_1    dim_2    dim_3    dim_4    dim_5    dim_6    dim_7
   <chr>             <dbl>    <dbl>    <dbl>    <dbl>    <dbl>    <dbl>    <dbl>
 1 "To renew Ame… -2.25e-2 -0.0336  -0.0105   2.37e-2  0.0603  -1.12e-1  9.42e-3
 2 "Actual event…  1.67e-2 -0.0587  -0.00838  6.88e-2  0.0559  -9.04e-2 -3.42e-2
 3 "We can gain … -1.17e-2 -0.0244  -0.00678  4.06e-2 -0.0129  -3.32e-2 -2.70e-2
 4 "As early as … -1.45e-2  0.0235  -0.00560  4.22e-2 -0.0105  -8.61e-2  3.67e-2
 5 "Just as Amer… -4.32e-3 -0.0319  -0.00575 -3.59e-3  0.00715 -8.32e-2  3.65e-5
 6 "If this is t…  3.53e-2 -0.0804  -0.0102  -2.30e-3 -0.0433  -1.20e-1  4.58e-2
 7 "I am certain…  1.18e-2 -0.0690  -0.00791  1.57e-2 -0.0134  -1.07e-1  1.70e-2
 8 "And, I belie… -3.41e-2 -0.0132  -0.00555  1.30e-2  0.00736 -2.51e-2 -5.87e-2
 9 "It is safe t… -4.17e-2 -0.0704  -0.00512  1.63e-3  0.0233  -6.04e-2 -4.09e-2
10 "Friends and …  3.86e-2 -0.0758  -0.00910 -4.15e-3 -0.0486  -1.01e-1  4.90e-2
11 "The Governme… -2.02e-2 -0.0154  -0.00691 -2.68e-2  0.0515  -6.24e-2 -1.96e-2
12 "In doing thi…  1.89e-2 -0.100   -0.00635 -5.79e-3  0.0120  -7.62e-2 -1.32e-2
13 "Rather, it h… -1.19e-2 -0.0370  -0.0129   5.96e-2 -0.00316 -1.13e-1 -1.66e-2
14 "I have appro…  3.14e-2 -0.0557  -0.00991 -1.11e-2  0.0578  -3.22e-2 -2.34e-2
15 "We will repa…  7.85e-2 -0.0361  -0.0111  -3.51e-2  0.0818  -9.00e-2  3.13e-3
16 "Finally, to …  3.59e-2 -0.00509 -0.0110   1.55e-3  0.00858 -9.71e-2 -9.65e-3
17 "My best effo… -1.18e-2 -0.0647  -0.0116  -4.49e-3  0.0310   6.54e-4  5.29e-2
18 "And we have …  8.67e-3 -0.0619  -0.00824 -2.13e-2  0.0371  -6.51e-2 -3.48e-2
19 "And we will … -1.34e-2 -0.0580  -0.00585 -1.78e-2  0.0334  -7.92e-2 -1.78e-2
20 "And in any c… -2.45e-2 -0.0793  -0.00348 -3.17e-2  0.0811  -8.45e-3 -4.05e-2
21 "Today I also…  2.14e-2 -0.0480  -0.00588 -2.26e-2  0.0523  -1.41e-1  9.61e-2
22 "We must keep… -1.12e-2 -0.0462  -0.00956 -1.11e-4  0.0505  -8.39e-2  4.17e-2
23 "What makes u…  5.86e-3 -0.0482  -0.00943  1.50e-2  0.0267  -9.91e-2  9.89e-4
24 "The prayers …  2.14e-2 -0.0116  -0.00460 -2.24e-2  0.0174   1.77e-4  3.41e-2
25 "The whole sy…  5.67e-4 -0.139   -0.0122   5.36e-3 -0.00749 -7.54e-2  4.23e-2
26 "War never le… -2.68e-2 -0.0443  -0.00852  1.92e-2  0.0485  -9.41e-2 -2.44e-2
27 "I would like… -7.08e-2  0.0102  -0.0102   4.41e-2  0.0500  -1.11e-1  1.65e-2
28 "Small wonder… -8.62e-3 -0.0385  -0.00951  2.99e-2  0.0197  -8.60e-2  4.91e-2
29 "We are told …  3.23e-2 -0.0597  -0.00792  8.33e-2  0.0175  -9.01e-2 -4.42e-2
30 "Whoever woul… -5.41e-2 -0.00349 -0.00527  4.77e-2 -0.0681  -1.14e-2 -3.52e-2
31 "If we meet t…  1.49e-2 -0.0626  -0.0102   1.43e-2 -0.00148 -8.34e-2  6.12e-3
32 "The North an… -2.56e-2 -0.0886  -0.00970  3.97e-2  0.0424  -8.56e-2  1.74e-2
33 "This trial c…  2.23e-2 -0.0138  -0.00806  1.32e-1  0.0343  -3.52e-2 -1.37e-2
34 "Ours is a la… -5.39e-2 -0.0437  -0.00798  1.44e-2  0.0395  -8.75e-2  3.89e-2
35 "But we shall…  3.04e-3 -0.0521  -0.00774 -3.27e-3  0.0128  -1.83e-2  4.81e-2
36 "Experience h…  6.17e-3 -0.102   -0.00788  5.72e-2 -0.0193  -9.25e-2  6.29e-2
37 "In our prese…  1.59e-2 -0.0554  -0.00385 -9.71e-3  0.0357  -3.56e-2  9.67e-3
38 "Now the very… -4.03e-2 -0.116   -0.00688 -2.93e-3  0.00555  1.12e-2  3.17e-2
39 "Wise counsel…  1.13e-4 -0.0791  -0.00624  2.66e-2 -0.0249  -6.19e-2  5.10e-2
40 "So much has … -7.68e-2 -0.0377  -0.0110   3.19e-2  0.0662  -3.03e-2  1.97e-2
41 "They have as…  4.38e-2 -0.0985  -0.00940  7.07e-3  0.0153  -4.62e-2  4.15e-2
42 "We have beat…  7.35e-2 -0.0288  -0.0126   5.14e-2  0.0444  -3.66e-2 -7.72e-3
43 "Fortunately … -5.41e-2 -0.0927  -0.00480  3.58e-2  0.0426  -1.31e-2  3.39e-2
44 "We contempla… -1.86e-2 -0.0711  -0.01000 -1.17e-3  0.0277  -7.34e-2  6.52e-2
45 "They made a …  1.10e-2 -0.0527  -0.00426 -3.91e-2  0.00186 -9.29e-2  3.01e-2
46 "Some may sti… -5.47e-2  0.0630  -0.00754  8.45e-2  0.0419  -6.70e-2  4.23e-2
47 "President Re… -3.50e-2  0.0179  -0.00685 -4.03e-2  0.00825 -4.26e-2 -1.05e-2
48 "It fails to …  9.40e-3 -0.0660  -0.00717  4.52e-2  0.0375  -1.52e-1 -3.48e-2
49 "But we have … -2.42e-2 -0.0749  -0.00827  1.92e-2  0.0213   7.21e-2  4.12e-4
50 "A bridge wid…  1.20e-2 -0.0154  -0.00821  1.84e-2  0.0260  -8.72e-2  6.55e-2
# ℹ 250 more rows
# ℹ 1,017 more variables: dim_8 <dbl>, dim_9 <dbl>, dim_10 <dbl>, dim_11 <dbl>,
#   dim_12 <dbl>, dim_13 <dbl>, dim_14 <dbl>, dim_15 <dbl>, dim_16 <dbl>,
#   dim_17 <dbl>, dim_18 <dbl>, dim_19 <dbl>, dim_20 <dbl>, dim_21 <dbl>,
#   dim_22 <dbl>, dim_23 <dbl>, dim_24 <dbl>, dim_25 <dbl>, dim_26 <dbl>,
#   dim_27 <dbl>, dim_28 <dbl>, dim_29 <dbl>, dim_30 <dbl>, dim_31 <dbl>,
#   dim_32 <dbl>, dim_33 <dbl>, dim_34 <dbl>, dim_35 <dbl>, dim_36 <dbl>, …

Embeddings: Beispiel

Sentence Embeddings: Ähnlichkeit

sentence_sim_matrix <- simil(as.matrix(select(sentence_emb, -sentence)), method = "cosine")
sentence_sim_matrix |>
  as.matrix() |>
  as_tibble() |>
  set_names(sentence_sample) |>
  mutate(sentence = sentence_sample) |>
  gather("sentence2", "similarity", -sentence) |>
  filter(sentence > sentence2) |>
  arrange(desc(similarity)) |>
  slice_head(n = 20) |>
  knitr::kable()

Embeddings: Beispiel

sentence sentence2 similarity
Just as America’s role is indispensable in preserving the world’s peace, so is each nation’s role indispensable in preserving its own peace. It is important that we understand both the necessity and the limitations of America’s role in maintaining that peace. 0.8076691
There is so much to be done. For everywhere we look, there is work to be done. 0.7877597
This faith is the abiding creed of our fathers. That’s what will lend meaning to the creed our fathers once declared. 0.7758326
Yet we endured and we prevailed. We have beaten back despair and defeatism. 0.7722788
We must act, knowing that our work will be imperfect. We must act on what we know. 0.7379628
It is our glory that whilst other nations have extended their dominions by the sword we have never acquired any territory except by fair purchase or, as in the case of Texas, by the voluntary determination of a brave, kindred, and independent people to blend their destinies with our own. Foreign powers should therefore look on the annexation of Texas to the United States not as the conquest of a nation seeking to extend her dominions by arms and violence, but as the peaceful acquisition of a territory once her own, by adding another member to our confederation, with the consent of that member, thereby diminishing the chances of war and opening to them new and ever-increasing markets for their products. 0.7312595
Now more than ever, we must do these things together, as one nation and one people. For any one of us to succeed, we must succeed as one America. 0.7287994
That’s America. And, I believe America is better than this. 0.7244906
To renew America we must be bold. A spring reborn in the world’s oldest democracy, that brings forth the vision and courage to reinvent America. 0.7093138
We will carry on. We sing it still. 0.7087671
We will repair our alliances and engage with the world once again. To the world, too, we offer new engagement and a renewed vow: We will stay strong to protect the peace. 0.7057480
I will fight for you with every breath in my body - and I will never, ever let you down. I will always level with you. 0.7056759
We must act, knowing that our work will be imperfect. For everywhere we look, there is work to be done. 0.7034697
To renew America we must be bold. And by our dreams and labors we will redeem the promise of America in the 21st century. 0.6984393
To renew America we must be bold. Above all, my message to Americans today is that it is time for us to once again act with courage, vigor, and the vitality of history’s greatest civilization. 0.6948836
This is our summons to greatness. Now we must step up. 0.6942487
That’s America. America stands alone as the world’s indispensable nation. 0.6926186
Yet we endured and we prevailed. We will carry on. 0.6907114
We have been carried in safety through a perilous crisis. Our crisis today is the reverse. 0.6886360
Texas was once a part of our country - was unwisely ceded away to a foreign power - is now independent, and possesses an undoubted right to dispose of a part or the whole of her territory and to merge her sovereignty as a separate and independent state in ours. It is our glory that whilst other nations have extended their dominions by the sword we have never acquired any territory except by fair purchase or, as in the case of Texas, by the voluntary determination of a brave, kindred, and independent people to blend their destinies with our own. 0.6858266

Fragen?

Typologie: Unsupervised vs. Supervised

Unsupervised vs. Supervised Methods

  • Unsupervised (nicht überwacht):
    • Meist explorative und deskriptive Forschungsfragen
    • Bei Klassifikationsaufgaben: Kategorien vor Analyse unbekannt
    • Relativ wenig Informationsinput durch Forschende
    • Typisches Verfahren: Topic Model
  • Supervised (überwacht):
    • Auch deduktive und inferentielle Forschungsziele
    • Bei Klassifikationsaufgaben: Kategorien vor Analyse bekannt
    • Relativ großer Informationsinput durch Forschende
    • Typisches Verfahren: Supervised machine learning classification
  • Heuristische Kategorien, nicht trennscharf: Nicht alle Verfahren lassen sich hinsichtlich jeder ihrer Eigenschaften eindeutig zuordnen

Unsupervised: Textstatistiken

  • BoW: z.B. Worthäufigkeiten, formale Verständlichkeitsindices
  • Embedding: z.B. Ähnlichkeit von Texten (siehe Einführung in Embeddings)

Verständlichkeit von Parteiprogrammen

(Brettschneider & Thoms, 2025)

Fragen?

Unsupervised: Topic Models

  • Identifikation von “Topics” in Dokumenten; Beschreibung durch typische Begriffe und typische Texte
  • Gut geeignet, um sehr große Dokumentmengen ohne große Vorkenntnisse zu struktureren
  • Sehr verbreitet in der Kommunikationswissenschaft: Deskriptiv nützlich, theoretisch umstritten (Chen et al., 2023)
  • BoW: Latent Dirichlet Allocation (LDA; z.B. Maier et al. (2018)) oder Strucural Topic Models mit dem R Paket {stm} (Roberts et al., 2019).
  • Embedding: Aktuell im Trend: Python Paket BERTopic (Grootendorst, 2022)

BERTopic

Topics

(Simon et al., 2025)

Fragen?

Supervised: Dictionaries (Wörterbucher)

  • Liste von Begriffen, die ein Konstrukt definieren
  • Vorteil: Einfach zu verstehen; Nachteil: Schwerer zu erstellen, als man vielleicht denken würde.
  • BoW: Exaktes Vorkommen eines Begriffs aus dem Wörterbuch → Konstrukt liegt vor (einfach durchzuführen z.B. in {quanteda} oder {tidytext})
  • Embedding: Je näher der Vektor eines Konstrukts am Text ist, desto typischer ist der Text für das Konstrukt (Distributed Dictionary Representations, Garten et al., 2018); Verfügbar in {dictvectoR} (Thiele, 2022).

Wörterbuch Sentiment in Politischen Texten

Dictionary object with 4 key entries.
- [negative]:
  - a lie, abandon*, abas*, abattoir*, abdicat*, aberra*, abhor*, abject*, abnormal*, abolish*, abominab*, abominat*, abrasiv*, absent*, abstrus*, absurd*, abus*, accident*, accost*, accursed* [ ... and 2,838 more ]
- [positive]:
  - ability*, abound*, absolv*, absorbent*, absorption*, abundanc*, abundant*, acced*, accentuat*, accept*, accessib*, acclaim*, acclamation*, accolad*, accommodat*, accomplish*, accord, accordan*, accorded*, accords [ ... and 1,689 more ]
- [neg_positive]:
  - best not, better not, no damag*, no no, not ability*, not able, not abound*, not absolv*, not absorbent*, not absorption*, not abundanc*, not abundant*, not acced*, not accentuat*, not accept*, not accessib*, not acclaim*, not acclamation*, not accolad*, not accommodat* [ ... and 1,701 more ]
- [neg_negative]:
  - not a lie, not abandon*, not abas*, not abattoir*, not abdicat*, not aberra*, not abhor*, not abject*, not abnormal*, not abolish*, not abominab*, not abominat*, not abrasiv*, not absent*, not abstrus*, not absurd*, not abus*, not accident*, not accost*, not accursed* [ ... and 2,840 more ]

Fragen?

Supervised Machine Learning

Quelle: Computational Analysis of Digital Communication by VU Amsterdam (CC BY 4.0)

Supervised Machine Learning

Supervised Machine Learning

(Laurer et al., 2024)

Fragen?

LLM-basierte Zero|Few-Shot-Klassifikation

  • Seit der Veröffentlichung von Chatbots auf Basis von generativen Large Language Models (ChatGPT, 2023) äußerst populär.
  • Einfache Grundidee: Beschreibung des Konstrukts und Codiereinheit an LLM → Klassifikation als Antwort
  • Anfangs enthusiastischer Hype, mittlerweile Phase der systematischen Evaluation von Chancen und Grenzen
  • Gute praktische Einführungen von Törnberg (2024) und Kravets-Meinke et al. (2025).
  • Siehe Seminarplan AI-powered content analysis für weitere Literatur.
  • Separate Sitzung im Sommersemester 2026 🤓

Fragen?

Fazit & Ausblick

Fazit & Ausblick

  • CTA hat sich von einfachen Wortzählungen zu komplexen semantischen Modellen entwickelt.
  • Die Wahl der Repräsentation (BoW vs. Embeddings) beeinflusst die Analysemöglichkeiten und Fähigkeiten stark. Embedding-basierte Analysen werden zunehmend zum Standard, v.a. für Supervised Methods.
  • Zunehmende Nutzung von Large Language Models, Transfer Learning und Zero- bzw. Few-Shot-Klassifikation in den Sozialwissenschaften. Auch für interessierte Sozialwissenschaftler:innen ohne Informatik-Hintergrund zugänglich.

Fragen?

Hausaufgabe

  1. Lesen Sie Bachl & Scharkow (2024).
  2. Denken Sie an den Bericht aller Hausaufgaben für die Aktive Teilnahme.

Schöne Ferien 🏝️ und eine gute vorlesungsfreihe Zeit

Danke

Marko Bachl

Literatur

Atteveldt, W. H. van. (2008). Semantic network analysis: Techniques for extracting, representing, and querying media content [Phdthesis]. https://research.vu.nl/files/75843774/complete%20dissertation.pdf
Bachl, M., & Scharkow, M. (2024). Computational text analysis. OSF. https://doi.org/10.31219/osf.io/3yhu8
Benoit, K., Watanabe, K., Wang, H., Nulty, P., Obeng, A., Müller, S., & Matsuo, A. (2018). quanteda: An R package for the quantitative analysis of textual data. Journal of Open Source Software, 3(30), 774. https://doi.org/10.21105/joss.00774
Brettschneider, F., & Thoms, C. (2025). Die Wahlprogramme zur Bundestagswahl 2025 im Vergleich aller Bundestagswahlen seit 1949. Universität Hohenheim. https://www.uni-hohenheim.de/uploads/media/Wahlprogramm-Check_BTW_2025.pdf
Chan, C. (2023). grafzahl: fine-tuning Transformers for text data from within R. Computational Communication Research, 5(1), 76–84. https://doi.org/10.5117/CCR2023.1.003.CHAN
Chen, Y., Peng, Z., Kim, S.-H., & Choi, C. W. (2023). What we can do and cannot do with topic modeling: A systematic review. Communication Methods and Measures, 17(2), 111–130. https://doi.org/10/gr4345
Cuilenburg, J. J. van, Kleinnijenhuis, J., & Ridder, J. A. de. (1988). Artificial intelligence and content analysis. Quality and Quantity, 22(1), 65–97. https://doi.org/10/btfvq3
Denny, M. J., & Spirling, A. (2018). Text preprocessing for unsupervised learning: Why it matters, when it misleads, and what to do about it. Political Analysis, 26(2), 168–189. https://doi.org/10/gdjsqk
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). Pre-training of deep bidirectional transformers for language understanding. arXiv. https://doi.org/10.48550/arXiv.1810.04805
DeWeese, L. C. I. (1977). Computer content analysis of „day-old“ newspapers: A feasibility study. Public Opinion Quarterly, 41(1), 91–94. https://doi.org/10/cpvddz
Garten, J., Hoover, J., Johnson, K. M., Boghrati, R., Iskiwitch, C., & Dehghani, M. (2018). Dictionaries and distributions: Combining expert knowledge and large scale textual data content analysis. Behavior Research Methods, 50(1), 344–361. https://doi.org/10/gc3z28
Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30), e2305016120. https://doi.org/10/gsqx5m
Grimmer, J., & Stewart, B. M. (2013). Text as data: The promise and pitfalls of automatic content analysis methods for political texts. Political Analysis, 21(3), 267–297. https://doi.org/10/f458q9
Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794.
Gruber, J. B., & Weber, M. (2025). rollama: Communicate with ’Ollama’ to Run Large Language Models Locally. https://doi.org/10.32614/CRAN.package.rollama
Hase, V., Mahl, D., & Schäfer, M. S. (2023). The "computational turn": an "interdisciplinary turn"? A systematic review of text as data approaches in journalism studies. Online Media and Global Communication, 2(1), 122–143. https://doi.org/10/gt2kx9
Kjell, O., Giorgi, S., & Schwartz, H. A. (2023). The text-package: An R-package for Analyzing and Visualizing Human Language Using Natural Language Processing and Deep Learning. Psychological Methods. https://doi.org/10.1037/met0000542
Kravets-Meinke, D., Schmid-Petri, H., Niemann, S., & Schmid, U. (2025). Generative large language models (gLLMs) in content analysis: A practical guide for communication research. arXiv. https://doi.org/10.48550/arXiv.2510.24337
Kuhn, M., & Wickham, H. (2020). tidymodels: A collection of packages for modeling and machine learning using tidyverse principles. https://www.tidymodels.org
Laurer, M., Atteveldt, W. van, Casas, A., & Welbers, K. (2024). Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI. Political Analysis, 32(1), 84–100. https://doi.org/10/gsgptm
Maier, D., Waldherr, A., Miltner, P., Wiedemann, G., Niekler, A., Keinert, A., Pfetsch, B., Heyer, G., Reber, U., Häussler, T., Schmid-Petri, H., & Adam, S. (2018). Applying LDA topic modeling in communication research: Toward a valid and reliable methodology. Communication Methods and Measures, 12(2-3), 93–118. https://doi.org/10/gfwckq
Manning, C. D., & Schütze, H. (1999). Foundations of statistical natural language processing. MIT Press.
Pipal, C., Song, H., & Boomgaarden, H. G. (2023). If you have choices, why not choose (and share) all of them? A multiverse approach to understanding news engagement on social media. Digital Journalism, 11(2), 255–275. https://doi.org/10/gt2w9r
Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf
Rajapakse, T. (2021). Simple Transformers. In Simple Transformers. https://simpletransformers.ai/
Roberts, M. E., Stewart, B. M., & Tingley, D. (2019). stm: An R package for structural topic models. Journal of Statistical Software, 91(1), 1–40. https://doi.org/10/ggc8cz
Silge, J., & Robinson, D. (2016). tidytext: Text Mining and Analysis Using Tidy Data Principles in R. Journal of Open Source Software, 1(3). https://doi.org/10.21105/joss.00037
Simon, M., Zannettou, S., Welbers, K., Kroon, A. C., & Trilling, D. (2025). Tracing information flows in the hybrid media system: The agenda-setting role of dark platforms surrounding the Ukraine invasion discourse. International Journal of Communication, 19, 25. https://ijoc.org/index.php/ijoc/article/view/23212
Stone, P. J., Bales, R. F., Namenwirth, J. Z., & Ogilvie, D. M. (1962). The General Inquirer: A computer system for content analysis and retrieval based on the sentence as a unit of information. Behavioral Science, 7(4). https://doi.org/10/b2p9mg
Thiele, D. (2022). thieled/dictvectoR: v0.0.1. Zenodo. https://doi.org/10.5281/ZENODO.7079599
Törnberg, P. (2024). Best practices for text annotation with large language models. Sociologica, 18(2), 67–85. https://doi.org/10/g9vgm7
Wijffels, J., & Watanabe, K. (2025). word2vec: Distributed Representations of Words. https://doi.org/10.32614/CRAN.package.word2vec
Young, L., & Soroka, S. (2012). Affective news: The automated coding of sentiment in political texts. Political Communication, 29(2), 205–231. https://doi.org/10/gft3hs