LECTURE SUMMARY - Computation Analysis of Digital Communication lecture notes
+ EXAM QUESTIONS.
Please don't mind the English gramma since I wanted to make it as compact as possible :) (from 51 pages to 19 pages!!)
LECTURE 1 - What is computational social sciences? 2
PRELIMINARY SUMMARY 4
LECTURE 2 - Text as Data - Basics of Automatic Text Analysis 4
Text as Data - How can we analyze texts with computers? 4
Automated Text Analysis 5
Automatic text analysis steps - van Atteveldt, Welbers, & Van der Velden, 2019 5
Deductive Approaches: Dictionary-based Text Analysis 6
Example - state of union speech corpus 6
LECTURE 3 - Machine Learning - Supervised Text Classification 8
What is machine learning? 8
Supervised text classification 9
Principles of supervised text classification 9
Validation 11
Example - prediction musing genre from lyrics (homework 2/3A) 11
Conclusion 12
LECTURE 4 - Machine Learning - Unsupervised Topic Modeling 13
What is topic modeling? - based on example: nuclear technology from 1945 - 2013 13
Topic Modeling as Dimensionality Reduction 14
Latent Dirichlet Allocation - LDA topic modeling 14
Conclusion and outlook 18
Conclusion 18
EXAMPLE EXAM QUESTION (MULTIPLE CHOICE) 23
EXAMPLE EXAM QUESTION (OPEN FORMAT) 23
, CADC 2022 - KAYI MAN
LECTURE 1 - What is computational social sciences?
● Field of Social Science that uses algorithmic tools and large/unstructured data to understand
human and social behavior
● Computational methods as “microscope”: Methods are not the goal, but contribute to
theoretical development and/or data generation
● Complements rather than replaces traditional methodologies
● Includes methods such as, e.g.,:
○ Advanced data wrangling/data science
○ Combining of different data sets
○ Automated Text Analysis
○ Machine Learning (supervised and unsupervised)
○ Actor-based modeling
○ Simulations
○ …
● To better understand text, results. “how do we understand large data set”
● How can we work with data this large
● To big to put in excel-sheet
● Unsupervised vs. Supervised
TYPICAL WORKFLOW
1. Identification problem/purpose =
2. Data acquisitions = Different way to get existing data
3. Data wrangling = What do i have to transfor, add, delete the data to make it useful
4. Data analysis & modeling = statistical analysis or creating algorithms
5. Reporting = using data to communicate …. ?
WHY IS THIS IMPORTANT NOW?
- Collecting data used to be expensive (surveys, observations)
- Digital age: behaviors of billions are recorded, stored and therefore analyzable
- Digital record of behavior is created by everytime/thing you click/call/pay
- (meta-)data are byproduct of peeps everyday actions aka digital traces
- Big data = often called large-scale records of peeps/businesses
10 CHARACTERISTICS OF BIG DATA (Salganik, 2017, chap. 2.3)
1. Big = scale / volume of current datasets is often impressive
2. Always-on = big data systems are constantly collecting data (FB = always-on)
3. Non reactive = subjects are non reactive and not aware of the collecting (ethical?) of zijn zo
gewend dat het hun behavior niet veranderd
4. Incomplete = most big data sources are incomplete, don't have info that you want to
research. Because data was created for other purposes than research.
5. Inaccessible = Data held by companies/governments are difficult for researchers to access.
6. Non representative = not representative of certain populations
7. Drifting = systems are changing constantly, difficult for long-term study trends. The way they
do it, changes
8. Algorithmically confounded = behavior in big data is not natural; driving by engineering
goals. Weird algorithm implemented by FB, predetermines how the data is gonna look like.
Record produced by the system that is built by platform.
9. Dirty = Big data includes noise (junk, spam)
10. Sensitive = some info that companies/governments have, are sensitive (ethical?)
Privacy issues
, CADC 2022 - KAYI MAN
TYPICAL COMPUTATIONAL RESEARCH STRATEGIES
1. Counting things (how often do peep use phones per day? What topics do news sites cover most?
2. Forecasting and nowcasting (predictions both present and in future; crime prediction…)
3. Approximating experiments (investigate potential nudges to make user select certain news)
ADVANTAGES AND DISADVANTAGES
Advantages of Computational Methods
- Actual behavior vs. self-report (because biased)
- Social context vs. lab setting
- Small N to large N
Disadvantages of Computational Methods
- Techniques often complicated
- Data often proprietary (=eigendomsrecht)
- Samples often biased
- Insufficient metadata (we have data but don’t know who they are)
DEFINITION Van Atteveldt & Peng, 2018
“Computational Communication Science is the
- label applied to the emerging subfield that investigates
- the use of computational algorithms
- to gather and analyze big and often semi- or unstructured data sets
- to develop and test communication science theories”
PROMISES
Three developments fueled the computational methods of communication sciences
1. Vast amounts of digitally available data
2. Improved tools to analyze big data (auto text analysis methods) changes fast!
3. Powerful and cheap processing power & easy computing infrastructure (Github)
ETHICS OF ‘BIG DATA’ AND COMPUTATIONAL RESEARCH
THE “FACEBOOK MOOD MANIPULATION” STUDY (Kramer et al., 2014)
● Massive online experiment (N ~ 700k)
● Main Research Question: Is emotion contagious?
● Experimental groups: positive / negative / control
● Stimulus: Hide (negative / positive / random) messages from FB timeline
● Measurement / dependent variables: sentiment of posts by user
Question: Do you think these studies are problematic? If yes, why?
● No consent is given
● Shared with third party
COMPUTATIONAL TECHNIQUE: SENTIMENT ANALYSIS
● Count occurrences of words in both categories, subtract negative
from positive
Positive words reduced in feed = more negative words used
Negative words reduced in feed = more positive words used
IS THIS GOOD SCIENCE? WHY NOT?
● What’s cool?
○ Potentially interesting research question
○ actual behavior measured as well as self-report measures
● What’s not so cool? A lot…
○ No informed consent, not replicable, manipulation
○ Low internal validity
■ Is sentiment of posts indicative of mood?
■ Does change in sentiment originate in contagion of mood?
○ Low measurement accuracy
■ Are word counts indicative of sentiment?
The benefits of buying summaries with Stuvia:
Guaranteed quality through customer reviews
Stuvia customers have reviewed more than 700,000 summaries. This how you know that you are buying the best documents.
Quick and easy check-out
You can quickly pay through credit card or Stuvia-credit for the summaries. There is no membership needed.
Focus on what matters
Your fellow students write the study notes themselves, which is why the documents are always reliable and up-to-date. This ensures you quickly get to the core!
Frequently asked questions
What do I get when I buy this document?
You get a PDF, available immediately after your purchase. The purchased document is accessible anytime, anywhere and indefinitely through your profile.
Satisfaction guarantee: how does it work?
Our satisfaction guarantee ensures that you always find a study document that suits you well. You fill out a form, and our customer service team takes care of the rest.
Who am I buying these notes from?
Stuvia is a marketplace, so you are not buying this document from us, but from seller pikayichu. Stuvia facilitates payment to the seller.
Will I be stuck with a subscription?
No, you only buy these notes for $6.40. You're not tied to anything after your purchase.