4 mini-projects · 30% of final grade
Assignments
The submission part of the course consists of four assignments, solved in groups of 1–3 students. The collected assignments are handed in as one submission at the end of the semester. During the semester each group presents and gets feedback on their work.
Assignment 1: Boolean search, topic models, and sentiment
Do State of the Union speeches predict GDP growth? Boolean search, text regression, LDA topic modelling, and sentiment scoring on two centuries of speeches
Assignment 2: Word embeddings
Train, compare, and interrogate word embeddings on the State of the Union corpus — word2vec, PMI+SVD, analogies, and semantic change over two centuries
Assignment 3: Fine-tuning a language model
Fine-tune BERT for text classification — first on Yelp reviews, then on your own State of the Union sentiment labels — and benchmark it against the simple baselines
Assignment 4: Prompt engineering and a marketing use-case
Systematic prompt experimentation, then a critical reflection: where would text as data actually create value in a marketing scenario — and what could go wrong?