Software, Data & Demos

Standalone resources with a home of their own.

Most code and datasets are released alongside a specific paper and are linked from that paper — start with the publications.

Software

Toolkits and reference implementations released by the group.

Neural UMLS Concept Linker

Given a clinical note and span (mention) annotations over the note, this system normalizes each span to it's corresponding UMLS Concept with high accuracy. Elliot Schumacher, Andriy Mulyar, Mark Dredze. Clinical Concept Linking with Contextualized Neural Representations. Association for Computational Linguistics (ACL), 2020.

Cross-Lingual Entity Linker

The architecture is described in Cross-Lingual Transfer in Zero-Shot Cross-Language Entity Linking, Schumacher et al 2020 (to Appear in Findings of ACL 2021, https://arxiv.org/abs/2010.09828).

Do Models of Mental Health Based on Social Media Data Generalize?

Code and data for "Do Models of Mental Health Based on Social Media Data Generalize?" in Findings of EMNLP (2020)

semantic-text-similarity with BERT

An easy-to-use interface to fine-tuned BERT models for computing semantic similarity. This was used for our submission to the N2C2 2019 shared task. https://n2c2.dbmi.hms.harvard.edu/track1

BERT Long Document Classification for Clinical Phenotyping

An easy-to-use interface to fully trained BERT based models for multi-class and multi-label long document classification. Pre-trained models are currently available for two clinical note (EHR) phenotyping tasks: smoker identification and obesity detection. This model was used in our 2019 ML4H workshop paper: https://arxiv.org/abs/1910.13664

Deep DMR

Implementation of Deep Dirichlet Multinomial Regression in python + cython. Adrian Benton, Mark Dredze. Deep Dirichlet Multinomial Regression. North American Chapter of the Association for Computational Linguistics (NAACL), 2018.

Demographer: Gender Identification for Social Media

Demographer is a Python package that identifies demographic characteristics based on a name. It's designed for Twitter, where it takes the name of the user and returns information about his or her likely demographics.

BitbucketPre-2023 Twitter data

SPRITE: Structured PRIor Topic modEls

A general purpose topic modeling package that implements SPRITE from our TACL 2015 paper. The package supports multi-threaded training and makes implementing new models easier: Michael J Paul, Mark Dredze. SPRITE: Generalizing Topic Models with Structured Priors. Transactions of the Association for Computational Linguistics (TACL), 2015.
This also contains a detailed readme and tutorial (see Wiki) on how to use the code.

PARMA: A Predicate Argument Aligner

PARMA is our predicate argument aligner: Travis Wolfe, Benjamin Van Durme, Mark Dredze, Nicholas Andrews, Charley Beller, Chris Callison-Burch, Jay DeYoung, Justin Snyder, Jonathan Weese, Tan Xu, Xuchen Yao. PARMA: A Predicate Argument Aligner. Association for Computational Linguistics (ACL) (short paper), 2013.

Confidence Weighted Learning Library

We have collected most of the core algorithms in the confidence weighted learning framework for release as a software library. Please email me for the code.

Contact me for access

Golden Horse

Code for named entity recognition using embeddings, focused on Chinese social media (Weibo). This code implements the methods in our paper: Nanyun Peng, Mark Dredze. Named Entity Recognition for Chinese Social Media with Jointly Trained Embeddings. Empirical Methods in Natural Language Processing (EMNLP) (short paper), 2015.

GitHubPre-2023 Twitter data

Carmen: Geolocation for Twitter

Carmen is a library for geolocating tweets. Given a tweet, Carmen will return Location objects that represent a physical location. Carmen uses both coordinates and other information in a tweet to make geolocation decisions. It's not perfect, but this greatly increases the number of geolocated tweets over what Twitter provides. The Python and Java versions don't give exactly the same results due to differences in the dependencies. If you use Carmen, please cite: Mark Dredze, Michael J Paul, Shane Bergsma, Hieu Tran. Carmen: A Twitter Geolocation System with Applications to Public Health. AAAI Workshop on Expanding the Boundaries of Health Informatics Using AI (HIAI), 2013.

GitHubPre-2023 Twitter data

Twitter Stream Downloader

Code for downloading data using the Twitter streaming API.

GitHubPre-2023 Twitter data

Mingpipe

Code for Chinese name matching. Given two Chinese person names, assigns a score based on how likely the two names refer to the same person.

csLDA

Cross language topic models based on code-switched documents. Documents can be in different languages and some "glue" documents contain multiple languages. csLDA learns topics for each language and aligns topics across languages.

Multiview Representations of Twitter Users

Code and data for our paper: Adrian Benton, Raman Arora, and Mark Dredze. Learning Multiview Representations of Twitter Users. Association for Computational Linguistics (ACL), 2016.

LinkPre-2023 Twitter data

Automated Reviewer Assignment (used in multiple ACL affiliated conferences)

I authored a process for automatically assigning reviewers to an area for ACL affiliated conferences. The code is freely available for others to use. Let me know if you plan on using this system; I'm happy to answer questions.

Data

Annotated corpora and shared-task resources.

Entries marked “Pre-2023 Twitter data” below predate X ending free API access in 2023 and can no longer be rehydrated from tweet IDs by anyone, including the people who created them. They stay listed for citation continuity.

Coronavirus Twitter Data: A collection of COVID-19 tweets with automated annotations

This dataset contains tweets related to COVID-19. The dataset contains Twitter ids, from which you can download the original data directly from Twitter. Additionally, we include the date, keywords related to COVID-19 and the inferred geolocation. Check detailed information at http://twitterdata.covid19dataresources.org/index.

ZenodoPre-2023 Twitter data

Demographics Race/Ethnicity Training Data

Our paper, “Using Noisy Self-Reports to Predict Twitter User Demographics,” produced a dataset of Twitter users whose profile descriptions may self-report their race or ethnicity. We used this dataset to train classifiers for these demographic labels, and showed that models trained on the collected data perform better on gold standard survey data than models trained only on crowd-sourced data. We distribute the trained models.

LinkPre-2023 Twitter data

Civil Unrest on Twitter (CUT)

Tweets labeled with information related to protest, riots and civil unrest on Twitter. Based on Justin Sech, Alexandra DeLucia, Anna L Buczak, Mark Dredze. Civil Unrest on Twitter (CUT): A Dataset of Tweets to Support Research on Civil Unrest. EMNLP Workshop on Noisy User-generated Text (W-NUT), 2020

GitHubPre-2023 Twitter data

Named Entity Recognition for Chinese Social Media (Weibo)

This dataset contains messages selected from Weibo and annotated according to the DEFT ERE annotation guidelines. Annotations include both name and nominal mentions. The corpus contains 1,890 messages sampled from Weibo between November 2013 and December 2014.

GitHubPre-2023 Twitter data

Vaccine Related English Tweets from the United States: 2016 to 2018

A list of tweet ids sampled from the 1% feed that contain vaccine related keywords. All tweets are in English (according to the Twitter provided metadata) and have been geolocated to the United States (using Carmen).

DownloadPre-2023 Twitter data

Annotated Gun Control/Rights Tweets

This dataset contains 50k (automatically) annotated tweets about gun control and gun rights. It was used in our paper: Adrian Benton, Mark Dredze. Using Author Embeddings to Improve Tweet Stance Classification. EMNLP Workshop on Noisy User-generated Text (W-NUT), 2018.

DownloadPre-2023 Twitter data

Annotations for "Weaponized Health Communication"

This contains 10k tweet annotations for our paper: Weaponized Health Communication: Twitter Bots and Russian Trolls Amplify the Vaccine Debate

DownloadPre-2023 Twitter data

CLPsych Shared Task

The Computational Linguistics and Clinical Psychology (CLPsych) workshop has hosted shared and unshared tasks for several years. In 2015 the shared task used data from Twitter users who state a diagnosis of depression or post traumatic stress disorder (PTSD) along with demographically-matched community controls. The shared task provided an apples-to-apples comparisons of various approaches to modeling language relevant to mental health from social media. The shared task consisted of three binary classification experiments: (1) depression versus control, (2) PTSD versus control, and (3) depression versus PTSD.

LinkPre-2023 Twitter data

RateMD Dataset

This dataset contains the reviews from RateMD with aspect ratings that was used in this paper:
Byron C. Wallace, Michael J. Paul, Urmimala Sarkar, Thomas A. Trikalinos, Mark Dredze. A Large-Scale Quantitative Analysis of Latent Factors and Sentiment in Online Doctor Reviews. Journal of the American Medical Informatics Association (JAMIA), 2014;21(6):1098--1103.

Zika Conspiracy Tweets

This dataset contains annotations for whether a tweet about Zika contains pseudo-scientific information. Analysis of this dataset was published in:
Mark Dredze, David A Broniatowski, Karen M Hilyard. Zika Vaccine Misconceptions: A social media analysis. Vaccine, 2016.

DownloadPre-2023 Twitter data

Vaccination Sentiment and Relevance Tweets

This dataset contains annotations for whether a tweet is relevant to the topic of vaccinations, and if the author is expressing a positive or negative view about vaccines. Analysis of this dataset was published in:
Michael Smith, David A. Broniatowski, Mark Dredze. Using Twitter to Examine Social Rationales for Vaccine Refusal. International Engineering Systems Symposium (CESUN), 2016.
Mark Dredze, David A. Broniatowski, Michael Smith, Karen M. Hilyard. Understanding Vaccine Refusal: Why We Need Social Media Now. American Journal of Preventive Medicine, 2015.

DownloadPre-2023 Twitter data

Flu Vaccination Tweets

This dataset contains annotations for whether a tweet is relevant to the topic of flu vaccination, and if the author intends to receive a flu vaccine. Analysis of this dataset was published in:
Xiaolei Huang, Michael C. Smith, Michael Paul, Dmytro Ryzhkov, Sandra Quinn, David Broniatowski, Mark Dredze. Examining Patterns of Influenza Vaccination in Social Media. AAAI Joint Workshop on Health Intelligence (W3PHIAI), 2017.

DownloadPre-2023 Twitter data

Named Entity Recognition and Entity Linking for Speech

This corpus contains broadcast news transcripts annotated for named entities and entity linking against the TAC KBP 2009 corpus. This was used in our NAACL 2015 paper "Entity Linking for Spoken Language" and in our 2011 Interspeech paper:
Carolina Parada, Mark Dredze, Frederick Jelinek. OOV Sensitive Named-Entity Recognition in Speech. International Speech Communication Association (INTERSPEECH), 2011.

Twitter Grammy XDoc Corpus: Entity Linking and Disambiguation

This corpus contains tweets about the Grammy Award ceromony annotated for entity linking and cross document coreference resolution (entity disambigutation). The corpus is described in our paper:
Mark Dredze, Nicholas Andrews, Jay DeYoung. Twitter at the Grammys: A Social Media Corpus for Entity Linking and Disambiguation. EMNLP Workshop on Natural Language Processing for Social Media, 2016.

BitbucketPre-2023 Twitter data

Twitter Health Keywords

These files contain the keywords we use to collect and identify health related tweets.

FigsharePre-2023 Twitter data

Health Twitter Annotations

These annotations were created for the paper:
Michael J. Paul, Mark Dredze. A Model for Mining Public Health Topics from Twitter. Technical Report -, Johns Hopkins University, 2011.
The annotations label tweets as they relate to health. The annotations are described on page 2 of the paper. The file includes tweet ids which you can use to download the data.

DownloadPre-2023 Twitter data

Influenza Twitter Annotations

These annotations were created for the paper:
Alex Lamb, Michael J. Paul, Mark Dredze. Separating Fact from Fear: Tracking Flu Infections on Twitter. North American Chapter of the Association for Computational Linguistics (NAACL), 2013.
The annotations label tweets are related to influenza, awareness vs. infection and if the tweet is about the author or someone else. The files include tweet ids which you can use to download the data.

DownloadPre-2023 Twitter data

Twitter Hurricane Sandy Dataset

A collection of tweets from areas hit by hurricane Sandy (2012) in the United States. This dataset is meant for research in social media disaster response.
Haoyu Wang, Eduard Hovy, Mark Dredze. The Hurricane Sandy Twitter Corpus. AAAI Workshop on the World Wide Web and Public Health Intelligence, 2015.

GitHubPre-2023 Twitter data

Twitter First Name, Last Name, and Location Clusters

A set of clusters extracted from Twitter that contains firstnames, lastnames, and locations. We used this in our NAACL 2013 paper:
Shane Bergsma, Mark Dredze, Benjamin Van Durme, Theresa Wilson, David Yarowsky. Broadly Improving User Classification via Communication-Based Name and Location Clustering on Twitter. North American Chapter of the Association for Computational Linguistics (NAACL), 2013.

LinkPre-2023 Twitter data

Enron Attachment Prediction Email

Enron emails annotated with attachment information and cleaned of numerous artifacts inserted by email programs. Unfortunately, I don't have a copy of the attachments. Very few groups had that data, and I am not aware of anyone who currently has a copy. Email me for the data.

Contact me for access

Multi-Domain Sentiment Dataset

Product reviews from several different product types taken from Amazon.com. This dataset is from:
John Blitzer, Mark Dredze, Fernando Pereira. Biographies, Bollywood, Boom-boxes and Blenders: Domain Adaptation for Sentiment Classification. Association for Computational Linguistics (ACL), 2007.

Image Spam Dataset

A collection of ham and spam images taken from real user email. This dataset is from:
Mark Dredze, Reuven Gevaryahu, Ari Elias-Bachrach. Learning Fast Classifiers for Image Spam. Conference on Email and Anti-Spam (CEAS), 2007.

TAC 2009 Entity Linking

A collection of manually linked training examples to supplement those provided in the TAC 2009 KBP task. These are described in my Coling 2010 paper on entity linking. If you use this data, please cite:
Mark Dredze, Paul McNamee, Delip Rao, Adam Gerber, Tim Finin. Entity Disambiguation for Knowledge Base Population. Conference on Computational Linguistics (Coling), 2010.

Demos

Public systems built out of the research, several maintained with collaborators.

Twitter Social Mobility Index

We are monitoring public Tweets to gather geolocation data of users to see how people are socially distancing after the March 16th guidelines announced nationally by the White House. On average, there is a 54.97% reduction in movement within the United States, sampled by Tweets from public Twitter users. By tracking geolocation data, we can find which state, city, or region within the US is following the guidelines and reducing travel the most and how that could locally link to reduction or surges of COVID-19.

No longer available

Health Tweets

HealthTweets.org is a research platform for sharing the latest developments in mining health trends from Twitter and other social media sites. HealthTweets.org allows public health practitioners and researchers direct access to the output of the group’s data mining algorithms.

LinkNo longer available

Tobacco Watcher

Tobacco Watcher is an automated surveillance system for tobacco focused media analytics. Our computers search the web in dozens of languages for news articles, and identifies each article by tobacco control topic, location, date and language to make media monitoring easier and more effective. By operating continuously in real time, Tobacco Watcher can help you stay on top of a fast changing global landscape by freely delivering and analyzing the world’s tobacco-related media.