Johns Hopkins University · Whiting School of Engineering

Mark Dredze

My Research

I build and evaluate AI systems for language, and study what happens when they reach the places decisions actually get made — medicine, public health, and policy.

01

What does a language model actually know, and how could we tell?

Factuality, context–memory conflict, and whether a benchmark score predicts anything about real use.

Task Matters: Knowledge Requirements Shape LLM Responses to Context–Memory Conflict

Kaiser Sun, Fan Bai, Mark Dredze

Association for Computational Linguistics (ACL) (Findings), 2026.

ACL Anthology
02

How do you keep an AI system aligned with public health goals over its whole lifecycle?

A framework borrowed from epidemiology for spotting and mitigating AI misalignment in medicine and public health.

Guiding artificial intelligence in public health and medicine with epidemiology: A lifecycle framework for mitigating AI misalignment

Ahmed Hassoon, Christine Lin, Hyun Yi (Jacqualine) Woo, Ruxandra Irimia, Jill A. Marsteller, Anthony Li, Antonio Bander, Hubert Leo, Xiaoyi Peng, David Rastall, Mark Dredze

Annals of Epidemiology, 2025.

Link
03

Do the metrics we use to measure readability actually measure readability?

A critical look at standard readability metrics, and whether they capture what they claim to.

Evaluating the Evaluators: Are readability metrics good measures of readability?

Isabel Cachola, Daniel Khashabi, Mark Dredze

Empirical Methods in Natural Language Processing (EMNLP), 2025.

ACL Anthology
04

What goes wrong when generative AI is deployed in finance, and how do you catch it?

Risks specific to financial services, and practical approaches to mitigating them.

Understanding and Mitigating Risks of Generative AI in Financial Services

Sebastian Gehrmann, Claire Huang, Xian Teng, Sergei Yurovski, Iyanuoluwa Shode, Chirag S. Patel, Arjun Bhorkar, Naveen Thomas, John Doucette, David Rosenberg, Mark Dredze, David Rabinowitz

ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2025.

arXiv
05

How does what a model learns in pre-training shape what fine-tuning can later achieve?

Tracing the relationship between pre-training and fine-tuning over a model's development.

Amuro & Char: Analyzing the Relationship between Pre-Training and Fine-Tuning of Large Language Models

Kaiser Sun, Mark Dredze

NAACL Workshop on Representation Learning for NLP (RepL4NLP), 2025.

ACL Anthology
06

Does retrieval make language models safer?

A safety analysis showing retrieval-augmented generation doesn't reliably improve safety, and can make it worse.

RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models

Bang An, Shiyue Zhang, Mark Dredze

Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), 2025.

07

Can a model not just answer a hard medical question, but explain it?

Benchmarking LLMs on answering and explaining challenging medical questions, not just guessing correctly.

Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

Hanjie Chen, Zhouxiang Fang, Yash Singla, Mark Dredze

Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), 2025.

ACL Anthology

View All Publications →

Research Group

Alexandra DeLucia

Alexandra DeLucia

PhD Student
Jie Gao

Jie Gao

Postdoc
Bernal Jimenez Gutierrez

Bernal Jimenez Gutierrez

Postdoc
Ahmed Hassoon

Ahmed Hassoon

PhD Student
Jen-Tse Huang

Jen-Tse Huang

Postdoc
Heyuan Huang

Heyuan Huang

PhD Student
Sonal Joshi

Sonal Joshi

Postdoc
Kaiser Sun

Kaiser Sun

PhD Student
Miriam Wanner

Miriam Wanner

PhD Student
Didi Zhou

Didi Zhou

PhD Student

Group & Alumni →

Recent News

Latest

I have been appointed the Inaugural Director of the Data Science and AI Institute.

Our paper on RAG LLMs are Not Safer was covered by Waters Technology, Radical Data Science, and Dataconomy.

Alex Spangher, Violet Peng, Sebastian Gehrmann and I received an Outstanding Paper Award at EMNLP for our paper Do LLMs Plan Like Human Writers? Comparing Journalist Coverage of Press Releases with LLMs.

I gave a seminar "Applications of Artificial Intelligence (AI) in Tobacco Regulatory Science" to the Center for Coordination of Analytics, Science, Enhancement , and Logistics (CASEL).

News Archive →

About

Biography, Speaking, and Press.

I am the John C. Malone Professor of Computer Science at Johns Hopkins University and the inaugural Director of the Johns Hopkins Data Science and AI Institute.

I develop artificial intelligence systems based on natural language processing and explore applications to public health and medicine — tobacco control, vaccination, infectious disease surveillance, mental health, drug use, and gun violence prevention — alongside methods for clinical NLP on medical records. My work is regularly covered by NPR, The New York Times, and CNN.

I obtained my PhD from the University of Pennsylvania in 2009.

Speaker Bio & Photographs

Short, medium, and long biographies for event programs, plus high-resolution headshots.

Bios & photographs →
Media Requests

I speak regularly on AI capability and evaluation, and on AI in medicine and public health.

Subject: MEDIA REQUEST mdredze@cs.jhu.edu →
Get in Touch

Get in Touch

Prospective Students

Thinking about a PhD or postdoc

How I advise, what I look for, and how to decide between programs. Please read it before emailing — it answers most of what people ask.

Join the Group →
Media

Journalists and press

I speak regularly on AI capability and evaluation, and on AI in medicine and public health.

Subject: MEDIA REQUEST
mdredze@cs.jhu.edu →
Collaboration

Researchers and partners

Clinical and public-health collaborations, industry research partnerships, and other opportunities in my group.

How to work with me →