Johns Hopkins University · Whiting School of Engineering

Mark Dredze

My Research

I build and evaluate AI systems for language, and study what happens when they reach the places decisions actually get made — medicine, public health, and policy.

01

What does a language model actually know, and how could we tell?

Factuality, context–memory conflict, and whether a benchmark score predicts anything about real use.

Task Matters: Knowledge Requirements Shape LLM Responses to Context–Memory Conflict

Kaiser Sun, Fan Bai, Mark Dredze

Association for Computational Linguistics (ACL) (Findings), 2026.

ACL Anthology
02

How do you keep an AI system aligned with public health goals over its whole lifecycle?

A framework borrowed from epidemiology for spotting and mitigating AI misalignment in medicine and public health.

Guiding artificial intelligence in public health and medicine with epidemiology: A lifecycle framework for mitigating AI misalignment

Ahmed Hassoon, Christine Lin, Hyun Yi (Jacqualine) Woo, Ruxandra Irimia, Jill A. Marsteller, Anthony Li, Antonio Bander, Hubert Leo, Xiaoyi Peng, David Rastall, Mark Dredze

Annals of Epidemiology, 2025.

Link
03

Does a summary give this particular reader what they need?

A reader-centered way to evaluate summarization, showing that standard and LLM-as-judge metrics miss what a reader's background and needs require.

Best Paper Award

Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation

Isabel Cachola, William Walden, Reno Kriz, Mark Dredze

ACM Conference on Human-AI Complementarity and Alignment (HCOMP), 2026.

arXiv
04

What goes wrong when generative AI is deployed in finance, and how do you catch it?

Risks specific to financial services, and practical approaches to mitigating them.

Understanding and Mitigating Risks of Generative AI in Financial Services

Sebastian Gehrmann, Claire Huang, Xian Teng, Sergei Yurovski, Iyanuoluwa Shode, Chirag S. Patel, Arjun Bhorkar, Naveen Thomas, John Doucette, David Rosenberg, Mark Dredze, David Rabinowitz

ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2025.

arXiv
05

How does what a model learns in pre-training shape what fine-tuning can later achieve?

Tracing the relationship between pre-training and fine-tuning over a model's development.

Amuro & Char: Analyzing the Relationship between Pre-Training and Fine-Tuning of Large Language Models

Kaiser Sun, Mark Dredze

NAACL Workshop on Representation Learning for NLP (RepL4NLP), 2025.

ACL Anthology
06

Does retrieval make language models safer?

A safety analysis showing retrieval-augmented generation doesn't reliably improve safety, and can make it worse.

RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models

Bang An, Shiyue Zhang, Mark Dredze

Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), 2025.

07

What is an AI agent actually doing when it runs for hours on its own?

A taxonomy of agent actions, built with Grounded Theory, that turns long agent trajectories into behavioral profiles you can compare and use to spot failure modes.

How to Interpret Agent Behavior

Jie Gao, Kaiser Sun, Jen-tse Huang, Katherine Van Koevering, Sijie Ji, Heyuan Huang, Weiyan Shi, Zhuoran Lu, Ziang Xiao, Daniel Khashabi, Mark Dredze

Neural Information Processing Systems (NeurIPS) (Evaluation & Datasets Track), 2026.

arXiv

View All Publications →

Research Group

Alexandra DeLucia

Alexandra DeLucia

PhD Student
Jie Gao

Jie Gao

Postdoc
Bernal Jimenez Gutierrez

Bernal Jimenez Gutierrez

Postdoc
Ahmed Hassoon

Ahmed Hassoon

PhD Student
Jen-Tse Huang

Jen-Tse Huang

Postdoc
Heyuan Huang

Heyuan Huang

PhD Student
Kaiser Sun

Kaiser Sun

PhD Student
Miriam Wanner

Miriam Wanner

PhD Student
Didi Zhou

Didi Zhou

PhD Student

Group & Alumni →

Recent News

Latest

Isabel Cachola, William Walden, Reno Kriz and I received a Best Paper Award at HCOMP 2026 for our paper Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation.

I have been appointed the Inaugural Director of the Data Science and AI Institute.

Our paper on RAG LLMs are Not Safer was covered by Waters Technology, Radical Data Science, and Dataconomy.

Alex Spangher, Violet Peng, Sebastian Gehrmann and I received an Outstanding Paper Award at EMNLP for our paper Do LLMs Plan Like Human Writers? Comparing Journalist Coverage of Press Releases with LLMs.

News Archive →

About

Biography, Speaking, and Press.

I am the John C. Malone Professor of Computer Science at Johns Hopkins University and the inaugural Director of the Johns Hopkins Data Science and AI Institute.

I develop artificial intelligence systems based on natural language processing and explore applications to public health and medicine — tobacco control, vaccination, infectious disease surveillance, mental health, drug use, and gun violence prevention — alongside methods for clinical NLP on medical records. My work is regularly covered by NPR, The New York Times, and CNN.

I obtained my PhD from the University of Pennsylvania in 2009.

Speaker Bio & Photographs

Short, medium, and long biographies for event programs, plus high-resolution headshots.

Bios & photographs →
Media Requests

I speak regularly on AI capability and evaluation, and on AI in medicine and public health.

Subject: MEDIA REQUEST mdredze@cs.jhu.edu →
Get in Touch

Get in Touch

Prospective Students

Thinking about a PhD or postdoc

How I advise, what I look for, and how to decide between programs. Please read it before emailing — it answers most of what people ask.

Join the Group →
Media

Journalists and press

I speak regularly on AI capability and evaluation, and on AI in medicine and public health.

Subject: MEDIA REQUEST
mdredze@cs.jhu.edu →
Collaboration

Researchers and partners

Clinical and public-health collaborations, industry research partnerships, and other opportunities in my group.

How to work with me →