Published:
Chien-Ming Huang presents "Conversational Agents for Behavioral Medicine" on TV screens in front of the audience at the in-person convening of the Johns Hopkins Workgroup on AI and Healthcare.
Image Credit: Hopkins Business of Health Initiative

The Johns Hopkins Workgroup on AI and Healthcare had been running solely on Zoom for three and a half years before it finally convened in person on June 4.

About 60 faculty, researchers, and clinicians from the schools of Medicine, Public Health, Engineering, Nursing, and Business filled a conference room on the terrace level of the university’s Mt. Washington campus, where they gathered for a half-day round of lightning talks co-hosted by the Hopkins Business of Health Initiative (HBHI) and the Johns Hopkins Data Science and AI Institute (DSAI). The workgroup itself started in January 2023, two months after the public release of ChatGPT, as a monthly conversation about how AI was changing medical practice.

Mark Dredze, the inaugural director of the DSAI and a John C. Malone Professor of Computer Science, told the room that the partnerships DSAI most wants to make possible were exactly the kind on display that morning.

“The gap between AI that’s technically impressive and AI that’s actually clinically trustworthy is critical,” he said, “and Hopkins has to be a leader in that space.”

Michael Oberst gestures.

Michael Oberst. Image Credit: Hopkins Business of Health Initiative

Michael Oberst, an assistant professor of computer science and member of the Malone Center for Engineering in Healthcare, opened the first session on AI in clinical care with the observation that a randomized clinical trial of a medical AI tool takes about a year to run. By the time the results come in, the developers have usually moved on to a newer version of the model.

That’s why his group is working on methods for using the data from older trials to bound the likely effects of newer ones. These methods have to handle a finding that surprised his group: Higher-accuracy models do not always lead to better patient outcomes, because clinicians who learn to trust a model start to defer to it.

“Even a very high-performing model, if it makes an incorrect prediction, can lead to worse outcomes because people trust it too much,” Oberst said.

Chien-Ming Huang presents to an audience in front of a TV screen with three bar graphs that reads "Conversational sleep diary: Capturing richer context."

Chien-Ming Huang. Image Credit: Hopkins Business of Health Initiative

John C. Malone Associate Professor of Computer Science Chien-Ming Huang spoke later about a project on conversational AI in behavioral sleep medicine.

Patients with sleep problems are typically asked to keep a paper sleep diary for two weeks, but compliance is low; many patients fill out the whole diary the morning before their next appointment—which is why Huang’s group has built a conversational AI agent that patients can talk to each morning.

In a small study, the agent collected more and more detailed information than the paper diary did. Patients told the agent why they had slept badly, and the sleep specialist’s role ended up changing, too.

“The specialist brings their insights by asking questions,” Huang explained, “and the system brings the data to support these insights.”

Excerpted from the HBHI website »