EN.601.762 · JOHNS HOPKINS UNIVERSITY · FALL 2026

Multimodal Understanding and Generation

Instructor: Jaemin Cho (jaemin@jhu.edu) · Mon/Wed 1:30–2:45 pm · Hackerman 320 · Office hours: by appointment

This course provides a deep dive into modern Multimodal AI, focusing on models that integrate vision and language data. Topics include visual question answering, generative media (images/video), neuro-symbolic AI, and embodied AI agents. Through weekly paper reviews and a hands-on independent research project, students will gain the technical skills to understand and advance the state-of-the-art in multimodal deep learning.

Required Course Background: at least one upper-level/grad course in vision, NLP or machine learning.

Paper discussion

Each topic day covers 2–3 papers. Everyone reads all of that day's papers before class. Every student holds three roles over the semester:

Presenter (1–2×)

Present one or two of the day's papers — method, results, and answers to the class's submitted questions. 45 min shared across the day's papers, usually with a co-presenter; a couple of tightly paired sessions are covered by one presenter.

Archaeologist (≤1×)

A 10-minute survey of the day's papers' ancestors and successors — what came before, what built on them, where they sit.

Hacker (≤1×)

A 10-minute Colab implementing a small piece of one paper on a toy problem, walked through live.

Reader (everyone, every class)

Submit a one-page review of one paper (guideline below) via Gradescopeevery class, including days you present, archaeologist, or hack. Due 11:59 pm EST: previous Wednesday for Monday classes, previous Friday for Wednesday classes. No late submissions; your three lowest-scored reviews are dropped.

All submissions — paper reviews, project abstract, midpoint and final reports — go through Gradescope (entry code: B4P7WR), in PDF format — reviews, reports, and presentation slides alike.

After each class, the day's presenters, archaeologist, and hacker upload their materials (slides as PDF / Colab link) to the shared OneDrive folder — the link will be shared after the first class. Course announcements and questions happen on MS Teams; role assignments and Teams invitations will also go out after the first class.

Paper review guideline

One page per review, in four sections:

  1. Summary — what the paper does and claims
  2. Strengths
  3. Weaknesses
  4. Questions for presenters

After each deadline, the instructor compiles the submitted reviews and shares them on MS Teams — presenters address the questions during their session.

Research project

This course is heavily project-based: the project accounts for more than half of your grade. Teams of 2 recommended; any topic within the course's scope.

Deliverables: proposal · mid report · mid presentation · final poster session · final report.

A full-score final project is expected to be comparable in quality to a paper at a top AI conference (NeurIPS, CVPR, ICML, ICLR, ICCV, ACL, COLM, etc.).

Mid and final reports follow the NeurIPS 2026 format. Page limits — mid: 4 pages, final: 8 pages — including all main text, figures, and tables; references excluded.

Sep 16 — proposal

One-pager per team, due 11:59 pm:

  • A 200–300 word project abstract
  • Planned mid deliverables
  • Planned final deliverables
  • Team members and external collaborators (if any)

Only one submission per team is needed.

Oct 19 — mid report & presentation

All projects present in class; mid report in NeurIPS 2026 format, up to 4 pages.

Dec 9 — final poster session

All teams, in class.

Dec 15 — final report

NeurIPS 2026 format, up to 8 pages.

Schedule

Role assignments below are tentative and may change until sign-up closes on Teams. A cell reading TBD is an unfilled seat. If you would prefer your name not appear on this public page, email the instructor and it will be removed.

All deadlines are 11:59 pm EST — reviews are due the previous Wednesday for Monday classes, and the previous Friday for Wednesday classes.

#Date Review due TopicPresenter ArchaeologistHackerPapers
1Mon, Aug 31Course introduction — logistics, grading, roles, project expectations, role sign-upJaemin Cho
2Wed, Sep 2Understanding: Foundations (vision-language pretraining)Jaemin ChoJaemin ChoJaemin Cho
Mon, Sep 7No class — Labor Day
3Wed, Sep 9Guest lectureYushi Hu (Meta Superintelligence Labs)
4Mon, Sep 14Wed, Sep 9Generation: Text-to-image foundationsAsma Akel Alkhaldi · Jiwoo NohLurui WangTBD
5Wed, Sep 16Fri, Sep 11Understanding: Text-to-image/video groundingLukas William Geer · Ram Sai GaneshTBDZhanpeng Luo
Wed, Sep 16Project abstract due, 11:59 pm — one page per team
6Mon, Sep 21Wed, Sep 16Understanding: Modern vision-language modelsJingyang Chen · Shreyas PhegadeLin LongMatt Wang
7Wed, Sep 23Fri, Sep 18Generation: Modern text-to-image architecturesTengyang Deng · Xavier VelezRam Sai GaneshAsma Akel Alkhaldi
8Mon, Sep 28Wed, Sep 23Generation: Evaluation of image/video generatorsRuijie Tao · Zijian ZhangTBDKatherine M Guerrerio
9Wed, Sep 30Fri, Sep 25Understanding: Evaluation of multimodal understanding modelsKehao Xu · Timing YangZhanpeng LuoShenghan Zhou
10Mon, Oct 5Wed, Sep 30Efficient adaptation of understanding/generation modelsRui Liu · TBDTBDLurui Wang
11Wed, Oct 7Fri, Oct 2Generation: Editing and controlled generationMatt Wang · TBDJiayin LiRuijie Tao
12Mon, Oct 12Wed, Oct 7Generation: Video generationZhanpeng Luo (solo — covers both papers)Adway U KanhereLin Long
13Wed, Oct 14Fri, Oct 9Generation: Motion and musicShenghan Zhou · TBDXavier VelezShreyas Phegade
14Mon, Oct 19Mid-project presentation — all projectsAll students
15Wed, Oct 21Fri, Oct 16Understanding: Neuro-symbolicAdway U Kanhere · Lukas William GeerRuijie TaoJiayin Li
16Mon, Oct 26Wed, Oct 21Generation: Neuro-symbolicTBD · Timing YangChenyu LiKehao Xu
17Wed, Oct 28Fri, Oct 23Embodied AI: Vision-language-action modelsLin Long (solo — covers both papers)Shreyas PhegadeTBD
18Mon, Nov 2Wed, Oct 28Embodied AI: Learning actions from videosRui Liu · TBDShenghan ZhouJiwoo Noh
19Wed, Nov 4Fri, Oct 30Embodied AI: Spatial understanding and gaze in videoLurui Wang · Xavier VelezTengyang DengLukas William Geer
20Mon, Nov 9Wed, Nov 4Advanced: Unified understanding and generationChenyu Li · Jiayin LiKehao XuTiming Yang
21Wed, Nov 11Fri, Nov 6Advanced: Long video understandingAsma Akel Alkhaldi · Jingyang ChenZijian ZhangRui Liu
22Mon, Nov 16Wed, Nov 11Advanced: Multimodal Web AgentsJiwoo Noh · Ram Sai GaneshKatherine M GuerrerioAdway U Kanhere
23Wed, Nov 18Fri, Nov 13Advanced: World modelsMatt Wang · Zijian ZhangJingyang ChenChenyu Li
Mon, Nov 23No class — Thanksgiving
Wed, Nov 25No class — Thanksgiving
24Mon, Nov 30Wed, Nov 25Advanced: Multimodal rewards and feedbackKatherine M Guerrerio · Tengyang DengTBDTBD
25Wed, Dec 2Guest lectureTBD
26Mon, Dec 7Guest lectureTBD
27Wed, Dec 9Final project poster sessionAll students
Tue, Dec 15Final project report due

Grading

WeightComponent
30%Role-playing paper discussion (presenter / archaeologist / hacker)
15%Paper review assignments
15%Mid proposal & report
20%Final presentation
20%Final report