EN.601.762 · JOHNS HOPKINS UNIVERSITY · FALL 2026

Multimodal Understanding and Generation

Instructor: Jaemin Cho (jaemin@jhu.edu) · Mon/Wed 1:30–2:45 pm · Hackerman 320 · Office hours: by appointment

This course provides a deep dive into modern Multimodal AI, focusing on models that integrate vision and language data. Topics include visual question answering, generative media (images/video), neuro-symbolic AI, and embodied AI agents. Through weekly paper reviews and a hands-on independent research project, students will gain the technical skills to understand and advance the state-of-the-art in multimodal deep learning.

Required Course Background: at least one upper-level/grad course in vision, NLP or machine learning.

Paper discussion

Each topic day covers 2–3 papers. Everyone reads all of that day's papers before class. Every student holds three roles over the semester:

Presenter (1–2×)

Present one paper as part of that day's team — method, results, and answers to the class's submitted questions. 45 min shared across the day's papers.

Archaeologist (≤1×)

A 10-minute survey of the day's papers' ancestors and successors — what came before, what built on them, where they sit.

Hacker (≤1×)

A 10-minute Colab implementing a small piece of one paper on a toy problem, walked through live.

Reader (everyone, every class)

Submit a one-page review of one paper (guideline below) via Gradescopeevery class, including days you present, archaeologist, or hack. Due 11:59 pm EST: previous Wednesday for Monday classes, previous Friday for Wednesday classes. No late submissions; your three lowest-scored reviews are dropped.

All submissions — paper reviews, project abstract, midpoint and final reports — go through Gradescope (entry code: B4P7WR), in PDF format — reviews, reports, and presentation slides alike.

After each class, the day's presenters, archaeologist, and hacker upload their materials (slides as PDF / Colab link) to the shared OneDrive folder — the link will be shared after the first class. Course announcements and questions happen on MS Teams; role assignments and Teams invitations will also go out after the first class.

Paper review guideline

One page per review, in four sections:

  1. Summary — what the paper does and claims
  2. Strengths
  3. Weaknesses
  4. Questions for presenters

After each deadline, the instructor compiles the submitted reviews and shares them on MS Teams — presenters address the questions during their session.

Research project

This course is heavily project-based: the project accounts for more than half of your grade. Teams of 2 recommended; any topic within the course's scope.

Deliverables: proposal · mid report · mid presentation · final poster session · final report.

A full-score final project is expected to be comparable in quality to a paper at a top AI conference (NeurIPS, CVPR, ICML, ICLR, ICCV, ACL, COLM, etc.).

Mid and final reports follow the NeurIPS 2026 format. Page limits — mid: 4 pages, final: 8 pages — including all main text, figures, and tables; references excluded.

Sep 16 — proposal

One-pager per team, due 11:59 pm:

  • A 200–300 word project abstract
  • Planned mid deliverables
  • Planned final deliverables
  • Team members and external collaborators (if any)

Only one submission per team is needed.

Oct 19 — mid report & presentation

All projects present in class; mid report in NeurIPS 2026 format, up to 4 pages.

Dec 9 — final poster session

All teams, in class.

Dec 15 — final report

NeurIPS 2026 format, up to 8 pages.

Schedule

All deadlines are 11:59 pm EST — reviews are due the previous Wednesday for Monday classes, and the previous Friday for Wednesday classes.

#Date Review due TopicPresenter ArchaeologistHackerPapers
1Mon, Aug 31Course introduction — logistics, grading, roles, project expectations, role sign-upJaemin Cho
2Wed, Sep 2Understanding: Foundations (vision-language pretraining)Jaemin ChoJaemin ChoJaemin Cho
Mon, Sep 7No class — Labor Day
3Wed, Sep 9Guest lectureYushi Hu (Meta Superintelligence Labs)
4Mon, Sep 14Wed, Sep 9Generation: Text-to-image foundationsTBDTBDTBD
5Wed, Sep 16Fri, Sep 11Understanding: Text-to-image/video groundingTBDTBDTBD
Wed, Sep 16Project abstract due, 11:59 pm — one page per team
6Mon, Sep 21Wed, Sep 16Understanding: Modern vision-language modelsTBDTBDTBD
7Wed, Sep 23Fri, Sep 18Generation: Modern text-to-image architecturesTBDTBDTBD
8Mon, Sep 28Wed, Sep 23Generation: Evaluation of image/video generatorsTBDTBDTBD
9Wed, Sep 30Fri, Sep 25Understanding: Evaluation of multimodal understanding modelsTBDTBDTBD
10Mon, Oct 5Wed, Sep 30Efficient adaptation of understanding/generation modelsTBDTBDTBD
11Wed, Oct 7Fri, Oct 2Generation: Editing and controlled generationTBDTBDTBD
12Mon, Oct 12Wed, Oct 7Generation: Video generationTBDTBDTBD
13Wed, Oct 14Fri, Oct 9Generation: Motion and musicTBDTBDTBD
14Mon, Oct 19Mid-project presentation — all projectsAll students
15Wed, Oct 21Fri, Oct 16Understanding: Neuro-symbolicTBDTBDTBD
16Mon, Oct 26Wed, Oct 21Generation: Neuro-symbolicTBDTBDTBD
17Wed, Oct 28Fri, Oct 23Embodied AI: Vision-language-action modelsTBDTBDTBD
18Mon, Nov 2Wed, Oct 28Embodied AI: Learning actions from videosTBDTBDTBD
19Wed, Nov 4Fri, Oct 30Embodied AI: Spatial understanding and gaze in videoTBDTBDTBD
20Mon, Nov 9Wed, Nov 4Advanced: Unified understanding and generationTBDTBDTBD
21Wed, Nov 11Fri, Nov 6Advanced: Long video understandingTBDTBDTBD
22Mon, Nov 16Wed, Nov 11Advanced: Visual prompting and GUI agentsTBDTBDTBD
23Wed, Nov 18Fri, Nov 13Advanced: World modelsTBDTBDTBD
Mon, Nov 23No class — Thanksgiving
Wed, Nov 25No class — Thanksgiving
24Mon, Nov 30Wed, Nov 25Advanced: Multimodal rewards and feedbackTBDTBDTBD
25Wed, Dec 2Guest lectureTBD
26Mon, Dec 7Guest lectureTBD
27Wed, Dec 9Final project poster sessionAll students
Tue, Dec 15Final project report due

Grading

WeightComponent
30%Role-playing paper discussion (presenter / archaeologist / hacker)
15%Paper review assignments
15%Mid proposal & report
20%Final presentation
20%Final report