This course provides a deep dive into modern Multimodal AI, focusing on models that integrate vision and language data. Topics include visual question answering, generative media (images/video), neuro-symbolic AI, and embodied AI agents. Through weekly paper reviews and a hands-on independent research project, students will gain the technical skills to understand and advance the state-of-the-art in multimodal deep learning.
Required Course Background: at least one upper-level/grad course in vision, NLP or machine learning.
Paper discussion
Each topic day covers 2–3 papers. Everyone reads all of that day's papers before class. Every student holds three roles over the semester:
Presenter (1–2×)
Present one paper as part of that day's team — method, results, and answers to the class's submitted questions. 45 min shared across the day's papers.
Archaeologist (≤1×)
A 10-minute survey of the day's papers' ancestors and successors — what came before, what built on them, where they sit.
Hacker (≤1×)
A 10-minute Colab implementing a small piece of one paper on a toy problem, walked through live.
Reader (everyone, every class)
Submit a one-page review of one paper (guideline below) via Gradescope — every class, including days you present, archaeologist, or hack. Due 11:59 pm EST: previous Wednesday for Monday classes, previous Friday for Wednesday classes. No late submissions; your three lowest-scored reviews are dropped.
All submissions — paper reviews, project abstract, midpoint and final reports — go through Gradescope (entry code: B4P7WR), in PDF format — reviews, reports, and presentation slides alike.
After each class, the day's presenters, archaeologist, and hacker upload their materials (slides as PDF / Colab link) to the shared OneDrive folder — the link will be shared after the first class. Course announcements and questions happen on MS Teams; role assignments and Teams invitations will also go out after the first class.
Paper review guideline
One page per review, in four sections:
- Summary — what the paper does and claims
- Strengths
- Weaknesses
- Questions for presenters
After each deadline, the instructor compiles the submitted reviews and shares them on MS Teams — presenters address the questions during their session.
Research project
This course is heavily project-based: the project accounts for more than half of your grade. Teams of 2 recommended; any topic within the course's scope.
Deliverables: proposal · mid report · mid presentation · final poster session · final report.
A full-score final project is expected to be comparable in quality to a paper at a top AI conference (NeurIPS, CVPR, ICML, ICLR, ICCV, ACL, COLM, etc.).
Mid and final reports follow the NeurIPS 2026 format. Page limits — mid: 4 pages, final: 8 pages — including all main text, figures, and tables; references excluded.
Sep 16 — proposal
One-pager per team, due 11:59 pm:
- A 200–300 word project abstract
- Planned mid deliverables
- Planned final deliverables
- Team members and external collaborators (if any)
Only one submission per team is needed.
Oct 19 — mid report & presentation
All projects present in class; mid report in NeurIPS 2026 format, up to 4 pages.
Dec 9 — final poster session
All teams, in class.
Dec 15 — final report
NeurIPS 2026 format, up to 8 pages.
Schedule
All deadlines are 11:59 pm EST — reviews are due the previous Wednesday for Monday classes, and the previous Friday for Wednesday classes.
| # | Date | Review due | Topic | Presenter | Archaeologist | Hacker | Papers |
|---|---|---|---|---|---|---|---|
| 1 | Mon, Aug 31 | Course introduction — logistics, grading, roles, project expectations, role sign-up | Jaemin Cho | ||||
| 2 | Wed, Sep 2 | Understanding: Foundations (vision-language pretraining) | Jaemin Cho | Jaemin Cho | Jaemin Cho | ||
| Mon, Sep 7 | No class — Labor Day | ||||||
| 3 | Wed, Sep 9 | Guest lecture | Yushi Hu (Meta Superintelligence Labs) | ||||
| 4 | Mon, Sep 14 | Wed, Sep 9 | Generation: Text-to-image foundations | TBD | TBD | TBD |
|
| 5 | Wed, Sep 16 | Fri, Sep 11 | Understanding: Text-to-image/video grounding | TBD | TBD | TBD |
|
| Wed, Sep 16 | Project abstract due, 11:59 pm — one page per team | ||||||
| 6 | Mon, Sep 21 | Wed, Sep 16 | Understanding: Modern vision-language models | TBD | TBD | TBD | |
| 7 | Wed, Sep 23 | Fri, Sep 18 | Generation: Modern text-to-image architectures | TBD | TBD | TBD | |
| 8 | Mon, Sep 28 | Wed, Sep 23 | Generation: Evaluation of image/video generators | TBD | TBD | TBD | |
| 9 | Wed, Sep 30 | Fri, Sep 25 | Understanding: Evaluation of multimodal understanding models | TBD | TBD | TBD | |
| 10 | Mon, Oct 5 | Wed, Sep 30 | Efficient adaptation of understanding/generation models | TBD | TBD | TBD | |
| 11 | Wed, Oct 7 | Fri, Oct 2 | Generation: Editing and controlled generation | TBD | TBD | TBD | |
| 12 | Mon, Oct 12 | Wed, Oct 7 | Generation: Video generation | TBD | TBD | TBD | |
| 13 | Wed, Oct 14 | Fri, Oct 9 | Generation: Motion and music | TBD | TBD | TBD | |
| 14 | Mon, Oct 19 | Mid-project presentation — all projects | All students | ||||
| 15 | Wed, Oct 21 | Fri, Oct 16 | Understanding: Neuro-symbolic | TBD | TBD | TBD | |
| 16 | Mon, Oct 26 | Wed, Oct 21 | Generation: Neuro-symbolic | TBD | TBD | TBD | |
| 17 | Wed, Oct 28 | Fri, Oct 23 | Embodied AI: Vision-language-action models | TBD | TBD | TBD | |
| 18 | Mon, Nov 2 | Wed, Oct 28 | Embodied AI: Learning actions from videos | TBD | TBD | TBD |
|
| 19 | Wed, Nov 4 | Fri, Oct 30 | Embodied AI: Spatial understanding and gaze in video | TBD | TBD | TBD | |
| 20 | Mon, Nov 9 | Wed, Nov 4 | Advanced: Unified understanding and generation | TBD | TBD | TBD | |
| 21 | Wed, Nov 11 | Fri, Nov 6 | Advanced: Long video understanding | TBD | TBD | TBD | |
| 22 | Mon, Nov 16 | Wed, Nov 11 | Advanced: Visual prompting and GUI agents | TBD | TBD | TBD | |
| 23 | Wed, Nov 18 | Fri, Nov 13 | Advanced: World models | TBD | TBD | TBD | |
| Mon, Nov 23 | No class — Thanksgiving | ||||||
| Wed, Nov 25 | No class — Thanksgiving | ||||||
| 24 | Mon, Nov 30 | Wed, Nov 25 | Advanced: Multimodal rewards and feedback | TBD | TBD | TBD |
|
| 25 | Wed, Dec 2 | Guest lecture | TBD | ||||
| 26 | Mon, Dec 7 | Guest lecture | TBD | ||||
| 27 | Wed, Dec 9 | Final project poster session | All students | ||||
| Tue, Dec 15 | Final project report due |
Grading
| Weight | Component |
|---|---|
| 30% | Role-playing paper discussion (presenter / archaeologist / hacker) |
| 15% | Paper review assignments |
| 15% | Mid proposal & report |
| 20% | Final presentation |
| 20% | Final report |