This course provides a deep dive into modern Multimodal AI, focusing on models that integrate vision and language data. Topics include visual question answering, generative media (images/video), neuro-symbolic AI, and embodied AI agents. Through weekly paper reviews and a hands-on independent research project, students will gain the technical skills to understand and advance the state-of-the-art in multimodal deep learning.
Required Course Background: at least one upper-level/grad course in vision, NLP or machine learning.
Paper discussion
Each topic day covers 2–3 papers. Everyone reads all of that day's papers before class. Every student holds three roles over the semester:
Presenter (1–2×)
Present one or two of the day's papers — method, results, and answers to the class's submitted questions. 45 min shared across the day's papers, usually with a co-presenter; a couple of tightly paired sessions are covered by one presenter.
Archaeologist (≤1×)
A 10-minute survey of the day's papers' ancestors and successors — what came before, what built on them, where they sit.
Hacker (≤1×)
A 10-minute Colab implementing a small piece of one paper on a toy problem, walked through live.
Reader (everyone, every class)
Submit a one-page review of one paper (guideline below) via Gradescope — every class, including days you present, archaeologist, or hack. Due 11:59 pm EST: previous Wednesday for Monday classes, previous Friday for Wednesday classes. No late submissions; your three lowest-scored reviews are dropped.
All submissions — paper reviews, project abstract, midpoint and final reports — go through Gradescope (entry code: B4P7WR), in PDF format — reviews, reports, and presentation slides alike.
After each class, the day's presenters, archaeologist, and hacker upload their materials (slides as PDF / Colab link) to the shared OneDrive folder — the link will be shared after the first class. Course announcements and questions happen on MS Teams; role assignments and Teams invitations will also go out after the first class.
Paper review guideline
One page per review, in four sections:
- Summary — what the paper does and claims
- Strengths
- Weaknesses
- Questions for presenters
After each deadline, the instructor compiles the submitted reviews and shares them on MS Teams — presenters address the questions during their session.
Research project
This course is heavily project-based: the project accounts for more than half of your grade. Teams of 2 recommended; any topic within the course's scope.
Deliverables: proposal · mid report · mid presentation · final poster session · final report.
A full-score final project is expected to be comparable in quality to a paper at a top AI conference (NeurIPS, CVPR, ICML, ICLR, ICCV, ACL, COLM, etc.).
Mid and final reports follow the NeurIPS 2026 format. Page limits — mid: 4 pages, final: 8 pages — including all main text, figures, and tables; references excluded.
Sep 16 — proposal
One-pager per team, due 11:59 pm:
- A 200–300 word project abstract
- Planned mid deliverables
- Planned final deliverables
- Team members and external collaborators (if any)
Only one submission per team is needed.
Oct 19 — mid report & presentation
All projects present in class; mid report in NeurIPS 2026 format, up to 4 pages.
Dec 9 — final poster session
All teams, in class.
Dec 15 — final report
NeurIPS 2026 format, up to 8 pages.
Schedule
Role assignments below are tentative and may change until sign-up closes on Teams. A cell reading TBD is an unfilled seat. If you would prefer your name not appear on this public page, email the instructor and it will be removed.
All deadlines are 11:59 pm EST — reviews are due the previous Wednesday for Monday classes, and the previous Friday for Wednesday classes.
| # | Date | Review due | Topic | Presenter | Archaeologist | Hacker | Papers |
|---|---|---|---|---|---|---|---|
| 1 | Mon, Aug 31 | Course introduction — logistics, grading, roles, project expectations, role sign-up | Jaemin Cho | ||||
| 2 | Wed, Sep 2 | Understanding: Foundations (vision-language pretraining) | Jaemin Cho | Jaemin Cho | Jaemin Cho | ||
| Mon, Sep 7 | No class — Labor Day | ||||||
| 3 | Wed, Sep 9 | Guest lecture | Yushi Hu (Meta Superintelligence Labs) | ||||
| 4 | Mon, Sep 14 | Wed, Sep 9 | Generation: Text-to-image foundations | Asma Akel Alkhaldi · Jiwoo Noh | Lurui Wang | Jaemin Cho |
|
| 5 | Wed, Sep 16 | Fri, Sep 11 | Understanding: Text-to-image/video grounding | Lukas William Geer · Ram Sai Ganesh | Jaemin Cho | Zhanpeng Luo | |
| Wed, Sep 16 | Project abstract due, 11:59 pm — one page per team | ||||||
| 6 | Mon, Sep 21 | Wed, Sep 16 | Understanding: Modern vision-language models | Jingyang Chen · Shreyas Phegade | Lin Long | Matt Wang | |
| 7 | Wed, Sep 23 | Fri, Sep 18 | Generation: Modern text-to-image architectures | Tengyang Deng | Ram Sai Ganesh | Asma Akel Alkhaldi | |
| 8 | Mon, Sep 28 | Wed, Sep 23 | Generation: Evaluation of image/video generators | Ruijie Tao · Zijian Zhang | Ruijie Tao | Shenghan Zhou | |
| 9 | Wed, Sep 30 | Fri, Sep 25 | Understanding: Evaluation of multimodal understanding models | Kehao Xu | Zhanpeng Luo | Katherine M Guerrerio | |
| 10 | Mon, Oct 5 | Wed, Sep 30 | Efficient adaptation of understanding/generation models | Adway U Kanhere · Junhyeok Lee | — | Lurui Wang | |
| 11 | Wed, Oct 7 | Fri, Oct 2 | Generation: Editing and controlled generation | Katherine M Guerrerio · Xuyang Wang | Jiayin Li | Ruijie Tao | |
| 12 | Mon, Oct 12 | Wed, Oct 7 | Generation: Video generation | Zhanpeng Luo | Adway U Kanhere | Lin Long | |
| 13 | Wed, Oct 14 | Fri, Oct 9 | Generation: Motion and music | Qi Chen · Shenghan Zhou | Lukas William Geer | Shreyas Phegade | |
| 14 | Mon, Oct 19 | Mid-project presentation — all projects | All students | ||||
| 15 | Wed, Oct 21 | Fri, Oct 16 | Understanding: Neuro-symbolic | Qi Chen | Qi Chen | Jiayin Li | |
| 16 | Mon, Oct 26 | Wed, Oct 21 | Generation: Neuro-symbolic | Chenyu Li · Matt Wang | — | Kehao Xu | |
| 17 | Wed, Oct 28 | Fri, Oct 23 | Embodied AI: Vision-language-action models | Lin Long | Shreyas Phegade | Xuyang Wang |
|
| 18 | Mon, Nov 2 | Wed, Oct 28 | Embodied AI: Learning actions from videos | Junhyeok Lee · Shenghan Zhou | — | Jiwoo Noh | |
| 19 | Wed, Nov 4 | Fri, Oct 30 | Embodied AI: Spatial understanding and gaze in video | Lurui Wang | Tengyang Deng | Lukas William Geer | |
| 20 | Mon, Nov 9 | Wed, Nov 4 | Advanced: Unified understanding and generation | Chenyu Li · Jiayin Li | — | Kehao Xu | |
| 21 | Wed, Nov 11 | Fri, Nov 6 | Advanced: Long video understanding | Asma Akel Alkhaldi · Jingyang Chen | — | Zijian Zhang | |
| 22 | Mon, Nov 16 | Wed, Nov 11 | Advanced: Computer-Use Agents | Jiwoo Noh · Ram Sai Ganesh | — | Adway U Kanhere | |
| 23 | Wed, Nov 18 | Fri, Nov 13 | Advanced: World models | Matt Wang · Zijian Zhang | Jingyang Chen | Chenyu Li | |
| Mon, Nov 23 | No class — Thanksgiving | ||||||
| Wed, Nov 25 | No class — Thanksgiving | ||||||
| 24 | Mon, Nov 30 | Wed, Nov 25 | Advanced: Multimodal rewards and feedback | Katherine M Guerrerio · Xuyang Wang | — | Junhyeok Lee |
|
| 25 | Wed, Dec 2 | Guest lecture | Keunhong Park (World Labs) | ||||
| 26 | Mon, Dec 7 | Guest lecture | TBD | ||||
| 27 | Wed, Dec 9 | Final project poster session | All students | ||||
| Tue, Dec 15 | Final project report due |
Grading
| Weight | Component |
|---|---|
| 30% | Role-playing paper discussion (presenter / archaeologist / hacker) |
| 15% | Paper review assignments |
| 15% | Mid proposal & report |
| 20% | Final presentation |
| 20% | Final report |