Same Transformer — different modality
Patch · mel frame · latent → predict next unit
224×224 → 196 patches → Transformer → class
HuggingFace AutoModelForImageClassification
SAM: click → mask (IoU) · SD: text → denoise (FID)
whisper.transcribe() · encoder-decoder
$(S+D+I)/N$ · jiwer · <5% = strong
Encoder → mel → vocoder · MOS eval
Image + text → fused tokens → answer
IoU · FID · WER · MOS · MMMU
ViT + Whisper + TTS + VLM + agent tools
Scroll: course-full-multimodal.html#multimodal-practice