LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Paper • 2407.07895 • Published Jul 10 • 40
TextToon: Real-Time Text Toonify Head Avatar from Single Video Paper • 2410.07160 • Published Sep 23 • 8