daily papers - a jmkim0309 Collection

jmkim0309 's Collections

daily papers

updated 2 days ago

GenTron: Delving Deep into Diffusion Transformers for Image and Video Generation

Paper • 2312.04557 • Published Dec 7, 2023 • 13
Smooth Diffusion: Crafting Smooth Latent Spaces in Diffusion Models

Paper • 2312.04410 • Published Dec 7, 2023 • 15
PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding

Paper • 2312.04461 • Published Dec 7, 2023 • 62
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively

Paper • 2401.02955 • Published Jan 5, 2024 • 22
Denoising Vision Transformers

Paper • 2401.02957 • Published Jan 5, 2024 • 29
SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven Generation

Paper • 2312.16272 • Published Dec 26, 2023 • 7
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion

Paper • 2312.16486 • Published Dec 27, 2023 • 7
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models

Paper • 2411.07126 • Published Nov 11, 2024 • 29
Motion Control for Enhanced Complex Action Video Generation

Paper • 2411.08328 • Published Nov 13, 2024 • 5
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Paper • 2411.07975 • Published Nov 12, 2024 • 30
Pyramidal Flow Matching for Efficient Video Generative Modeling

Paper • 2410.05954 • Published Oct 8, 2024 • 39
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

Paper • 2412.04432 • Published Dec 5, 2024 • 16
LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Paper • 2412.04814 • Published Dec 6, 2024 • 46
Mind the Time: Temporally-Controlled Multi-Event Video Generation

Paper • 2412.05263 • Published Dec 6, 2024 • 11
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Paper • 2412.01169 • Published Dec 2, 2024 • 13
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation

Paper • 2410.13861 • Published Oct 17, 2024 • 53
UIP2P: Unsupervised Instruction-based Image Editing via Cycle Edit Consistency

Paper • 2412.15216 • Published Dec 19, 2024 • 5
MotiF: Making Text Count in Image Animation with Motion Focal Loss

Paper • 2412.16153 • Published Dec 20, 2024 • 6
Large Motion Video Autoencoding with Cross-modal Video VAE

Paper • 2412.17805 • Published Dec 23, 2024 • 24
AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation

Paper • 2501.09503 • Published Jan 16 • 13
Do generative video models learn physical principles from watching videos?

Paper • 2501.09038 • Published Jan 14 • 32
Small Models Struggle to Learn from Strong Reasoners

Paper • 2502.12143 • Published 5 days ago • 22