12 61 6

Ming Li

limingcv

https://liming-ai.github.io

liming-ai

AI & ML interests

Computer Vision, AIGC, VLM/LLM

Recent Activity

upvoted a paper 4 days ago

Cosmos World Foundation Model Platform for Physical AI

new activity about 1 month ago

THUDM/CogVideoX1.5-5B-I2V:Could you provide the image used for the demo in README?

upvoted a paper about 2 months ago

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

View all activity

Organizations

limingcv's activity

upvoted a paper 4 days ago

Cosmos World Foundation Model Platform for Physical AI

Paper • 2501.03575 • Published 5 days ago • 52

upvoted a paper about 2 months ago

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Paper • 2411.10442 • Published Nov 15, 2024 • 71

upvoted a paper 2 months ago

GPT-4o System Card

Paper • 2410.21276 • Published Oct 25, 2024 • 83

upvoted a paper 3 months ago

Movie Gen: A Cast of Media Foundation Models

Paper • 2410.13720 • Published Oct 17, 2024 • 91

upvoted 2 papers 4 months ago

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models

Paper • 2409.17146 • Published Sep 25, 2024 • 106

OmniGen: Unified Image Generation

Paper • 2409.11340 • Published Sep 17, 2024 • 109

upvoted 2 papers 5 months ago

Imagen 3

Paper • 2408.07009 • Published Aug 13, 2024 • 61

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Paper • 2408.06072 • Published Aug 12, 2024 • 37

upvoted 2 papers 6 months ago

Qwen2-Audio Technical Report

Paper • 2407.10759 • Published Jul 15, 2024 • 55

Qwen2 Technical Report

Paper • 2407.10671 • Published Jul 15, 2024 • 160

upvoted 5 papers 7 months ago

Depth Anything V2

Paper • 2406.09414 • Published Jun 13, 2024 • 95

An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels

Paper • 2406.09415 • Published Jun 13, 2024 • 50

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Paper • 2406.06525 • Published Jun 10, 2024 • 66

Step-aware Preference Optimization: Aligning Preference with Denoising Performance at Each Step

Paper • 2406.04314 • Published Jun 6, 2024 • 28

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Paper • 2406.04325 • Published Jun 6, 2024 • 73

upvoted 4 papers 8 months ago

Phased Consistency Model

Paper • 2405.18407 • Published May 28, 2024 • 46

An Introduction to Vision-Language Modeling

Paper • 2405.17247 • Published May 27, 2024 • 87

Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Paper • 2405.01535 • Published May 2, 2024 • 120

LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report

Paper • 2405.00732 • Published Apr 29, 2024 • 119

upvoted a paper 9 months ago

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Paper • 2404.16821 • Published Apr 25, 2024 • 55