PixArt

community

https://pixart-alpha.github.io

PixArt-alpha

Activity Feed

AI & ML interests

None defined yet.

Recent Activity

sayakpaul authored a paper 27 days ago

A Noise is Worth Diffusion Guidance

Lawrence-cj authored a paper about 1 month ago

MetaBEV: Solving Sensor Failures for BEV Detection and Map Segmentation

Lawrence-cj authored a paper about 1 month ago

HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

View all activity

PixArt-alpha's activity

sayakpaul

posted an update 9 days ago

Post

3694

Commits speak louder than words 🤪

* 4 new video models
* Multiple image models, including SANA & Flux Control
* New quantizers -> GGUF & TorchAO
* New training scripts

Enjoy this holiday-special Diffusers release 🤗
Notes: https://github.com/huggingface/diffusers/releases/tag/v0.32.0

sayakpaul

posted an update 15 days ago

Post

1685

In the past seven days, the Diffusers team has shipped:

1. Two new video models
2. One new image model
3. Two new quantization backends
4. Three new fine-tuning scripts
5. Multiple fixes and library QoL improvements

Coffee on me if someone can guess 1 - 4 correctly.

1 reply

sayakpaul

posted an update 23 days ago

Post

2065

Introducing a high-quality open-preference dataset to further this line of research for image generation.

Despite being such an inseparable component for modern image generation, open preference datasets are a rarity!

So, we decided to work on one with the community!

Check it out here:
https://huggingface.co/blog/image-preferences

7 replies

sayakpaul

posted an update 24 days ago

Post

2113

The Control family of Flux from @black-forest-labs should be discussed more!

It enables structural controls like ControlNets while being significantly less expensive to run!

So, we're working on a Control LoRA training script 🤗

It's still WIP, so go easy:
https://github.com/huggingface/diffusers/pull/10130

sayakpaul

authored a paper 27 days ago

A Noise is Worth Diffusion Guidance

Paper • 2412.03895 • Published 28 days ago • 28

sayakpaul

posted an update about 1 month ago

Post

1479

Let 2024 be the year of video model fine-tunes!

Check it out here:
https://github.com/a-r-r-o-w/cogvideox-factory/tree/main/training/mochi-1

Lawrence-cj

authored 3 papers about 1 month ago

MetaBEV: Solving Sensor Failures for BEV Detection and Map Segmentation

Paper • 2304.09801 • Published Apr 19, 2023

HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Paper • 2410.10812 • Published Oct 14, 2024 • 16

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Paper • 2410.10629 • Published Oct 14, 2024 • 9

sayakpaul

posted an update about 2 months ago

Post

2616

It's been a while we shipped native quantization support in diffusers 🧨

We currently support bistandbytes as the official backend but using others like torchao is already very simple.

This post is just a reminder of what's possible:

1. Loading a model with a quantization config
2. Saving a model with quantization config
3. Loading a pre-quantized model
4. enable_model_cpu_offload()
5. Training and loading LoRAs into quantized checkpoints

Docs:
https://huggingface.co/docs/diffusers/main/en/quantization/bitsandbytes

1 reply

Lawrence-cj

updated a Space 3 months ago

Running on A10G

352

👀

Pixart-α

sayakpaul

posted an update 3 months ago

Post

2756

Did some little experimentation to resize pre-trained LoRAs on Flux. I explored two themes:

* Decrease the rank of a LoRA
* Increase the rank of a LoRA

The first one is helpful in reducing memory requirements if the LoRA is of a high rank, while the second one is merely an experiment. Another implication of this study is in the unification of LoRA ranks when you would like to torch.compile() them.

Check it out here:
sayakpaul/flux-lora-resizing

1 reply

Lawrence-cj

updated a Space 4 months ago

Runtime error

234

👁

PixArt Sigma

sayakpaul

authored a paper 4 months ago

LlamaDuo: LLMOps Pipeline for Seamless Migration from Service LLMs to Small-Scale Local LLMs

Paper • 2408.13467 • Published Aug 24, 2024 • 24

sayakpaul

posted an update 5 months ago

Post

2948

Here is a hackable and minimal implementation showing how to perform distributed text-to-image generation with Diffusers and Accelerate.

Full snippet is here: https://gist.github.com/sayakpaul/cfaebd221820d7b43fae638b4dfa01ba

With @JW17

sayakpaul

posted an update 5 months ago

Post

4481

Flux.1-Dev like images but in fewer steps.

Merging code (very simple), inference code, merged params: sayakpaul/FLUX.1-merged

Enjoy the Monday 🤗

4 replies

sayakpaul

posted an update 5 months ago

Post

3796

With larger and larger diffusion transformers coming up, it's becoming increasingly important to have some good quantization tools for them.

We present our findings from a series of experiments on quantizing different diffusion pipelines based on diffusion transformers.

We demonstrate excellent memory savings with a bit of sacrifice on inference latency which is expected to improve in the coming days.

Diffusers 🤝 Quanto ❤️

This was a juicy collaboration between @dacorvo and myself.

Check out the post to learn all about it
https://huggingface.co/blog/quanto-diffusers