tlwu commited on Dec 21, 2023

Commit

638944e

1 Parent(s): 7646089

models from Olive

Browse files

Files changed (21) hide show

ORT_CUDA/sd-xl-base-1.0/engine/clip2.ort_cuda.fp16/model.onnx.data +0 -3
ORT_CUDA/sd-xl-refiner-1.0/engine/unetxl.ort_cuda.fp16/model.onnx.data +0 -3
ORT_CUDA/sd-xl-refiner-1.0/engine/vae.ort_cuda.fp16/model.onnx +0 -3
README.md +13 -13
model_index.json +46 -0
scheduler/scheduler_config.json +21 -0
{ORT_CUDA/sd-xl-base-1.0/engine/clip.ort_cuda.fp16 → text_encoder}/model.onnx +2 -2
{ORT_CUDA/sd-xl-base-1.0/engine/unetxl.ort_cuda.fp16 → text_encoder_2}/model.onnx +2 -2
{ORT_CUDA/sd-xl-refiner-1.0/engine/clip2.ort_cuda.fp16 → text_encoder_2}/model.onnx.data +2 -2
tokenizer/merges.txt +0 -0
tokenizer/special_tokens_map.json +24 -0
tokenizer/tokenizer_config.json +30 -0
tokenizer/vocab.json +0 -0
tokenizer_2/merges.txt +0 -0
tokenizer_2/special_tokens_map.json +24 -0
tokenizer_2/tokenizer_config.json +38 -0
tokenizer_2/vocab.json +0 -0
{ORT_CUDA/sd-xl-refiner-1.0/engine/unetxl.ort_cuda.fp16 → unet}/model.onnx +2 -2
{ORT_CUDA/sd-xl-base-1.0/engine/unetxl.ort_cuda.fp16 → unet}/model.onnx.data +1 -1
{ORT_CUDA/sd-xl-base-1.0/engine/clip2.ort_cuda.fp16 → vae_decoder}/model.onnx +2 -2
{ORT_CUDA/sd-xl-refiner-1.0/engine/clip2.ort_cuda.fp16 → vae_encoder}/model.onnx +2 -2

ORT_CUDA/sd-xl-base-1.0/engine/clip2.ort_cuda.fp16/model.onnx.data DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:f928e0fb8826f641d36a14760546caa23d97366db4c15de1a5188802bd21e97e
-size 1389319680

ORT_CUDA/sd-xl-refiner-1.0/engine/unetxl.ort_cuda.fp16/model.onnx.data DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:94919ef99ac87418ad84b94ddc8581fc330390b3e3fd8839eee649b630eda78f
-size 4519331328

ORT_CUDA/sd-xl-refiner-1.0/engine/vae.ort_cuda.fp16/model.onnx DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:bbcf032f2aef098c298c82cff7e66160a8bc0937b99943914e8f75575a3a50fe
-size 99070466

README.md CHANGED Viewed

@@ -18,6 +18,11 @@ tags:
 This repository hosts the optimized versions of **Stable Diffusion XL 1.0** to accelerate inference with ONNX Runtime CUDA execution provider.
 See the [usage instructions](#usage-example) for how to run the SDXL pipeline with the ONNX files hosted in this repository.
 ## Model Description
@@ -35,16 +40,12 @@ The VAE decoder is converted from [sdxl-vae-fp16-fix](https://huggingface.co/mad
 Below is average latency of generating an image of size 1024x1024 using NVIDIA A100-SXM4-80GB GPU:
-| Engine      | Batch Size | PyTorch 2.1    | ONNX Runtime CUDA |
-|-------------|------------|----------------|-------------------|
-| Static      | 1          | N/A            | 3389 ms           |
-| Static      | 4          | N/A            | 12264 ms          |
-| Dynamic     | 1          | 3779 ms        | 3458 ms           |
-| Dynamic     | 4          | 13504 ms       | 12347 ms          |
-Static means the engine is built for the given batch size and image size combination, and CUDA graph is used to speed up.
-Dynamic means the engine is built to support dynamic batch size and image sizes.
 ## Usage Example
@@ -82,6 +83,8 @@ sh build.sh --config Release  --build_shared_lib --parallel --use_cuda --cuda_ve
 python3 -m pip install build/Linux/Release/dist/onnxruntime_gpu-*-cp310-cp310-linux_x86_64.whl --force-reinstall
 ```
 5. Install libraries and requirements
 ```shell
 python3 -m pip install --upgrade pip
@@ -94,8 +97,5 @@ python3 -m pip install --upgrade polygraphy onnx-graphsurgeon --extra-index-url
 ```shell
 python3 demo_txt2img_xl.py \
   "starry night over Golden Gate Bridge by van gogh" \
-  --width 1024         \
-  --height 1024        \
-  --denoising-steps 8  \
-  --work-dir /workspace/stable-diffusion-xl-1.0-onnxruntime
 ```

 This repository hosts the optimized versions of **Stable Diffusion XL 1.0** to accelerate inference with ONNX Runtime CUDA execution provider.
+The models are generated by [Olive](https://github.com/microsoft/Olive/tree/main/examples/stable_diffusion) with command like the following:
+```
+python stable_diffusion_xl.py --provider cuda --optimize --use_fp16_fixed_vae
+```
 See the [usage instructions](#usage-example) for how to run the SDXL pipeline with the ONNX files hosted in this repository.
 ## Model Description
 Below is average latency of generating an image of size 1024x1024 using NVIDIA A100-SXM4-80GB GPU:
+| Batch Size | PyTorch 2.1    | ONNX Runtime CUDA |
+|------------|----------------|-------------------|
+| 1          | 3779 ms        | 3389 ms           |
+| 4          | 13504 ms       | 12264 ms          |
+In this test, CUDA graph was used to speed up in both torch compile the unet and ONNX Runtime.
 ## Usage Example
 python3 -m pip install build/Linux/Release/dist/onnxruntime_gpu-*-cp310-cp310-linux_x86_64.whl --force-reinstall
 ```
+If the GPU is not A100, change CMAKE_CUDA_ARCHITECTURES=80 in the command line according to the GPU compute capacity (like 89 for RTX 4090, or 86 for RTX 3090). If your machine has less than 64GB memory, replace --parallel by --parallel 4 --nvcc_threads 1  to avoid out of memory.
 5. Install libraries and requirements
 ```shell
 python3 -m pip install --upgrade pip
 ```shell
 python3 demo_txt2img_xl.py \
   "starry night over Golden Gate Bridge by van gogh" \
+  --engine-dir /workspace/stable-diffusion-xl-1.0-onnxruntime
 ```

model_index.json ADDED Viewed

	@@ -0,0 +1,46 @@

+{
+  "_class_name": "ORTStableDiffusionXLPipeline",
+  "_diffusers_version": "0.24.0",
+  "_name_or_path": "stabilityai/stable-diffusion-xl-base-1.0",
+  "feature_extractor": [
+    null,
+    null
+  ],
+  "force_zeros_for_empty_prompt": true,
+  "image_encoder": [
+    null,
+    null
+  ],
+  "scheduler": [
+    "diffusers",
+    "EulerDiscreteScheduler"
+  ],
+  "text_encoder": [
+    "diffusers",
+    "OnnxRuntimeModel"
+  ],
+  "text_encoder_2": [
+    "diffusers",
+    "OnnxRuntimeModel"
+  ],
+  "tokenizer": [
+    "transformers",
+    "CLIPTokenizer"
+  ],
+  "tokenizer_2": [
+    "transformers",
+    "CLIPTokenizer"
+  ],
+  "unet": [
+    "diffusers",
+    "OnnxRuntimeModel"
+  ],
+  "vae_decoder": [
+    "diffusers",
+    "OnnxRuntimeModel"
+  ],
+  "vae_encoder": [
+    "diffusers",
+    "OnnxRuntimeModel"
+  ]
+}

scheduler/scheduler_config.json ADDED Viewed

	@@ -0,0 +1,21 @@

+{
+  "_class_name": "EulerDiscreteScheduler",
+  "_diffusers_version": "0.24.0",
+  "beta_end": 0.012,
+  "beta_schedule": "scaled_linear",
+  "beta_start": 0.00085,
+  "clip_sample": false,
+  "interpolation_type": "linear",
+  "num_train_timesteps": 1000,
+  "prediction_type": "epsilon",
+  "sample_max_value": 1.0,
+  "set_alpha_to_one": false,
+  "sigma_max": null,
+  "sigma_min": null,
+  "skip_prk_steps": true,
+  "steps_offset": 1,
+  "timestep_spacing": "leading",
+  "timestep_type": "discrete",
+  "trained_betas": null,
+  "use_karras_sigmas": false
+}

{ORT_CUDA/sd-xl-base-1.0/engine/clip.ort_cuda.fp16 → text_encoder}/model.onnx RENAMED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:80f597ccfe89500810537ce191b2501c0a2d832ac00e44aa307ff48da17f4ef8
-size 246165434

 version https://git-lfs.github.com/spec/v1
+oid sha256:661f552b00c76982e4a8e5d7a8814c33ff9354fb0298666ec875a1762f6d5076
+size 246178359

{ORT_CUDA/sd-xl-base-1.0/engine/unetxl.ort_cuda.fp16 → text_encoder_2}/model.onnx RENAMED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:3b65c5441a03b2e2bbbc246579f223b20375a7fc86391d17d71d9d263a792921
-size 705183

 version https://git-lfs.github.com/spec/v1
+oid sha256:6b9f04f2d71f0ae8cbb085a78e1e65e2509607d694a84c15dde6de1ce2db58e0
+size 1389427378

{ORT_CUDA/sd-xl-refiner-1.0/engine/clip2.ort_cuda.fp16 → text_encoder_2}/model.onnx.data RENAMED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:f928e0fb8826f641d36a14760546caa23d97366db4c15de1a5188802bd21e97e
-size 1389319680

 version https://git-lfs.github.com/spec/v1
+oid sha256:3da7ac65349fbd092e836e3eeca2c22811317bc804fd70af157b4550f2d4bcb5
+size 2778639360

tokenizer/merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer/special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,24 @@

+{
+  "bos_token": {
+    "content": "<|startoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": "<|endoftext|>",
+  "unk_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer/tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,30 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "49406": {
+      "content": "<|startoftext|>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "49407": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "<|startoftext|>",
+  "clean_up_tokenization_spaces": true,
+  "do_lower_case": true,
+  "eos_token": "<|endoftext|>",
+  "errors": "replace",
+  "model_max_length": 77,
+  "pad_token": "<|endoftext|>",
+  "tokenizer_class": "CLIPTokenizer",
+  "unk_token": "<|endoftext|>"
+}

tokenizer/vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_2/merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_2/special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,24 @@

+{
+  "bos_token": {
+    "content": "<|startoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": "!",
+  "unk_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer_2/tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,38 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "!",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "49406": {
+      "content": "<|startoftext|>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "49407": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "<|startoftext|>",
+  "clean_up_tokenization_spaces": true,
+  "do_lower_case": true,
+  "eos_token": "<|endoftext|>",
+  "errors": "replace",
+  "model_max_length": 77,
+  "pad_token": "!",
+  "tokenizer_class": "CLIPTokenizer",
+  "unk_token": "<|endoftext|>"
+}

tokenizer_2/vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff

{ORT_CUDA/sd-xl-refiner-1.0/engine/unetxl.ort_cuda.fp16 → unet}/model.onnx RENAMED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:af07cf55c9f0dde9f75ed66a3973cbe035af9411758ebf355cde79e198f58ab9
-size 531378

 version https://git-lfs.github.com/spec/v1
+oid sha256:de912f37c6c90e43bab63adf89515d76290d1eb60208b9a642795e347ff3701e
+size 736952

{ORT_CUDA/sd-xl-base-1.0/engine/unetxl.ort_cuda.fp16 → unet}/model.onnx.data RENAMED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:fc7a0e1b7d003e16c6797acff972ebae520de529303f93f5d318356679a381dc
 size 5135092480

 version https://git-lfs.github.com/spec/v1
+oid sha256:7cdae162dbf695bc3abb40653265bdabad17e1d9eef7c4e44beeb2a834b70cdc
 size 5135092480

{ORT_CUDA/sd-xl-base-1.0/engine/clip2.ort_cuda.fp16 → vae_decoder}/model.onnx RENAMED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:eed98f933a93c2f754d075c0a4d64a5039dfa713cfd3c180037d201f0f3be20d
-size 134411

 version https://git-lfs.github.com/spec/v1
+oid sha256:7987d20deef6934d7d30bd7486da698940765d5383a5ca009f0aad74c737ec70
+size 99072671

{ORT_CUDA/sd-xl-refiner-1.0/engine/clip2.ort_cuda.fp16 → vae_encoder}/model.onnx RENAMED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:f3504a45acc2b5e8b08c8dc4680a8f7eed54fc397c94c7a2dba6e817a0770225
-size 134410

 version https://git-lfs.github.com/spec/v1
+oid sha256:a56f9f96a763bc9995d032d6e03159cf433569047488e7594f0b15066cbed44f
+size 68412330