> ## Documentation Index
> Fetch the complete documentation index at: https://dripart-claude-comfy-concurrency-limits-page-s84u33.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Qwen-Image-2.1 ComfyUI Native Workflow Example

> Run Qwen-Image-2.1 in ComfyUI: one model for image generation and editing, native 2K output, professional typography, and alpha channel support.

**Qwen-Image-2.1** is the latest open-weight release in the Qwen-Image series from Alibaba's Qwen team. A single model covers both text-to-image generation and instruction-based image editing, with native 2K output, professional typography, and an alpha channel for transparent backgrounds.

**Key Features**:

* **Generation and editing in one model**: the same weights serve text-to-image prompts and editing instructions, so a workflow does not need to swap checkpoints
* **Native 2K output**: generate at up to 2048x2048 directly instead of upscaling a smaller result
* **Professional typography**: dense small text and complex layouts hold up, for infographics, slides, UI mockups, posters, and packaging designs
* **Alpha channel support**: the VAE carries four channels, so transparent-background images can be generated and edited directly instead of being cut out afterwards
* **Multi-image editing**: reference images are spliced into the text encoder in slot order. The Text Encode Qwen Image 2.1 node accepts up to 16 slots (`image_1` to `image_16`); the two edit templates on this page wire the first 10 (`image_1` to `image_10`). `image_1` is the image being edited and the rest supply content, so the prompt addresses them by index, for example `image_1` is written as `<image1>`
* **Localized edits**: describe the object or region to change and the rest of the image is preserved

**Related Links**:

* [GitHub Repository](https://github.com/QwenLM/Qwen-Image)
* [Hugging Face (Comfy-Org/Qwen-Image-2.1)](https://huggingface.co/Comfy-Org/Qwen-Image-2.1)
* [Qwen Image 2.1 on Comfy](https://comfy.org/qwen-image-2.1/)

## Qwen-Image-2.1 workflow

<Tip>
  <Tabs>
    <Tab title="Local users">
      Make sure your ComfyUI is updated.

      * [Download ComfyUI](https://www.comfy.org/download)
      * [Update Guide](/installation/update_comfyui)

      Workflows in this guide can be found in the [Workflow Templates](/interface/features/template).
      If you can't find them in the template, your ComfyUI may be outdated.

      If nodes are missing when loading a workflow, possible reasons:

      1. You are not using the latest ComfyUI version (Nightly version)
      2. Some nodes failed to import at startup
    </Tab>

    <Tab title="Cloud users">
      * [Cloud](https://cloud.comfy.org) will update after ComfyUI stable release.

      So, if you find any core node missing in this document, it might be because the new core nodes have not yet been released in the latest stable version. Please wait for the next stable release.
    </Tab>
  </Tabs>
</Tip>

<h3 id="image_qwen_image_2_1_t2i">
  Qwen Image 2.1 Text to Image
</h3>

Generate an image from a text prompt at the aspect ratio and megapixel target you select.

<img src="https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/image_qwen_image_2_1_t2i-1.webp" alt="Qwen-Image-2.1 text to image workflow preview" />

<CardGroup cols={2}>
  <Card title="Download Workflow" icon="download" href="https://github.com/Comfy-Org/workflow_templates/blob/main/templates/image_qwen_image_2_1_t2i.json">
    Download JSON or search "Qwen-Image-2.1" in Template Library
  </Card>

  <Card title="Run on Comfy Cloud" icon="cloud" href="https://cloud.comfy.org/?template=image_qwen_image_2_1_t2i&utm_source=docs&utm_medium=referral&utm_campaign=qwen-image-2-1">
    Run ComfyUI online with zero setup
  </Card>
</CardGroup>

**Example output**

![Qwen-Image-2.1 text to image example output](https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/output/image_qwen_image_2_1_t2i.png)

<h3 id="image_qwen_image_2_1_image_edit">
  Qwen Image 2.1 Image Edit
</h3>

Edit an image with an instruction. Add reference images when the edit needs content that is not in the source image, such as putting a garment from a second photo onto the person in the first.

<img src="https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/image_qwen_image_2_1_image_edit-1.webp" alt="Qwen-Image-2.1 image edit workflow preview" />

<CardGroup cols={2}>
  <Card title="Download Workflow" icon="download" href="https://github.com/Comfy-Org/workflow_templates/blob/main/templates/image_qwen_image_2_1_image_edit.json">
    Download JSON or search "Qwen-Image-2.1" in Template Library
  </Card>

  <Card title="Run on Comfy Cloud" icon="cloud" href="https://cloud.comfy.org/?template=image_qwen_image_2_1_image_edit&utm_source=docs&utm_medium=referral&utm_campaign=qwen-image-2-1">
    Run ComfyUI online with zero setup
  </Card>
</CardGroup>

**Input materials**

Upload these files to the matching `LoadImage` nodes:

<CardGroup cols={2}>
  <Card title="portrait_model_denim.png" icon="image" href="https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/input/portrait_model_denim.png">
    `LoadImage` node 470 · `portrait_model_denim.png`
  </Card>

  <Card title="clothing_light_blue_denim_shirt.png" icon="image" href="https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/input/clothing_light_blue_denim_shirt.png">
    `LoadImage` node 475 · `clothing_light_blue_denim_shirt.png`
  </Card>
</CardGroup>

**Example output**

<div style={{display: 'grid', gridTemplateColumns: 'repeat(2, minmax(0, 1fr))', gap: '1rem', alignItems: 'start'}}>
  <img src="https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/input/portrait_model_denim.png" alt="Input image" style={{width: '100%', height: 'auto', objectFit: 'contain'}} />

  <img src="https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/output/image_qwen_image_2_1_image_edit.png" alt="Qwen-Image-2.1 image edit example output" style={{width: '100%', height: 'auto', objectFit: 'contain'}} />
</div>

<h3 id="image_qwen_image_2_1_background_removal">
  Remove Background: Qwen Image 2.1
</h3>

Remove the background from a photo with an edit instruction. The workflow reuses the image edit subgraph with the prompt `Remove the background, and output a PNG image`, then compares the result with the original.

<img src="https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/templates/image_qwen_image_2_1_background_removal-1.webp" alt="Qwen-Image-2.1 background removal workflow preview" />

<CardGroup cols={2}>
  <Card title="Download Workflow" icon="download" href="https://github.com/Comfy-Org/workflow_templates/blob/main/templates/image_qwen_image_2_1_background_removal.json">
    Download JSON or search "Qwen-Image-2.1" in Template Library
  </Card>

  <Card title="Run on Comfy Cloud" icon="cloud" href="https://cloud.comfy.org/?template=image_qwen_image_2_1_background_removal&utm_source=docs&utm_medium=referral&utm_campaign=qwen-image-2-1">
    Run ComfyUI online with zero setup
  </Card>
</CardGroup>

**Input materials**

Upload this file to the matching `LoadImage` node:

<Card title="angry_broccoli.png" icon="image" href="https://raw.githubusercontent.com/Comfy-Org/workflow_templates/main/input/angry_broccoli.png">
  `LoadImage` node 470 · `angry_broccoli.png`
</Card>

## Model links

The three workflows share the same diffusion model, VAE, and base text encoder. The text-to-image and image-edit templates each add a separate prompt-enhancement text encoder: `qwen3.5_9b_qwen_image_2.1_pe_t2i` for text to image and `qwen3.5_9b_qwen_image_2.1_pe_i2i` for image edit.

**text\_encoders**

* [qwen3vl\_8b\_int8\_convrot.safetensors](https://huggingface.co/Comfy-Org/Qwen-Image-2.1/blob/main/text_encoders/qwen3vl_8b_int8_convrot.safetensors) (loaded by the templates, lower memory)
* [qwen3vl\_8b\_bf16.safetensors](https://huggingface.co/Comfy-Org/Qwen-Image-2.1/blob/main/text_encoders/qwen3vl_8b_bf16.safetensors) (full precision, needs more memory)
* [qwen3.5\_9b\_qwen\_image\_2.1\_pe\_t2i.int8\_convrot.safetensors](https://huggingface.co/Comfy-Org/Qwen-Image-2.1/blob/main/text_encoders/qwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors) (prompt enhancement in the text to image template)
* [qwen3.5\_9b\_qwen\_image\_2.1\_pe\_i2i.int8\_convrot.safetensors](https://huggingface.co/Comfy-Org/Qwen-Image-2.1/blob/main/text_encoders/qwen3.5_9b_qwen_image_2.1_pe_i2i.int8_convrot.safetensors) (prompt enhancement in the image edit template)

**diffusion\_models**

* [qwen\_image\_2.1\_int8\_convrot.safetensors](https://huggingface.co/Comfy-Org/Qwen-Image-2.1/blob/main/diffusion_models/qwen_image_2.1_int8_convrot.safetensors) (loaded by the templates, lower memory)
* [qwen\_image\_2.1\_bf16.safetensors](https://huggingface.co/Comfy-Org/Qwen-Image-2.1/blob/main/diffusion_models/qwen_image_2.1_bf16.safetensors) (full precision, needs more memory)

**vae**

* [qwen\_image\_2.1\_vae\_bf16.safetensors](https://huggingface.co/Comfy-Org/Qwen-Image-2.1/blob/main/vae/qwen_image_2.1_vae_bf16.safetensors)

**Model Storage Location**

```
📂 ComfyUI/
├── 📂 models/
│   ├── 📂 text_encoders/
│   │      ├── qwen3vl_8b_int8_convrot.safetensors
│   │      ├── qwen3vl_8b_bf16.safetensors
│   │      ├── qwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors
│   │      └── qwen3.5_9b_qwen_image_2.1_pe_i2i.int8_convrot.safetensors
│   ├── 📂 diffusion_models/
│   │      ├── qwen_image_2.1_int8_convrot.safetensors
│   │      └── qwen_image_2.1_bf16.safetensors
│   └── 📂 vae/
│          └── qwen_image_2.1_vae_bf16.safetensors
```

## Workflow settings

### Sampler settings

All three workflows sample with `steps` 25, `cfg` 1, the `euler` sampler, and the `simple` scheduler.

At `cfg` 1 ComfyUI skips the negative conditioning pass, so a negative prompt has no effect on these workflows. The template keeps `cfg` 1, which is the path Qwen-Image-2.1 is published for; raise it only when you want the negative prompt to take effect. `cfg` 2 follows dense prompts more closely, including small text and numbers, at the cost of over-sharpened edges. Higher values shift exposure toward over-bright or over-dim, `cfg` 5 degrades quality badly, and `cfg` 0.5 breaks the image. Change one value at a time, and compare at a fixed seed.

`steps` is the second lever. The published pipeline for Qwen-Image-2.1 uses about 40 to 50 steps with the `euler` sampler, and the template starts lower at 25. Simple localized edits hold up at 4 to 8 steps, so a request such as changing one garment's color runs several times faster than at the default, while edits that rewrite the whole frame lose coherence and need the full 25. Stubborn fine detail such as hands and fingers settles by about 30 steps, and going from 25 to 40 steps reduces fizzle in detailed areas. Adjust `steps` only when something specific fails to resolve.

### Resolution

**Text to image**: the **Resolution Selector** node sets the aspect ratio and a megapixel target, where 1.0 MP is about 1024x1024. Qwen-Image-2.1 generates natively at 2K, so set the target to about 4.0 MP for a 2048x2048 square output.

**Image edit**: the canvas comes from the `resolution` control on the Image Edit subgraph node (range `0` to `4096`, step `32`). The node's own default is `1024`, and the template ships `0`. With `custom_size` off, which is the template default, the edit is generated at the size of the first reference image, `image_1`:

* `resolution` `0` keeps each reference at its own pixel size, rounded to a multiple of 32, so a 3000x4000 photo gives a 3008x4000 canvas
* `resolution` above `0` scales `image_1`'s aspect ratio to a `resolution` x `resolution` pixel budget, which is a total pixel count rather than a width or height, rounded to multiples of 32, so `resolution` `1024` on the same photo gives about 896x1184

Turn on `custom_size` to take the canvas from the **Resolution Selector** instead, and keep it close to the resized `image_1` size, otherwise the edit can shift.

The canvas size is what drives the cost: a 3000x4000 reference with `resolution` 0 samples at 3008x4000 (about 12 MP), while the same edit at `resolution` 1024 samples at about 896x1184 (about 1 MP). On an RTX 5090 that is roughly 6 s/it against about 0.3 s/it. Lowering `resolution` speeds an edit up by generating a smaller result, not by adding detail. If an edit feels slow, check `resolution` and the pixel size of the reference images before changing anything else. Several large references slow the edit further, and the KV cache settings below affect that cost. Targets well above the model's 2K trained size, such as 4K, lose prompt adherence.

### Prompting for image edit

Reference images are spliced into the text encoder in slot order, and the prompt addresses each one by index:

```
Keep the character and pose in <image1> unchanged, put this light blue denim shirt from <image2> on the character, preserve the original facial features, hair, body shape and pose
```

To point an edit at one region, mark that region on the image first. Open the Mask Editor on the `LoadImage` node that supplies `image_1` (node 470: right-click the node, then **Open in Mask Editor**), paint over the region with the **Paint Pen**, and save. Paint Pen strokes go on the image's RGB layer rather than into the mask, so saving writes the marked image back into the node and the marks become part of the reference the edit reads. Naming the mark's color in the prompt points the instruction at that region, for example "change the jacket in the red area". Keep the mark inside the region you want changed, and state the color you want in the prompt as well when the mark's color appears in the output.

### KV cache

The edit workflows include the **Qwen Image 2.1 Cache** node, which keeps the cached text and reference prefix in memory between sampling steps. The template defaults work for most setups.

The node is experimental, and it exposes two controls:

| Control | Values | Effect |
| - | - | - |
| `device` | `auto` (default), `gpu`, `cpu`, `off` | `auto` uses spare VRAM, then RAM. `cpu` (RAM) is prefetched behind compute and costs little speed. `off` recomputes the prefix every step, which is slower but rules the cache out when debugging |
| `dtype` | `default` (default), `int8`, `int4` | Storage precision. `default` is lossless. `int8` halves the cache at about bf16 accuracy, and `int4` quarters it but roughly doubles the per-step error |

### Prompt enhancement

The text to image and image edit templates include an optional prompt enhancement step. A dedicated text encoder rewrites the prompt before sampling, turning a short request into a longer, more detailed one.

Both templates expose the controls on the Qwen Image 2.1 subgraph node:

* `refine_prompt`: off by default in both templates. Turn it on to have the image model sample with the rewritten prompt instead of the prompt you typed
* `PE_model`: the text encoder that performs the rewrite, `qwen3.5_9b_qwen_image_2.1_pe_t2i.int8_convrot.safetensors` for text to image and `qwen3.5_9b_qwen_image_2.1_pe_i2i.int8_convrot.safetensors` for image edit
* `thinking_mode`: off by default in both templates. Turn it on together with `refine_prompt` to let the rewriting model reason before it writes

A `Preview Any` node inside the subgraph shows the prompt that actually reaches the image model, so you can read the rewrite before deciding whether to keep it. The instruction the rewriting model follows lives in a `Text (System Prompt)` node that you can edit. Because the enhancement branch is only evaluated when `refine_prompt` is on, the enhancement text encoder is optional in both templates; turn `refine_prompt` on to use it. The background removal template does not include this step.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.