Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

Generate, Edit and Make Transparent Images with Qwen-Image-2.1

Qwen-Image-2.1 combines text-to-image, image editing and native RGBA output in one diffusers pipeline. Here is how to run it.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Qwen-Image-2.1 is the Qwen team’s open text-to-image and image editing model. One pipeline does both. Its visual generation component has 7B parameters across 32 Single-Stream DiT layers. Native transparency is one of the four improvements the card lists for this release. The model can generate RGBA images from a prompt and edit transparent layers. The card also says it can extract subjects from photographs, but it gives no code or prompt for extraction, so this post doesn’t cover it. If you make stickers, product cutouts or UI assets, the model may let you skip the usual “generate, then matte” pipeline. The card doesn’t publish any evaluation of alpha or matting quality, though. Check the edges on your own outputs before you rely on them in production.

Key specs

Visual generation component 7B parameters, 32 Single-Stream DiT layers
Tasks Text-to-image, image editing, transparent (RGBA) generation and editing, subject extraction (no usage documented)
Reference images for editing Up to 10
Local edit controls Circles, painted annotations, or separate masks
Supported sizes 7 supported sizes, ~4 MP each, 1:1 to 16:9/9:16
Efficiency features Mixed-granularity attention, prefix KV cache reuse
Library diffusers (QwenImage21Pipeline)
License Qwen Research License

The card also lists improved typography, portrait lighting and fine detail as improvements in this release. It doesn’t publish any benchmark numbers, so judge quality on your own prompts or in the demo Space.

Install

The card installs diffusers from source. The version specifiers are quoted here, because your shell would otherwise treat > as an output redirect.

Terminal window
pip install "torch>=2.4.0"
pip install "transformers>=5.17"
pip install git+https://github.com/huggingface/diffusers
pip install accelerate pillow

Run it

t2i.py
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
image = pipe(
prompt="A neon shop sign that reads \"QWEN IMAGE 2.1\", rainy night, reflections on wet pavement",
width=2048, height=2048,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("t2i_example.png")

If the model doesn’t fit on your GPU, drop .to("cuda") and turn on CPU offload. It runs slower but uses less peak VRAM:

t2i_offload.py
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
image = pipe(
prompt="A neon shop sign that reads \"QWEN IMAGE 2.1\", rainy night, reflections on wet pavement",
width=2048, height=2048,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("t2i_offload.png")

Batch-generate transparent sticker assets

For transparent images, the card gives one example prompt and calls it the recommended prompt format:

This is an RGBA image with transparency. A cute cartoon dragon sticker. The image has alpha channel and the background is transparent.

The card doesn’t mark a slot for the subject or say which parts must stay fixed. Our reading is that the first and last sentences carry the transparency instruction and the middle sentence describes the subject. The script below follows that reading: it swaps in a different subject each time and keeps the other two sentences unchanged. It makes a set of sticker assets from a list of subjects and gives each one a fixed seed so you can reproduce it:

stickers.py
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
subjects = [
"A cute cartoon dragon sticker",
"A cartoon coffee cup with a smiling face, sticker style",
"A flat illustration of a paper airplane, sticker style",
]
for i, subject in enumerate(subjects):
prompt = (
"This is an RGBA image with transparency. "
f"{subject}. "
"The image has alpha channel and the background is transparent."
)
image = pipe(
prompt=prompt,
width=2048, height=2048,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42 + i),
).images[0]
print(f"{i}: mode={image.mode}, size={image.size}")
image.save(f"sticker_{i:02d}.png")

Save to PNG, as the card does, because JPEG has no alpha channel. The card doesn’t document the output mode, so the script prints image.mode. Check that it says RGBA before you pass the files further down your pipeline.

Bulk background replacement for product photos

To edit, pass a PIL image as image and describe the change in prompt. The card says the model preserves identity for people and products. That fits catalog work, where you want the same product in a new scene. This script processes every PNG and JPEG in a folder. Files with other extensions are skipped:

rebackground.py
from pathlib import Path
import torch
from PIL import Image
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
src_dir = Path("products")
out_dir = Path("products_edited")
out_dir.mkdir(exist_ok=True)
prompt = "Change the background to a sunset beach"
extensions = {".png", ".jpg", ".jpeg"}
def load_flat_rgb(path):
# Composite any transparency onto white before dropping alpha,
# so transparent pixels don't turn black or into junk.
im = Image.open(path).convert("RGBA")
background = Image.new("RGBA", im.size, "white")
return Image.alpha_composite(background, im).convert("RGB")
for path in sorted(src_dir.iterdir()):
if path.suffix.lower() not in extensions:
continue
input_image = load_flat_rgb(path)
image = pipe(
prompt=prompt,
image=input_image,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
if path.suffix.lower() in {".jpg", ".jpeg"}:
image = image.convert("RGB")
image.save(out_dir / path.name)
print(f"edited {path.name}")

Product PNGs often open as RGBA, P or LA images, and the card doesn’t say which input modes the editing path accepts. It does say the model can edit transparent layers, so the model might treat an RGBA input as a transparent-layer edit rather than a regular photo edit. The script flattens every input to RGB to avoid this. A plain .convert("RGB") would discard the alpha channel without compositing, and the pixels that were transparent could reach the model as black or garbage. So the script first composites each image onto a white background. The card’s own editing example passes Image.open(...) straight through with no conversion. Skip the flattening only if you want to edit transparent layers. The script saves results to a separate products_edited folder under the original filenames, so it never overwrites your source files. JPEG outputs are converted to RGB before saving because JPEG can’t store alpha.

The card also says the model supports up to 10 reference images and local edits marked with circles, painted annotations or separate masks. The quick-start code doesn’t show the call signature for any of these, so check the GitHub repo before you build on them.

Text-heavy graphics at every aspect ratio

Typography is one of the listed improvements, and the card includes text rendering examples. A common job is making one banner in several formats for different placements. Use the card’s table of supported resolutions rather than choosing your own sizes:

banners.py
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
aspect_ratios = {
"1:1": (2048, 2048),
"4:3": (2400, 1792),
"3:4": (1792, 2400),
"3:2": (2528, 1696),
"2:3": (1696, 2528),
"16:9": (2752, 1536),
"9:16": (1536, 2752),
}
prompt = (
"A minimalist conference poster with large bold text that reads "
"\"BUILD WITH OPEN MODELS\", subtitle \"October 2026\", "
"clean geometric shapes, soft gradient background"
)
for ratio in ["1:1", "16:9", "9:16"]:
width, height = aspect_ratios[ratio]
image = pipe(
prompt=prompt,
width=width, height=height,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save(f"banner_{ratio.replace(':', 'x')}.png")
print(f"saved {ratio} at {width}x{height}")

We suggest putting the exact text you want rendered in quotes inside the prompt, as the card’s neon-sign example does. This is our suggestion, not a rule from the card. The backslashes are only Python string escaping. Every run reuses the same seed, but a different canvas shape gives a different composition, so the layouts won’t match across ratios.

Gotchas

  • Diffusers from source. The card installs diffusers from GitHub, plus transformers>=5.17 and torch>=2.4.0. Quote the version specifiers in your shell. Diffusers main changes often, so pin a commit if you need reproducible builds.
  • Use bfloat16. Every example on the card loads with torch_dtype=torch.bfloat16. Keep that setting unless you have a reason to change it.
  • Large recommended canvas. All the supported resolutions are about 4 megapixels (2048×2048 and similar), so each image is expensive. The card gives no VRAM requirement. On smaller cards, start with enable_model_cpu_offload().
  • Supported resolutions only. The card lists a fixed set of supported sizes. Other sizes may work, but the card doesn’t say they do.
  • RGBA depends on the prompt. The only RGBA method the card documents is its recommended prompt format. It shows no other option. Without that wording you may get an opaque background. The card shows only one example, so treating its first and last sentences as a reusable wrapper is our reading.
  • Subject extraction isn’t documented. The card says the model can extract subjects from photographs but gives no code or prompt for it.
  • Flatten RGBA inputs carefully when editing. The card doesn’t say how the editing path handles RGBA or palette inputs. If you convert to RGB, composite onto a solid background first, or transparent areas may turn black or into junk.
  • Editing output size isn’t documented. The card’s editing example passes no width/height. Compare the output size with your input before you use the results.
  • No published benchmarks. The card makes quality claims but gives no scores, including none for alpha quality. Test it on your own prompts.
  • License. The model uses the Qwen Research License, not Apache or MIT. Read the LICENSE before any commercial or production use.

When to pick it

Pick Qwen-Image-2.1 if you need transparent assets straight from a prompt, or if you want generation and instruction-based editing in one model and one diffusers pipeline. The 10-reference-image limit and identity preservation make it worth testing for product and portrait compositing. It also runs at roughly 2K resolutions out of the box.

Skip it, or wait, if:

  • you need a permissive license for a commercial product and the Qwen Research License doesn’t cover your use;
  • you need a stable, released diffusers version in production rather than a source install;
  • you need published benchmark numbers to justify the choice, because the card has none;
  • you need a multi-reference, mask-based editing or subject-extraction API today and can’t confirm how to call it from the GitHub repo.

It’s worth an afternoon of testing for prototyping sticker packs, transparent assets, text-heavy graphics and batch background swaps on a single CUDA GPU.

Related