For the complete documentation index, see llms.txt. This page is also available as Markdown.

Qwen-Image-2.1: How to Run Locally

Run Qwen-Image-2.1 locally with Unsloth FP8 and GGUF quants.

Qwen-Image-2.1 is a new 7B parameter text-to-image generation and image editing model by Qwen which runs locally on 11GB VRAM with GGUFs and 24GB with INT8/FP8. This guide shows you how to run Qwen-Image-2.1 for image gen along with memory requirements, recommended settings and more.

Qwen-Image-2.1 combines a 7B image generator with a Qwen3-VL 8B encoder. The model supports native 2K generation, text rendering, transparent images and editing with up to 10 reference images.

The quants use Unsloth Dynamic methodology which upcasts important layers to higher precision for more accuracy. Run them in Unsloth Desktop, diffusers, stable-diffusion.cpp and more. Unsloth uploads: GGUF • FP8

Image editing by FP8 Qwen-Image-1.2 via Unsloth

🖥️ Memory requirements

Use the table below to choose a starting configuration for your hardware. Generate one image at the listed resolution, then increase it if memory allows.

FP8 quants are recommended for GPUs with 24 GB of VRAM or more. It may also be preferable for any setup with a GPU, as it can offer faster inference than GGUF models even with offloading. Just ensure you have enough RAM for offloading.

Use GGUF models mostly if you have CPU RAM, or use a unified-memory system such as a Mac.

Memory figures are estimates, not tested minimums. Memory usage varies with resolution and offloading settings.

Available hardware
Starting configuration

12–16 GB VRAM

GGUF Q4_K_M, 1024×1024, batch 1

24 GB VRAM

INT8/FP8: 512 × 512 GGUF Q4_K_M: 1024 × 1024

6GB VRAM

You can run FP8 on just 6GB of VRAM using offloading. Inference will be <2x slower.

CPU-only: 12–16 GB RAM

GGUF Q4_K_M with the Q4_K_XL text encoder

Apple Silicon: 12–16 GB+ unified memory

GGUF with a compatible native backend

CPU offloading can help when VRAM is limited, but requires additional system RAM and may slow generation.

Quantization Analysis

We show that INT8 has the lowest reasonable LPIPS (lower is better) than even FP8, so we default to using INT8.

scheme
file size
LPIPS mean
LPIPS max
SSIM mean

int8

7.26 GB

0.064

0.167

0.936

fp8

7.12 GB

0.112

0.329

0.899

⚡ Quickstart

1

Setup Unsloth

The easiest way to get started is by downloading the Unsloth Desktop app. Works on macOS, Windows, and Linux.

Download Unsloth

Or, if you prefer to install manually:

MacOS, Linux, WSL:

Windows PowerShell:

2

Setup Qwen-Image-2.1

In this release, Qwen-Image-2.1 is available in Unsloth for text-to-image generation only.

Select Images from the sidebar. The Create an image workflow opens by default.

Open Select image model in there or go to Model hub and search for your desired Qwen-Image-2.1 GGUF or FP8 quant from the selector which will download the model.

3

Generate your image

Describe the image you want to generate. You can keep the recommended settings for your first image.

Click Generate. The first generation may take longer while the model downloads and loads. Your result will appear in the gallery, where you can view its settings or download it.

Image Editing

You can easily also edit photos with transparent backgrounds using Qwen-Image-2.1. Remove elements, add elements, change the vibe while keeping the same foundation/structure. Just select the Edit tab inside of Unsloth and write your prompt.

Use the recipe for the backend actually running the model, not just its FP8/GGUF file extension.

Setting
Unsloth GPU / Diffusers
Native stable-diffusion.cpp

Steps

40 - Qwen/Diffusers baseline

20 - CLI default and Unsloth's recorded test

Guidance / CFG

1.0 - CFG disabled

6.0 - upstream sd.cpp example

Negative prompt

Leave blank; set guidance to 1.0 explicitly

Leave blank to reproduce the example

Sampler / scheduler

Keep the model's pipeline scheduler

Euler

Flow shift

Keep the pipeline's scheduler configuration

Automatic; omit --flow-shift

First-run resolution

1024×1024

1024×1024

Batch size

1

1

Seed

42, or another fixed comparison seed

42, or another fixed comparison seed

Resolution and aspect ratios

Use dimensions divisible by 32. Qwen's native presets include:

Shape
Smaller test size
Qwen native preset

Square

1024×1024

2048×2048

Landscape 4:3

1152×864

2400×1792

Portrait 3:4

864×1152

1792×2400

Photo 3:2

1248×832

2528×1696

Portrait 2:3

832×1248

1696×2528

Widescreen 16:9

1536×864

2752×1536

Vertical 9:16

864×1536

1536×2752

Unsloth supports dimensions up to 2048 pixels per side. Start at 1024 × 1024 and increase the resolution if memory allows. Lowering the step count speeds up generation but may not resolve memory limits.

LoRAs and ControlNet

Load compatible LoRAs to add a character, subject or visual style to your generations. Adjust the LoRA weight to control how strongly it affects the result.

Supported models can also use ControlNet to guide the structure of an image. Adjust its strength to control how closely the result follows the guidance image.

Generated images are saved to the local gallery. Open a result to view its prompt and settings, select Recipe to restore its generation setup or download the image.

Supported models and workflows

What you want to do
Models to start with

Generate or transform images

Z-Image, Qwen-Image, FLUX.1, SDXL

Edit using instructions

Qwen-Image-Edit, FLUX.1 Kontext

Use reference images

FLUX.2 klein

More models are available in the model picker. Formats and workflows vary by model, and unsupported workflows are hidden automatically.

Generation settings

The default settings are a good starting point. However, it is possible to adjust settings to get the best result, these are the main parameters you may want to change:

Setting
What it changes

Prompt

Describes the image you want to create. In Edit, this becomes the instruction to follow.

Aspect ratio and resolution

Set the shape and size of the image. Larger images use more memory and take longer.

Steps

Control how long the model spends generating. More steps do not always improve the result.

Guidance

Controls how strongly the prompt guides the result. Some models work best at 0.

Seed

Helps recreate a result. Leave it empty to use a random seed.

Advanced generation settings

These controls change how the model runs on your device. Adjust them to reduce memory usage, improve performance or troubleshoot generation.

Setting
What it changes

Speed

Controls compilation and performance optimizations. Compiled modes may take longer during the first generation.

Precision

Changes the numerical format used to run the model. Lower precision can reduce memory usage. Available options depend on the model and hardware.

Attention

Selects the attention implementation used during generation.

Memory

Balances generation speed against GPU memory usage. Try Low VRAM if a model is close to your memory limit.

Step cache

Reuses some calculations between diffusion steps to improve generation speed.

CPU offload

Moves parts of the model into system memory. This reduces GPU memory usage but may make generation slower.

⚠️ Troubleshooting

Images does not appear

Update to the latest version, then restart the app. See Updating Unsloth for instructions.

A model will not download or load

Some image models are large and take time to download. Check that you have enough free storage and a stable internet connection.

Some models also require you to accept their licence on Hugging Face and add a Hugging Face token before downloading them.

The model runs out of memory

Try a smaller GGUF quantization or a 4-bit version of the model. For GGUF models, start with the size marked recommended.

You can also:

  • Lower the image resolution.

  • Keep the batch size at 1.

  • Close other applications using the GPU.

  • Choose a smaller model.

A TIGHT model may use system memory and run more slowly. A model marked OOM is unlikely to fit.

A workflow is missing

The available workflows depend on the loaded model. Unsupported workflows will not appear.

For example, instruction-based editing requires a model such as Qwen-Image-Edit or FLUX.1 Kontext.

❓ FAQ

Does image generation run locally?

Yes. Once a model is downloaded, image generation runs on your device.

What is the difference between Transform and Edit?

Transform redraws an existing image using a prompt that describes the complete result.

Edit follows instructions such as changing a background or adding an object. It requires a compatible editing model.

Can I train my own image LoRA?

Yes. Fine-tune a supported model on your own images, then load the finished LoRA in Create.

See Fine-tune an image model to learn more.

Last updated

Was this helpful?