Using Krea with reference images
Qwen3-VL encoder. This allows us to pass an image along with the prompt. In simple terms, we can just skip the whole description of an image we want and pass one instead.
Currently, I know of three ways of doing this, each one for a different use case:
This is a brief post about the three modes we can use with Krea-2, using Modular Diffusers, custom blocks, LoRAs, and SDNQ. If you just want to play with them, you can use this Space, or if you just want the code, you can go to the Diffusers recipes and get it there.
Vision only
This method doesn't require any LoRA. We just need to encode the image and pass the embeddings to the model along with the prompt.
Using this image as a reference:
Original photo by Microsoft Copilot on Unsplash
With just this code:
import torch
from sdnq import SDNQConfig # noqa: F401
from diffusers.modular_pipelines import ModularPipelineBlocks
from diffusers.utils import load_image
blocks = ModularPipelineBlocks.from_pretrained("OzzyGT/krea2_reference_blocks", trust_remote_code=True)
pipe = blocks.init_pipeline("OzzyGT/Krea_2_Turbo_sdnq_dynamic_8bit")
pipe.load_components(dtype=torch.bfloat16)
pipe.to("cuda")
image = pipe(
prompt="photo",
reference_images=load_image(
"https://huggingface.co/datasets/OzzyGT/diffusers-examples/resolve/main/krea2_reference/20260907154202_edit_image_0.png"
),
output="images",
)[0]
which basically uses the the prompt "photo" and the image, we get this result:
As you can see, it does a pretty good job, but it also misses some details that the model probably can't see. So if you want a more precise replica of the image, you can add the missing details to the prompt, like ethnicity or what each element on the bed is. Using this method, you can also do some very light editing of the image.
For example, with this prompt: "change the color of the sweater to light blue and replace the popcorn with french fries and scones with a hamburger" you will get this image:
As you can see, it does change the image somehow, but every time we do it, the model produces a different result, and it's also not as precise. This is equivalent to prompting something like this: "a photo of an orange car in the city, change the color of the car to red", so it does kind of make sense that it works, with the first part of the prompt being the embeddings from the image.
One cool feature that I haven't seen anyone using is that, if you use a mask, you can prevent the model from seeing the parts that you don't want or need. So, for example, if I mask everything but the man, we get this image:
Or if I mask just the laptop:
The script used for these images is this one: reference_vision_only.py
Style reference
One thing that doesn't work with just vision is transferring the style from the image to the generation. For this, Ostris trained a LoRA that does this pretty well.
This one is pretty straightforward to use. You just prompt like normal and it will use the style of the image.
Using this image:
Photo by HYEWON HWANG on Unsplash
with this code:
import torch
from sdnq import SDNQConfig # noqa: F401
from diffusers.modular_pipelines import ModularPipelineBlocks
from diffusers.utils import load_image
blocks = ModularPipelineBlocks.from_pretrained("OzzyGT/krea2_reference_blocks", trust_remote_code=True)
pipe = blocks.init_pipeline("OzzyGT/Krea_2_Turbo_sdnq_dynamic_8bit")
pipe.load_components(dtype=torch.bfloat16)
pipe.load_lora_weights("ostris/krea2_turbo_style_reference", weight_name="krea2_style_reference.safetensors")
pipe.to("cuda")
image = pipe(
prompt="a capybara",
reference_images=load_image(
"https://huggingface.co/datasets/OzzyGT/diffusers-examples/resolve/main/krea2_reference/20260908045229_edit_image_0.png"
),
reference_mode="append",
output="images",
)[0]
Will produce this image:
It kind of works, but this is a bad image to test it with since it has three subjects too close to each other, so it mixes them into one. This is where it gets interesting to use a mask, since you can isolate the specific style and subject you want from any image. If we mask only the red plushie in the middle, we get this image:
Pretty cool, and it allows for some really creative experimentation, since you don't need to mask a specific subject or a clearly defined part of an image.
The script used for these images is this one: reference_style.py
Identity edit
This one is just like any other edit model. You pass an image and ask the model to edit something in it, and it should preserve most of the image.
Using this image:
Photo by Alvin David on Unsplash
and this code:
import torch
from sdnq import SDNQConfig # noqa: F401
from diffusers.modular_pipelines import ModularPipelineBlocks
from diffusers.utils import load_image
blocks = ModularPipelineBlocks.from_pretrained("OzzyGT/krea2_reference_blocks", trust_remote_code=True)
pipe = blocks.init_pipeline("OzzyGT/Krea_2_Turbo_sdnq_dynamic_8bit")
pipe.load_components(dtype=torch.bfloat16)
pipe.load_lora_weights("conradlocke/krea2-identity-edit", weight_name="krea2_identity_edit_v1_2.safetensors")
pipe.to("cuda")
image = pipe(
prompt="add sunglasses and a hat to the capybara",
reference_images=load_image(
"https://huggingface.co/datasets/OzzyGT/diffusers-examples/resolve/main/krea2_reference/alvin-david--cHj_FC6ukw-unsplash_croped.jpg"
),
reference_mode="prepend",
mask_reference_latents=True,
output="images",
)[0]
we get this image:
I don't see that much use for masks with this one, but you still can if you want, for example, to crop the capybara out of the image.
The script used for these images is this one: reference_identity_edit.py
Hope you liked it! These three techniques can help improve your workflow and add a little extra to your generations.
If you have any questions or suggestions, please don't hesitate to reach out to me on X OzzyGT or in the Diffusers channel on our Discord.










