# Wan 2.6 Online - Text, Image, and Video Reference Controls

> Canonical: https://stable-diffusion-web.com/wan

## Wan 2.6 - text, image, and video reference generator

The current form supports text, one-image, or up-to-three-video reference modes, model-specific duration controls, 720p or 1080p, shot type and seed where exposed, and paid access.

## Wan 2.6 video-reference examples

The gallery illustrates character, sequence, dialogue, camera, vertical, and product directions; it is not evidence of multi-shot planning, audio sync, or identity consistency.

### Reference Character

### Sequence Concept

### Dialogue Scene

### Cinematic Motion

### Vertical Short

### Product Scene

## Why Wan 2.6

The live form owns the current input, duration, ratio, resolution, shot-type, seed, and reference-video controls. It does not expose automatic multi-shot or generated-audio controls.

### Three input modes

Text mode requires a prompt, image mode requires one image plus a prompt, and reference mode requires video input plus a prompt.

### Text and image durations

Text and image modes offer 5, 10, or 15 seconds. The selected mode and downstream constraints determine which combination can run.

### Reference-video durations

Reference mode offers 5 or 10 seconds and preprocesses source-video duration before submission.

### 720p and 1080p settings

All three modes expose 720p and 1080p. Inspect delivered dimensions, compression, and artifacts after every run.

### Mode-specific ratio, shot, and seed controls

Text mode offers 16:9, 9:16, or 1:1 plus shot type and seed. Image mode exposes shot type and seed; reference mode does not expose those fields.

### Explicit upload limits

Image mode accepts one JPEG, PNG or WebP up to 10 MB. Reference mode accepts up to three MP4, MOV, or MKV files up to 100 MB each.

## Who uses Wan 2.6 video references, and for what?

Video creators, marketers, storyboard artists, and product teams can use these candidate workflows to evaluate text, image, and reference-video controls; every clip still needs review.

### Prompt-led motion tests

Use text mode for one focused scene, then inspect subject motion, anatomy, object contact, framing, text, and artifacts.

### Single-image animation studies

Use one clear source image and compare the clip for identity, product, background, crop, and composition drift.

### Reference-video experiments

Supply accepted video files when source motion should condition a candidate, then verify what was preserved or changed.

### Vertical, square, and landscape tests

Text mode can compare 9:16, 1:1, and 16:9 while holding the prompt and seed constant.

### Duration and resolution comparisons

Compare 5, 10, or 15 seconds where available and 720p versus 1080p, recording debit, delivered files, and artifacts.

### Product and campaign concepts

Treat generated clips as drafts for post-production and reject changes to labels, proportions, colors, trademarks, or regulated claims.

## What a run costs and returns

| Measure | Detail |
| --- | --- |
| 1 model | Available in the form today, each priced in the same credits. |
| Every paid plan | Includes a commercial licence for what you make. |
| JPEG, PNG, and WebP | Accepted for upload, up to 10 MB per file. |
| 5,000 characters of prompt | The longest prompt every model on this page accepts. |

## Wan 2.6 reference-to-video FAQ

Answers about three input modes, upload limits, duration, ratio, resolution, shot type, seed, paid access, credits, time, and review boundaries.

### What is Wan 2.6 on this page?

Wan 2.6 is a selectable paid-access video workflow with text, image, and reference-video modes. The live form defines which controls each mode exposes.

### Which input modes are available?

Choose a required prompt in text mode, one required image plus prompt in image mode, or accepted video references plus prompt in reference mode.

### What are the image and video upload limits?

Image mode accepts one JPEG, PNG or WebP up to 10 MB. Reference mode accepts up to three MP4, MOV, or MKV files up to 100 MB each.

### Which duration choices are available?

Text and image modes offer 5, 10, or 15 seconds. Reference-video mode offers 5 or 10 seconds.

### Which ratios and resolutions can I choose?

Text mode offers 16:9, 9:16, or 1:1. All modes expose 720p and 1080p; image and reference modes do not expose the same ratio selector.

### Where are shot type and seed available?

Text and image modes expose shot type and seed. Reference-video mode does not expose those controls in the current schema.

### How much does a run cost, and what access is required?

The catalogue marks the model as paid access and metadata starts at 80 credits with an estimated 120 seconds. Confirm the live estimate before submitting.

### Does this form expose generated audio or automatic multi-shot planning?

No. The current schema exposes neither a generated-audio control nor an automatic multi-shot storyboard control. Assemble accepted clips and audio in post-production.

### What should I review before using a clip?

Review prompt adherence, identity, anatomy, objects, text, framing, duration, resolution, reference drift, and artifacts. Generated clips are candidates rather than verified final footage.

### What limits should I check before using Wan 2.6?

Generated output from Wan 2.6 can contain visual, audio, or factual errors. Review each result before publishing, and upload only material you are authorized to use. Where the form accepts a reference image, supported files are JPEG, PNG, WebP up to 10 MB each.

### What happens to the files I upload to Wan 2.6?

An upload is stored against your own account, not pooled with other people's. Deleting the account removes it: the control is in Settings under Security, the request is confirmed by email, and the erasure removes the stored objects it owns. How long anything else is kept, and which processors handle it, is stated in the privacy policy.

### Can I use what Wan 2.6 makes commercially?

Every paid plan includes a commercial licence for what Wan 2.6 makes; the free tier is for personal use.

### How does Wan 2.6 compare, and what alternatives fit a different workflow?

Wan 2.6: Use Wan 2.6 with text, one image, or up to three video references; 5-15 second choices by mode, 720p/1080p, shot type, seed, and paid access. Related options are Stable Diffusion Models Hub: Compare the Stable Diffusion models and the other image and video models here: the inputs each takes, what a run costs, and the plan it needs, Z-Image Turbo: Use Z-Image Turbo with a required text prompt, ten ratios, 1-49 inference steps, guidance 0-20, seed, output count, and paid-model access, Veo 3.1: Use Veo 3.1 with text or first/end-frame image input, 16:9 or 9:16, 720p, 1080p or 4K, a seed control, fixed 8-second output, and generated audio, Seedream 5 Lite: Use Seedream 5 Lite with text or up to ten JPEG, PNG or WebP references, 2K or 3K output, eight ratios, seed, output count, and paid-model access, Seedream 4.5: Use Seedream 4.5 with text or up to ten JPEG, PNG or WebP references, 2K or 4K output, eight ratios, seed, output count, and paid-model access, and Seedance 2 Mini: Seedance 2 Mini is the lowest-cost Seedance 2 variant - the same multimodal ByteDance model for 1080p video with native audio, text-to-video. Choose by the required input, delivered output, controls, limits, current availability, and credit cost; the related-tools section links each canonical page.

## Test a current Wan 2.6 workflow

Open the generator, confirm paid-model access, choose a supported mode and settings, review the estimate, and inspect every returned clip.

## Related tools

### [Stable Diffusion Models Hub](https://stable-diffusion-web.com/models)

### [Z-Image Turbo](https://stable-diffusion-web.com/z-image)

### [Veo 3.1](https://stable-diffusion-web.com/veo-3)

### [Seedream 5 Lite](https://stable-diffusion-web.com/seedream-5-lite)

### [Seedream 4.5](https://stable-diffusion-web.com/seedream-4-5)

### [Seedance 2 Mini](https://stable-diffusion-web.com/seedance-2-mini)

_Last updated July 14, 2026_
