Wan 2.6 R2V: Complete Specifications, Pricing, API Access & Use Cases (2026)
What Is Wan 2.6 R2V?
Wan 2.6 R2V is Alibaba Cloud’s multimodal reference-to-video generation model, documented in the Wan 2.6 family by December 2025, with text, image, and video inputs, synchronized audiovisual output, a two-to-ten-second generation range, and per-second pricing as of July 2026.
"R2V" means reference-to-video. Unlike a text-to-video model that begins only with a written description, Wan 2.6 R2V accepts reference media that can guide a subject’s appearance, identity, movement, voice characteristics, or role in a newly generated scene.
Alibaba Cloud documents the model for single-role and multi-role video generation, multi-shot narratives, and audio-video synchronization. Output is available at 720p or 1080p, encoded as H.264 video in an MP4 container at 30 frames per second.
The model is relevant to creators and development teams evaluating character-consistent advertising clips, short narrative sequences, virtual presenters, product concepts, social-video assets, and other workflows that need stronger reference control than prompt-only generation provides.
What Are Wan 2.6 R2V’s Key Specifications and Pricing?
The following table separates Alibaba Cloud’s official model specifications from Gate.AI platform information. All values are current as of July 2026 unless stated otherwise.
| Specification | Verified Value |
|---|---|
| Provider | Alibaba Cloud / Tongyi Wanxiang (as of July 2026) |
| Model family | Wan 2.6 (as of July 2026) |
| Model type | Multimodal reference-to-video generation model (as of July 2026) |
| Model status | Wan 2.6 family model documented by December 2025; exact standalone launch date not specified in the reviewed official provider documentation (as of July 2026) |
| Context window | Not applicable as a language-model token context; prompt and media limits are defined through API parameters (as of July 2026) |
| Text-input pricing | No separate token-based input charge is documented (as of July 2026) |
| Reference-input pricing | Alibaba Cloud pricing can count reference-video duration toward billable seconds (as of July 2026) |
| Cached-input pricing | Not applicable or not specified (as of July 2026) |
| Gate.AI720p pricing | \$0.10 per second, as listed on theGate.AImodel card (as of July 2026) |
| Gate.AI1080p pricing | \$0.15 per second, as listed on theGate.AImodel card (as of July 2026) |
| Alibaba Cloud global 720p pricing | \$0.086012 per billable second (as of July 2026) |
| Alibaba Cloud global 1080p pricing | \$0.143353 per billable second (as of July 2026) |
| Pricing unit | Billable video second (as of July 2026) |
| Input modalities | Text, image, and video (as of July 2026) |
| Output modalities | Video with synchronized audio (as of July 2026) |
| Output resolutions | 720p and 1080p (as of July 2026) |
| Output duration | Any integer duration from 2 to 10 seconds in the documented deployment (as of July 2026) |
| Frame rate | 30 fps (as of July 2026) |
| File format | MP4 with H.264 encoding (as of July 2026) |
| Alibaba Cloud model ID | wan2.6-r2v (as of July 2026) |
| Gate.AImodel ID | alibaba/wan2.6-r2v, as listed on theGate.AImodel card (as of July 2026) |
| Provider API access | Alibaba Cloud Model Studio asynchronous video-generation API (as of July 2026) |
| Gate.AIaccess | Listed through theGate.AIModels catalog athttps://gate.ai/models(as of July 2026) |
| Regional availability | Global deployment documented for Germany and the United States; account and workspace configuration may affect access (as of July 2026) |
| Knowledge cutoff | Not applicable to this video generation model (as of July 2026) |
| Fine-tuning support | Not specified in the reviewed official documentation (as of July 2026) |
| Streaming generation | Not documented; the official workflow uses asynchronous tasks (as of July 2026) |
| Batch API support | Not confirmed from official sources (as of July 2026) |
| Tool or function calling | Not applicable (as of July 2026) |
| Structured output | API task metadata is structured, but language-model JSON generation is not applicable (as of July 2026) |
| Usage restrictions | Subject to platform terms, content rules, privacy requirements, and applicable law (as of July 2026) |
Alibaba Cloud’s current global pricing is slightly lower than the rounded Gate.AI model-card rates. These figures belong to separate platforms and should not be treated as interchangeable. Alibaba Cloud lists global rates of \$0.086012 per second at 720p and \$0.143353 per second at 1080p.
Alibaba Cloud calculates reference-to-video charges using billable input and output duration. Its pricing documentation indicates that reference-video duration can be counted, with the chargeable input component capped at five seconds. A request using five billable input seconds and generating ten output seconds may therefore be billed for fifteen seconds.
What Can Wan 2.6 R2V Do That Makes It Useful in Production?
Reference-guided subject consistency
Wan 2.6 R2V can use images and videos to establish the appearance of people, animals, products, or fictional characters. Reference media gives the model a stronger visual target than a text description alone, which can help recurring subjects remain recognizable across newly generated shots.
This capability may fit campaign characters, serialized social content, visual concept development, and authorized digital presenters. It does not guarantee frame-perfect identity preservation, so every output still requires review.
Single-role and multi-role generation
Alibaba Cloud documents support for both single-role and multi-role scenes. Multiple reference assets can be assigned to different entities and invoked in the prompt, enabling short conversations, interactions, reaction shots, or group compositions.
This is useful when a production concept requires more than one recognizable subject. Accuracy can decline as the number of characters, actions, and spatial relationships increases.
Multi-shot narrative generation
The model can create multiple compositions or narrative beats inside a short clip. A prompt may describe establishing shots, close-ups, subject actions, transitions, or camera movement rather than requesting one static composition.
Multi-shot support makes the model relevant to storyboards, previsualization, advertisements, trailers, and social-video drafts. Individual generations remain limited to short durations and normally require editing to form a complete sequence.
Synchronized video and audio
Wan 2.6 R2V generates audiovisual output rather than silent video alone. The documented feature set includes audio-video synchronization, which can support dialogue-led scenes, vocal performances, ambient sound, and short narrative clips.
Synchronization quality can vary with prompt clarity, source-media quality, character count, language, and scene complexity. Production teams should review lip movement, timing, pronunciation, and sound continuity before publishing.
Camera and action guidance
Prompts can guide actions, scene structure, framing, and camera behavior while reference media establishes the subjects. This combination is useful for converting a creative brief into a testable visual sequence.
The model should be treated as a generative production tool rather than a deterministic animation system. Exact trajectories, product geometry, hand placement, logo rendering, and camera paths may differ from the prompt.
What Are Wan 2.6 R2V’s Supported Modalities?
| Modality | Supported? | Notes |
|---|---|---|
| Text input | Yes | Describes the scene, actions, roles, dialogue, style, composition, and camera behavior |
| Image input | Yes | Can define the visual identity or appearance of a referenced subject |
| Video input | Yes | Can provide appearance, behavior, motion, and audiovisual reference information |
| Separate audio input | Not documented for Wan 2.6 R2V | Audio input is documented for newer Wan 2.7 R2V workflows, not as a standard Wan 2.6 R2V input |
| Video output | Yes | 720p or 1080p, 30 fps, MP4/H.264 |
| Audio output | Yes | Synchronized audiovisual output |
| Silent-video output | No for standard Wan 2.6 R2V | Silent output is specifically documented for wan2.6-r2v-flash, not the standard model |
| Text output | No | Responses include task metadata, but the generated media artifact is audiovisual video |
Alibaba Cloud’s current model overview lists text, image, and video as Wan 2.6 R2V inputs. It distinguishes the standard model from wan2.6-r2v-flash, which can generate either audiovisual or silent output.
The reference-to-video API supports multiple reference assets. Exact file-size, duration, resolution, and asset-count requirements should be checked against the deployment-specific API documentation before integration because the limits can change between model versions and regions.
Where Does Wan 2.6 R2V Fall Short?
Generated clips are short: Wan 2.6 R2V generates clips between two and ten seconds in the documented deployment. It is therefore designed for individual shots, short scenes, advertisements, and social-video segments rather than complete long-form productions.
Longer projects require multiple generations, continuity planning, editing, sound mixing, color work, and manual quality control.
Reference consistency is not absolute: Reference media improves control but does not guarantee identical faces, clothing, objects, voices, or proportions in every frame. Fine details can drift, particularly during fast movement, occlusion, camera transitions, or multi-character interaction.
This is a general generative-video limitation and is not unique to Wan 2.6 R2V.
Complex prompts can reduce instruction adherence: Prompts that combine multiple characters, dialogue lines, simultaneous actions, product-placement requirements, camera moves, and visual effects may create competing constraints. The model can omit, merge, or reinterpret details.
Production workflows generally benefit from dividing complex concepts into shorter, focused shots.
Billing includes more than output duration: Alibaba Cloud’s pricing can include the duration of the reference video as well as the generated output. Teams should calculate billable duration rather than estimating cost only from the length of the final clip.
Platform pricing, regional rates, retries, and workflow-level storage or processing costs should also be reviewed separately.
Latency and concurrency vary: Video generation is computationally intensive and uses an asynchronous task workflow. Completion time can vary according to resolution, duration, request complexity, service load, and account limits.
Alibaba Cloud applies rate limits at the account level across associated users, workspaces, and API keys. Requests exceeding an account limit may be rejected temporarily.
Rights and safety review remain necessary: Reference media can contain a person’s face, voice, performance, private information, trademarks, or copyrighted material. Teams must have appropriate rights and consent before uploading or reproducing those assets.
Generated video can also contain inaccurate, misleading, inappropriate, or visually defective material. It should not be used to impersonate people, fabricate evidence, or support legal, medical, financial, or safety-critical decisions without qualified human review.
What Is Wan 2.6 R2V Best Used For?
The model may fit the following scenarios when its short duration, reference controls, and audiovisual output align with the project.
| Use Case | Why Wan 2.6 R2V May Fit | Important Limitation |
|---|---|---|
| Recurring campaign characters | Reference assets can help preserve a recognizable subject across new scenes | Identity, clothing, and accessories can still drift |
| Short advertising concepts | Multi-shot and camera guidance can turn a brief into a visual prototype | Product shape, text, and logos require close review |
| Social-video segments | Two-to-ten-second clips fit intros, transitions, reactions, and short narrative beats | Complete posts usually require editing and captions |
| Authorized virtual presenters | Reference media can guide a presenter’s appearance and audiovisual behavior | Consent, disclosure, and voice rights must be managed |
| Dialogue prototypes | Multi-role support can test short interactions between referenced subjects | Lip synchronization and turn-taking may vary |
| Storyboarding | Rapid generations can help evaluate framing, mood, action, and shot order | Generated footage is not a deterministic storyboard |
| Character-led entertainment | Reference identity and multi-shot output can support short fictional scenes | Continuity across separate jobs requires manual control |
| Creative previsualization | Teams can test camera ideas before committing to production | Physics, timing, and geometry may not be production-accurate |
For prompt-only projects that do not require reference identities, the related Wan 2.6 text-to-video model may provide a more direct workflow.
How Does Wan 2.6 R2V Compare to Sora 2 and Hailuo 2.3?
Wan 2.6 R2V, Sora 2, and Hailuo 2.3 serve overlapping AI-video use cases, but their access methods, reference controls, duration limits, audiovisual features, and pricing structures should be compared at the product level.
| Comparison Area | Wan 2.6 R2V | Sora 2 | Hailuo 2.3 | Scenario Fit |
|---|---|---|---|---|
| Primary workflow | Reference-to-video generation | Broad generative-video creation | Broad generative-video creation | Wan 2.6 R2V may fit workflows centered on referenced subjects |
| Reference inputs | Text, image, and video | Current product-specific controls should be checked before selection | Current product-specific controls should be checked before selection | Verify whether the required reference type is supported |
| Character handling | Single-role and multi-role generation | Depends on current product features | Depends on current product features | Wan is relevant for short scenes involving several referenced entities |
| Audio output | Synchronized audiovisual output | Verify current product and API tier | Verify current product and API tier | Important for dialogue, voice, music, or sound-led clips |
| Output duration | Two to ten seconds | Varies by current product configuration | Varies by current product configuration | Wan is structured around short clips and shots |
| Resolution | 720p and 1080p | Verify current product-specific options | Verify current product-specific options | Wan provides explicitly documented resolution tiers |
| Pricing structure | Per billable second; reference duration can affect Alibaba Cloud cost | Product and access-tier dependent | Product and access-tier dependent | Compare total billable workflow cost rather than headline rates |
| API workflow | Asynchronous provider API;Gate.AImodel listing available | Access method depends on current offering | Access method depends on current offering | Integration requirements may determine suitability |
Teams evaluating Sora 2 or Hailuo 2.3 should compare the current documentation for reference-media support, clip duration, audio behavior, regional access, moderation requirements, concurrency, and effective cost.
No model is universally preferable. Wan 2.6 R2V may be especially relevant when the central requirement is generating short audiovisual scenes around one or more referenced subjects.
How Do I Access Wan 2.6 R2V Through Gate.AI?
Wan 2.6 R2V is available through the Gate.AI Models catalog under the model ID:
alibaba/wan2.6-r2v
As of July 2026, the Gate.AI model card lists:
- 720p output: \$0.10 per second
- 1080p output: \$0.15 per second
- Maximum generated duration: up to 10 seconds
- Output: synchronized audiovisual video
Gate.AI documents an asynchronous video-generation API at:
https://api.gate.ai/api/v1/videos
Requests use a Gate.AI API key in the Authorization: Bearer header. A successful submission returns a job ID and status URL. Applications can poll the status endpoint until the task is completed or failed. The completed response contains a temporary download URL and final billing information.
For reference-to-video generation, the request can include a publicly accessible HTTPS video URL in input_references, using video as the media type and reference_video as its role.
Python Example
The following example submits a Wan 2.6 R2V task, polls its status, and saves the completed video locally.
import osimport timefrom pathlib import Pathfrom typing import Anyimport requestsAPI_BASE_URL = "https://api.gate.ai"MODEL_ID = "alibaba/wan2.6-r2v"API_KEY = os.environ["GATEAI_API_KEY"]REFERENCE_VIDEO_URL = "https://your-cdn.example.com/reference-video.mp4"OUTPUT_PATH = Path("wan-2-6-r2v-output.mp4")HEADERS = {"Authorization": f"Bearer {API_KEY}","Content-Type": "application/json",}def parse_data(response: requests.Response) -> dict[str, Any]:"""Raise for HTTP errors and return the Gate.AI data object."""response.raise_for_status()payload = response.json()if not isinstance(payload, dict):raise RuntimeError("Gate.AI returned an unexpected response.")data = payload.get("data")if not isinstance(data, dict):raise RuntimeError(f"Gate.AI response did not contain a data object: {payload}")return datadef submit_video() -> dict[str, Any]:payload = {"model": MODEL_ID,"prompt": ("Use the referenced subject in a cinematic city scene. ""Preserve the subject's appearance and natural movement. ""The camera slowly tracks from left to right."),"duration": 10,"resolution": "1080p","aspect_ratio": "16:9","generate_audio": True,"seed": -1,"input_references": [{"type": "video","url": REFERENCE_VIDEO_URL,"role": "reference_video",}],"metadata": {"source": "wan-2-6-r2v-example",},}response = requests.post(f"{API_BASE_URL}/api/v1/videos",headers=HEADERS,json=payload,timeout=60,)return parse_data(response)def wait_for_completion(status_url: str) -> dict[str, Any]:while True:response = requests.get(status_url,headers={"Authorization": f"Bearer {API_KEY}"},timeout=30,)data = parse_data(response)status = data.get("status")print(f"Status: {status}; "f"estimated cost: {data.get('estimated_cost', 'not available')}")if status == "completed":return dataif status == "failed":raise RuntimeError(f"Video generation failed: {data}")if status not in {"pending", "in_progress"}:raise RuntimeError(f"Unexpected task status: {status}")time.sleep(5)def download_video(download_url: str) -> None:response = requests.get(download_url,headers={"Authorization": f"Bearer {API_KEY}"},timeout=120,allow_redirects=True,)response.raise_for_status()OUTPUT_PATH.write_bytes(response.content)def main() -> None:submitted = submit_video()job_id = submitted.get("job_id")status_url = submitted.get("status_url")if not job_id or not status_url:raise RuntimeError(f"Submission response is missing job details: {submitted}")print(f"Submitted job: {job_id}")completed = wait_for_completion(status_url)download_url = completed.get("download_url")if not download_url:raise RuntimeError(f"Completed task did not include a download URL: {completed}")download_video(download_url)print(f"Video saved to: {OUTPUT_PATH.resolve()}")print(f"Final billed cost: {completed.get('billed_cost', 'not available')}")if __name__ == "__main__":main()
curl Example
curl --request POST "https://api.gate.ai/api/v1/videos" \--header "Authorization: Bearer $GATEAI_API_KEY" \--header "Content-Type: application/json" \--header "Idempotency-Key: wan-r2v-$(date +%s)" \--data '{"model": "alibaba/wan2.6-r2v","prompt": "Use the referenced subject in a cinematic city scene. Preserve the subject appearance and natural movement while the camera slowly tracks from left to right.","duration": 10,"resolution": "1080p","aspect_ratio": "16:9","generate_audio": true,"seed": -1,"input_references": [{"type": "video","url": "https://your-cdn.example.com/reference-video.mp4","role": "reference_video"}],"metadata": {"source": "wan-2-6-r2v-example"}}'
Generation is asynchronous. A submitted request does not immediately return the video file. Applications must poll the returned status URL or configure a webhook callback.
Completed video URLs can expire. Production applications should copy the generated file into controlled, persistent storage shortly after the task completes.
Before deployment, confirm that alibaba/wan2.6-r2v remains enabled for the documented /api/v1/videos endpoint and that the model card still supports the selected duration, resolution, aspect ratio, audio option, and reference-media role.
FAQs
What resolution and duration does Wan 2.6 R2V support?
Wan 2.6 R2V supports 720p and 1080p audiovisual video at 30 frames per second. Alibaba Cloud’s documented deployment allows integer clip durations from two to ten seconds.
How much does Wan 2.6 R2V cost?
The Gate.AI model card lists \$0.10 per second at 720p and \$0.15 per second at 1080p. Alibaba Cloud lists separate global rates of \$0.086012 and \$0.143353 per billable second, respectively, as of July 2026.
Can developers access Wan 2.6 R2V through an API?
Yes. Gate.AI lists the model as alibaba/wan2.6-r2v. Alibaba Cloud Model Studio separately documents an asynchronous API using the provider model ID wan2.6-r2v. Developers should follow the current documentation for their selected platform.
What is Wan 2.6 R2V suitable for?
Wan 2.6 R2V may fit short character-consistent advertisements, authorized virtual-presenter clips, dialogue prototypes, multi-role scenes, social-video segments, storyboards, and creative previsualization. All outputs require review for visual defects, factual accuracy, consent, rights, and brand compliance.


