Wan 2.5 T2V Preview: Complete Specifications, Pricing, API Access & Use Cases (2026)
What Is Wan 2.5 T2V Preview?
Wan 2.5 T2V Preview is Alibaba’s text-to-video generation model, commercially launched as part of Wan2.5 Preview on November 11, 2025, featuring short video generation with synchronized audio, 480P/720P/1080P output options, and resolution-based per-second pricing as of July 2026.
Alibaba Cloud lists wan2.5-t2v-preview in the Wan video generation family. The official model documentation describes it as a "video with audio" text-to-video model with audio-video synchronization, text and audio input, 5-second and 10-second output durations, 30 fps video, and MP4 output using H.264 encoding.
The model is designed for short prompt-driven video generation rather than long-form production. Users typically evaluate it for creative concept clips, social media drafts, campaign storyboards, education visuals, and other workflows where text prompts need to produce brief video assets with matching audio. Teams comparing AI video generation model specs should review duration, output quality, audio support, pricing unit, prompt controls, and API availability before choosing a production model.
What Are Wan 2.5 T2V Preview’s Key Specifications and Pricing?
| Field | Verified Value |
|---|---|
| Provider | Alibaba / Alibaba Cloud Model Studio, as of July 2026 |
| Model Family | Wan / Wan2.5 Preview video generation family, as of July 2026 |
| Model Type | Text-to-video generation model with audio output, as of July 2026 |
| Release Date | Commercial Wan2.5 Preview launch event: November 11, 2025; exact first preview availability date is not separately confirmed from official sources as of July 2026 |
| Context Window | Not applicable as a video generation model; prompt length is the relevant input limit, as of July 2026 |
| Prompt Length | Up to 1,500 characters for Wan2.6 and Wan2.5 series text-to-video prompts, as of July 2026 |
| Input Pricing | Not token-priced; Alibaba Cloud states Wanx text-to-video billing charges output only, as of July 2026 |
| Cached Input Pricing | Not applicable / not confirmed from official sources as of July 2026 |
| Output Pricing | International deployment: 480P at \$0.05/second, 720P at \$0.10/second, 1080P at \$0.15/second, as of July 2026 |
| China Mainland Pricing | 480P at \$0.043006/second, 720P at \$0.086012/second, 1080P at \$0.143353/second, as of July 2026 |
| Pricing Unit | Per second of generated output video, based on output resolution, as of July 2026 |
| Supported Input Types | Text prompt and optional audio URL, as of July 2026 |
| Supported Output Types | MP4 video with audio, H.264 encoding, 30 fps, as of July 2026 |
| Resolution Options | 480P, 720P, and 1080P, as of July 2026 |
| Duration Options | 5 seconds or 10 seconds, as of July 2026 |
| Provider API Access | Alibaba Cloud Model Studio asynchronous API, as of July 2026 |
| Alibaba Model ID | wan2.5-t2v-preview, as of July 2026 |
| Gate.AIModel ID | alibaba/wan2.5-t2v-preview, as per theGate.AIlisting for this page, as of July 2026 |
| Gate.AIListing Pricing | 480P at \$0.05/second, 720P at \$0.10/second, 1080P at \$0.15/second, as per theGate.AIlisting for this page, as of July 2026 |
| Availability | Alibaba Cloud documents International and Chinese mainland deployment scopes for this model, as of July 2026 |
| Rate Limits | Not confirmed from official sources as of July 2026 |
| Fine-tuning Support | Model-specific fine-tuning support for wan2.5-t2v-preview is not confirmed from official sources as of July 2026 |
| Streaming Support | Not confirmed from official sources as of July 2026 |
| Batch API Support | Not confirmed from official sources as of July 2026 |
| Tool / Function Calling | Not applicable to this video generation model as documented, as of July 2026 |
| Structured Output / JSON Mode | Not applicable to generated video output; not confirmed as a model feature as of July 2026 |
| License / Usage Restrictions | Model-specific restrictions are not fully summarized in the available model documentation as of July 2026; users should review Alibaba Cloud andGate.AIpolicy terms before production deployment |
Alibaba Cloud states that the size and duration parameters affect cost, with the cost calculated as resolution-based unit price multiplied by duration in seconds. Gate.AI’s pricing page also states that image, audio, video, and similar multimodal capabilities may be billed by generation count, duration, resolution, or task specification, which is consistent with the per-second pricing model shown in the Gate.AI listing.
What Can Wan 2.5 T2V Preview Do That Makes It Useful in Production?
Wan 2.5 T2V Preview can generate short videos from text prompts. This is useful for teams that need visual motion drafts before committing budget to filming, animation, editing, or a longer creative pipeline. Because Alibaba documents 5-second and 10-second duration options, the model is best evaluated as a short-clip generator rather than a full scene or long-form video system.
The model supports video with audio. Alibaba Cloud’s model table lists "video with audio" and "audio-video synchronization" for wan2.5-t2v-preview, which makes it relevant for social clips, mood videos, audio-backed concepts, and short marketing drafts where sound and motion need to align.
Wan 2.5 T2V Preview also supports optional custom audio input through input.audio_url for Wan2.5 series models. Alibaba documents public HTTP/HTTPS audio URLs, WAV and MP3 formats, a 3-second to 30-second audio range, and a 15 MB file-size limit. If the audio is longer than the requested video duration, only the matching duration is used; if the audio is shorter, the remaining video is silent.
Resolution flexibility is another production advantage. The model supports 480P, 720P, and 1080P tiers, allowing teams to run lower-cost drafts at 480P and reserve 720P or 1080P for higher-quality preview assets. Alibaba’s API reference also lists common aspect-ratio sizes, including horizontal, vertical, and square formats across resolution tiers.
The model can fit prompt-controlled creative workflows, but it should not replace editorial review. Generated video may still need trimming, brand review, policy review, quality control, rights review, and post-production refinement before public release.
What Are Wan 2.5 T2V Preview’s Supported Modalities?
| Modality | Supported? | Notes |
|---|---|---|
| Text Input | Yes | Primary prompt input for text-to-video generation; Wan2.5 series prompts support up to 1,500 characters |
| Audio Input | Yes | Optional audio_url is supported for Wan2.5 series models; public HTTP/HTTPS URLs, WAV/MP3, 3–30 seconds, up to 15 MB |
| Image Input | No for this T2V model | Image input is documented for Wan image-to-video variants, not for wan2.5-t2v-preview |
| Video Input | No for this T2V model | Video input is documented for other Wan tasks, not for this text-to-video model |
| Video Output | Yes | MP4 video, H.264 encoding, 30 fps, 5s or 10s duration |
| Audio Output | Yes | The model is documented as "video with audio" with audio-video synchronization |
Where Does Wan 2.5 T2V Preview Fall Short?
Wan 2.5 T2V Preview is limited to short clips. Alibaba’s API reference lists valid duration values of 5 and 10 seconds for wan2.5-t2v-preview, so workflows requiring long scenes, multi-minute ads, or full narrative sequences need editing, stitching, or a different video model.
The model also has a defined prompt-length limit. Alibaba documents up to 1,500 characters for Wan2.6 and Wan2.5 series text-to-video prompts, with longer text automatically truncated. This means complex visual direction should be concise and structured.
Cost can rise quickly during iteration. Alibaba documents pricing by output resolution and duration, so repeated 1080P generations cost more than 480P draft passes. Teams using multimodal generation APIs should track resolution choice, retry volume, rejected generations, and review cycles rather than only the cost of the final accepted asset.
Gate.AI access details should also be handled carefully. The Gate.AI listing for this page identifies the model ID and pricing, while Gate.AI’s public documentation confirms OpenAI-compatible setup for supported models. However, the exact Gate.AI video-generation endpoint path, request schema, and response schema for this specific model are not visible in the public documentation reviewed for this rewrite.
This is a general AI limitation and is not model-specific unless stated by Alibaba: generated videos can contain artifacts, distorted motion, inaccurate visual details, unsafe content, misleading depictions, or poor text rendering. Human review is required for advertising, political, legal, medical, financial, identity-sensitive, copyrighted, or brand-regulated use.
What Is Wan 2.5 T2V Preview Best Used For?
| Use Case | Why Wan 2.5 T2V Preview May Fit | Important Limitation |
|---|---|---|
| Short creative concept clips | Converts text prompts into 5-second or 10-second video with audio | Not designed for long-form video generation |
| Social media draft videos | Supports 480P, 720P, and 1080P tiers for cost-quality control | Final formatting and platform compliance still need review |
| Audio-backed visual concepts | Supports synchronized audio output and optional custom audio URL input | Professional audio mixing may still be required |
| Advertising storyboard exploration | Helps teams test scene direction, motion, tone, and sound before production | Brand, likeness, copyright, and policy review remain necessary |
| Education and training visuals | Can create short explanatory clips from concise prompts | Factual accuracy and visual clarity require human verification |
| Product and campaign ideation | Useful for rapid exploration of mood, motion, and scene composition | Generated output should not be treated as final production material without QA |
Wan 2.5 T2V Preview is most suitable when the target output is a short AI-generated clip with audio. For longer narratives, complex multi-shot continuity, image-conditioned generation, or video-reference workflows, teams may need to evaluate other AI content generation models with verified duration, modality, and pricing requirements.
How Does Wan 2.5 T2V Preview Compare to Seedance 2.0 and Veo 3.1?
| Comparison Area | Wan 2.5 T2V Preview | Seedance 2.0 | Veo 3.1 | Scenario Fit |
|---|---|---|---|---|
| Provider | Alibaba | ByteDance Seed Team | Choose based on ecosystem, API access, compliance, output needs, and pricing | |
| Model Type | Text-to-video with audio | Audio-video generation and editing model | Video generation model with native audio | Wan 2.5 fits short text-to-video use; Seedance and Veo may fit broader video workflows |
| Inputs | Text and optional audio URL | Official launch material states text, image, audio, and video inputs | Google documents video generation workflows with native audio and image-based direction | Use modality requirements to narrow the choice |
| Output Duration | 5s or 10s | ByteDance describes up to 15-second high-quality multi-shot audio-video output | Duration depends on current Google API configuration and should be verified at implementation time | Wan 2.5 is best framed as a short-clip model |
| Resolution | 480P, 720P, 1080P | Verify current resolution options in official documentation before deployment | Verify current resolution options in Google documentation before deployment | Resolution affects cost and review quality |
| Pricing Clarity | Alibaba documents per-second pricing by resolution;Gate.AIlisting mirrors 480P/720P/1080P per-second pricing for this page | Not fully covered in the cited launch material | Check current Google billing documentation | Wan 2.5 has clear public per-second pricing in the cited Alibaba documentation |
| Main Constraint | Short duration and prompt-length limit | Official launch material notes continued model limitations | Google API limits and model availability should be checked before implementation | No universal winner; model fit depends on task constraints |
Wan 2.5 T2V Preview is a strong candidate for short text-to-video clips where synchronized audio, clear pricing, and 480P–1080P options are important. Seedance 2.0 and Veo 3.1 may be more relevant when a workflow needs broader multimodal references, longer clip options, or a different provider ecosystem, but production selection should be based on verified current documentation rather than headline capability claims.
How Do I Access Wan 2.5 T2V Preview Through Gate.AI?
As per the Gate.AI listing for this page, Wan 2.5 T2V Preview is available with the model ID alibaba/wan2.5-t2v-preview. The listing describes it as an Alibaba text-to-video model that can generate 5–10 second videos from text prompts with a native synchronized audio track. It supports 480P, 720P, and 1080P output tiers, priced at \$0.05/second for 480P, \$0.10/second for 720P, and \$0.15/second for 1080P, as of July 2026.
Gate.AI’s public documentation states that its OpenAI-compatible base URL is https://api.gate.ai/openai/v1, and that supported model IDs generally use the provider/model format. Gate.AI’s pricing page also states that video and similar multimodal capabilities may be billed by generation count, duration, resolution, or task specification.
The examples below use Gate.AI’s OpenAI-compatible setup, the model ID shown in the Gate.AI listing, and the listed resolution and duration options. For production use, developers should confirm the final accepted video-generation parameter names and response object in the Gate.AI console or model-specific documentation.
Python Example
from openai import OpenAIimport osclient = OpenAI(api_key=os.environ["GATE_AI_API_KEY"],base_url="https://api.gate.ai/openai/v1")response = client.chat.completions.create(model="alibaba/wan2.5-t2v-preview",messages=[{"role": "user","content": ("Generate a 10-second 720P cinematic video of a golden retriever ""running across a sunny beach at sunset, with gentle waves and ""a native synchronized audio track.")}],extra_body={"resolution": "720p","duration": 10})print(response)
This Python example configures the Gate.AI client with the OpenAI-compatible base URL and calls the listed model ID alibaba/wan2.5-t2v-preview. The resolution and duration fields reflect the Gate.AI listing’s supported output options. Before deploying this request, verify whether the model-specific Gate.AI schema uses these exact field names or an alternative video-generation payload format.
curl Example
curl https://api.gate.ai/openai/v1/chat/completions \-H "Authorization: Bearer $GATE_AI_API_KEY" \-H "Content-Type: application/json" \-d '{"model": "alibaba/wan2.5-t2v-preview","messages": [{"role": "user","content": "Generate a 10-second 720P cinematic video of a golden retriever running across a sunny beach at sunset, with gentle waves and a native synchronized audio track."}],"resolution": "720p","duration": 10}'
This curl example follows Gate.AI’s OpenAI-compatible API pattern and uses an environment variable for the API key. The request includes the listed model ID, a text prompt, a 720P resolution setting, and a 10-second duration setting. Developers should verify the final response format and any model-specific video parameters in Gate.AI before using the request in production.
FAQs
What is Wan 2.5 T2V Preview’s maximum video duration?
Alibaba Cloud documents 5-second and 10-second duration options for wan2.5-t2v-preview as of July 2026.
How much does Wan 2.5 T2V Preview cost?
Alibaba’s international pricing is \$0.05/second at 480P, \$0.10/second at 720P, and \$0.15/second at 1080P as of July 2026.
Does Wan 2.5 T2V Preview support API access?
Yes. Alibaba Cloud Model Studio documents asynchronous API access. As per the Gate.AI listing, the Gate.AI model ID is alibaba/wan2.5-t2v-preview.
What is Wan 2.5 T2V Preview useful for?
It is useful for short text-to-video clips with synchronized audio, including concept visuals, social drafts, ad storyboards, training clips, and motion experiments. Human review is still required before production use.


