Vofy
Vofy
BlogModelsAppsCampaignImageVideoPricing
BlogModelsAppsCampaignImageVideoPricing
Vofy

Your ALL-IN-ONE AI Creative Studio

Status unavailableJoin Discord
© 2026 Vofy. All rights reserved.
Product
  • Image
  • Video
  • Models
  • Rankings
  • Apps
  • Pricing
Company
  • Blog
  • Contact
Legal
  • Privacy
  • Terms

Wan 2.7 Image and Video Models Are Now Live on Vofy

Wan 2.7 image and video models are now live on Vofy. Explore reference workflows, supported formats, use cases, and how to start creating.

Try Vofy free →
Wan 2.7 Image and Video Models Are Now Live on Vofy - Featured visual guide
Vofy Team
Vofy Team•Editorial Team•Jul 9, 2026

Disclosure: This is a Vofy product announcement. Availability, controls, and workflow details reflect the Vofy product experience as of July 21, 2026 and may change as the models and platform evolve.

Wan 2.7 Video and Wan 2.7 Image are now available on Vofy. The two releases bring Alibaba's current Wan generation to both sides of a visual workflow: still-image creation and editing for the frames you need, plus short-form video generation for motion, narrative, and campaign delivery.

This is not one interchangeable model with two output buttons. Wan 2.7 Image is built around text-to-image, reference-guided editing, and coherent image sets, while Wan 2.7 Video supports text-to-video and reference-to-video workflows on Vofy. The practical benefit is a clearer path from a visual idea to a set of related images and then to motion, without forcing every stage into the same prompt or control scheme. Availability and supported settings below reflect Vofy as of July 2026.

TL;DR

  • Wan 2.7 Video on Vofy supports text-to-video and reference-to-video with 720p or 1080p output, five aspect ratios, and clips up to 15 seconds.
  • Wan 2.7 Image supports text-to-image and image-to-image creation, up to nine reference images, and 1K or 2K output.
  • Image set mode can request up to 12 related images, while standard mode creates up to four independent results.
  • Both models are live in Vofy's image and video creation workspaces, with the applicable credit cost shown before generation.

The shared version name matters less than choosing the right workflow for the job. Use the image model when composition, references, and a related set of stills are the main deliverable; use the video model when timing, motion, framing, and short-form playback are the main constraints.

1. What Is Launching on Vofy

Wan 2.7 is Alibaba's current visual generation family spanning image and video work. Alibaba announced Wan2.7-Image in April 2026 as a unified generation and editing model, then introduced Wan2.7-Video as a broader suite for generation and editing workflows. The provider-level releases cover more modes than any single product integration necessarily exposes, so it is important to separate the Wan family specification from the controls currently available on Vofy.

The table below summarizes the Vofy launch surface. These are the options documented on the live Vofy model pages as of July 2026, rather than a claim that every feature in Alibaba's broader model suite is present in each Vofy workspace.

CapabilityWan 2.7 Video on VofyWan 2.7 Image on Vofy
Core modesText-to-video, reference-to-videoText-to-image, image-to-image
Reference inputsFirst frame; up to five total images or videos; optional paired voice clipsUp to nine input images
Output size720p or 1080p1K or 2K
Quantity or duration2-15 seconds; up to 10 seconds when a reference video is included1-4 standard outputs; up to 12 requested in image set mode
Format control16:9, 9:16, 1:1, 4:3, 3:4Square, portrait, landscape, and widescreen presets

That distinction keeps production decisions straightforward. A creator planning a product launch can develop a related set of product stills with Wan 2.7 Image, then use selected frames or other approved assets as references for a Wan 2.7 video concept. The two workspaces remain independent, which lets each prompt focus on the controls that matter at that stage.

For background on the upstream releases, see Alibaba Cloud's Wan2.7-Image announcement and Wan2.7-Video announcement. The official Wan website remains the provider's destination for the wider model family.

2. Wan 2.7 Video: More Ways to Direct a Short Clip

The Vofy video integration is designed for two starting points. Text-to-video is appropriate when the scene can be described from scratch: subject, action, setting, camera movement, lighting, visual style, and pacing. Reference-to-video is the more controlled route when a subject, opening composition, movement cue, or established visual direction needs to survive the jump into motion.

On Vofy, a reference workflow can use a first frame and up to five total reference images or videos. Optional voice clips can be paired with referenced subjects when voice continuity is part of the brief. Output can be set to 720p or 1080p in 16:9, 9:16, 1:1, 4:3, or 3:4. Text-led and image-reference workflows support durations from 2 to 15 seconds, while a workflow containing a reference video supports up to 10 seconds. These limits make Wan 2.7 a practical fit for social ads, product reveals, music visuals, pitch-film moments, and short previsualization rather than an entire finished long-form sequence in one generation.

AI-generated Wan 2.7 cinematic video frame showing an ice caravan

AI-generated Wan 2.7 video example from Vofy showing a wide cinematic scene with multiple subjects.

Prompting should still be selective. A useful first pass names one primary action, one camera instruction, a setting, and a pacing cue instead of packing several scene changes into a short duration. For example: "A ceramic perfume bottle on black volcanic rock as a thin wave washes around it, slow dolly-in, overcast coastal light, restrained luxury commercial, deliberate pacing." If identity or product geometry matters, add references and describe what each one should control rather than expecting the model to infer their roles.

3. Wan 2.7 Image: References, Edits, and Related Sets

Wan 2.7 Image covers both new image generation and reference-guided editing in the Vofy image workspace. A prompt can begin without an input image, or it can combine as many as nine references to establish a subject, product, material, palette, composition, or brand direction. That makes the model useful when a brief is defined by several visual constraints that would be awkward to compress into text alone.

The output controls support 1K and 2K resolution across square, portrait, landscape, and widescreen presets. Standard mode returns between one and four independent images. When sequential image generation is enabled, a request can ask for up to 12 related images for campaign beats, product views, story development, or architectural exploration; the model determines the actual number returned. The distinction is important: image set mode aims for a coherent group, but it is not a guarantee that every requested frame will be returned or that every fine detail will remain identical across the set.

AI-generated Wan 2.7 Image series showing a consistent subject across greenhouse scenes

AI-generated Wan 2.7 Image example from Vofy demonstrating a related multi-image sequence.

Alibaba also reports broader provider-level strengths for Wan2.7-Image, including long text handling, multilingual rendering, personalization, and precise color direction. Those claims come from the upstream announcement, while Vofy's current launch page specifically documents the generation, editing, reference, resolution, and image set controls described above. For production work, test text-heavy layouts and exact brand-color requirements against your own source assets before committing to a full batch.

4. Where the Combined Workflow Helps

The strongest reason to have both Wan 2.7 models in one creative platform is not that every asset must use the same model family. It is that creators can separate look development from motion development while keeping source material close at hand. A marketer can explore a product's lighting and setting as stills, select the strongest composition, and then use an approved frame to guide a vertical teaser. A filmmaker can build visual beats as a related image set before translating one moment into a camera move. A social creator can create portrait directions first, then choose whether a square post, a vertical clip, or both are worth producing.

AI-generated Wan 2.7 product video frame featuring a perfume bottle and fluid motion

AI-generated Wan 2.7 video example from Vofy illustrating a vertical product-marketing concept.

There are also cases where another model is a better starting point. If a project already depends on a proven Kling workflow, the Kling 3.0 guide provides a separate route for multi-shot video work. Creators focused on prompt structure and character continuity can compare the approach with the Seedance 2 prompt playbook. For still images where text rendering is the central constraint, the GPT Image 2 guide is a useful alternative evaluation point. Model choice should follow the deliverable and source material, not a version number alone.

5. How to Start Creating

Both models are available from their dedicated Vofy workspaces. The quickest path is to choose the output type first, prepare the references you are authorized to use, and then set the delivery format before generating.

  1. Open Wan 2.7 Image in Vofy for still generation, edits, or image sets, or open Wan 2.7 Video in Vofy for motion.
  2. Write the prompt around the deliverable. For images, define the subject, composition, palette, material, and edit direction. For video, define the subject, action, camera, setting, and pacing.
  3. Add references only when they have a clear role, then choose the supported resolution, format, quantity, or duration.
  4. Review the credit estimate shown in the workspace and generate. Credit rates vary by model, resolution, duration, and selected settings as of July 2026.

Start with a small result count or short duration while testing prompt and reference behavior. Once the direction is stable, increase the delivery setting that matters for the final asset. This reduces wasted iterations and makes it easier to identify whether a mismatch came from the prompt, a reference, or an output setting.

6. Limits to Keep in Mind

Reference inputs improve control, but they do not turn generation into deterministic compositing. Small product details, typography, hands, repeated characters, or exact spatial relationships can still drift and should be inspected at full size. Coherent image sets likewise aim for related outputs rather than frame-perfect identity across every image. For video, a 15-second maximum is enough for a social beat or concept shot, but longer narratives still need multiple generations and editorial assembly.

The Vofy video integration currently presents text-to-video and reference-to-video, even though Alibaba describes a wider Wan2.7-Video suite that also includes image-to-video and video editing models. Do not assume every upstream mode is available through the same Vofy control. Check the live model pages for the current settings before planning a delivery, especially when a project depends on a specific input type, exact duration, or voice-reference behavior.

7. Conclusion

Wan 2.7 gives Vofy creators two distinct production tools under one model generation: a reference-aware image workspace for stills and related sets, and a short-form video workspace for text-led or reference-led motion. The useful question is not whether image or video is better. It is which uncertainty should be solved first: the look of the frame or the movement inside it. Choose that starting point, test with a focused prompt, and carry only the strongest assets into the next stage.

FAQ

Is Wan 2.7 available on Vofy now?

Yes. Wan 2.7 Video and Wan 2.7 Image are available on Vofy as of July 2026 through their dedicated model pages and creation workspaces.

What is the difference between Wan 2.7 and Wan 2.7 Image?

Wan 2.7 on Vofy is the video model, with text-to-video and reference-to-video workflows. Wan 2.7 Image is the image generation and editing model, with text-to-image, image-to-image, multi-reference, and coherent image set workflows.

How long can a Wan 2.7 video be on Vofy?

Text-led and image-reference workflows support clips from 2 to 15 seconds. Workflows that include a reference video support up to 10 seconds as of July 2026.

How many references can I use?

Wan 2.7 Video supports a first frame and up to five total reference images or videos, with optional paired voice clips. Wan 2.7 Image supports up to nine input images for reference-guided generation or editing.

Can Wan 2.7 Image create a consistent series?

Its sequential image generation mode is intended for coherent sets and can request up to 12 images, although the actual returned count and fine-detail consistency can vary. Treat the first run as visual development and review the set before using it in a campaign or storyboard.

Are Wan 2.7 credit costs fixed?

No. Vofy shows the applicable estimate in the creation workspace before generation. Rates vary by model, resolution, duration, quantity, and selected settings, so the live interface is the current source of truth.

Related Reading

  • Kling 3.0: A guide to its video workflows
  • Seedance 2 prompt guide and character workflow
  • GPT Image 2: What it is and when to use it

These guides provide comparison points when a Wan workflow does not match the source material or delivery format. Use them to evaluate the next model by the same criteria: inputs, controls, output constraints, and the amount of iteration the project can absorb.

More Articles

  • Kling 3.0 Complete Guide: Features, Pricing, Prompts, and Best Use Cases
    Kling 3.0 Complete Guide: Features, Pricing, Prompts, and Best Use CasesMar 17, 2026
  • Seedance 2.0 Prompt Guide for Better AI Video
    Seedance 2.0 Prompt Guide for Better AI VideoApr 3, 2026
  • GPT Image 2 Guide: What It Is, Key Features, and Use Cases
    GPT Image 2 Guide: What It Is, Key Features, and Use CasesApr 27, 2026
  • How to Use an AI Muscle Generator for Realistic Photos
    How to Use an AI Muscle Generator for Realistic PhotosJul 20, 2026
  • Wan 2.7
    Wan 2.7Jul 15, 2026
  • Wan 2.7 Image
    Wan 2.7 ImageJul 15, 2026

Try it yourself on Vofy

Generate AI images and videos with the best models - all in one studio.

Start for free →

Discover More

Kling 3.0 Complete Guide: Features, Pricing, Prompts, and Best Use Cases
Ryan MitchellRyan Mitchell•Mar 17, 2026

Kling 3.0 Complete Guide: Features, Pricing, Prompts, and Best Use Cases

Comprehensive guide to Kling 3.0 covering why it leads AI video generation right now, its core features, pricing tradeoffs, basic prompt structure, and best use cases.

Jul 15, 2026new

Wan 2.7

Alibaba's Wan 2.7 AI video generator supports text-to-video and reference-to-video workflows, 720p/1080p output, five aspect ratios, and clips up to 15 seconds.

Jul 15, 2026new

Wan 2.7 Image

Wan 2.7 Image is Alibaba's AI image generator and editor with up to nine references, 1K or 2K output, and coherent image sets of up to 12 images.