AI models on Vofy

Browse image and video models, compare capabilities, and open any model to create with it.

Image models

Compare image models for generation, editing, reference workflows, product shots, and design work.

Video models

Compare video models for text-to-video, image-to-video, motion control, short clips, and cinematic drafts.

Wan 2.6 Flash

Wan 2.6 Flash is a fast AI reference-to-video generator that creates consistent, multi-shot videos from image or video references, with optional audio.

AudioReferencesVideo generation

Wan 2.7

Alibaba's Wan 2.7 AI video generator supports text-to-video and reference-to-video workflows, 720p/1080p output, five aspect ratios, and clips up to 15 seconds.

Text to videoReferencesVideo generation

Gemini Omni Flash

Gemini Omni Flash brings Gemini's native multimodal intelligence to AI video creation and editing. Create 10-second videos from text, photos, references, or source footage, then refine the result with natural-language direction, native audio, and step-by-step creative control.

AudioReferencesVideo generation

Seedance 2.0 Mini

Seedance 2.0 Mini is a streamlined Seedance 2.0 video model for short-form generation, reference-guided motion, video edits, and clip extension. It supports text-to-video, image-to-video, first-and-last-frame generation, multimodal references, and audio-visual sync workflows.

Text to videoImage to videoAudio

Kling 2.6

Kling 2.6 is a balanced Kling video model on Vofy for short clips, motion-controlled video, interpolation, and optional audio workflows. The current Vofy setup supports text-to-video, image-to-video, interpolation, and motion control at 720p or 1080p.

Text to videoImage to videoAudio

Sora 2 Pro

Sora 2 Pro is OpenAI's higher-quality Sora 2 tier on Vofy for longer, more polished AI video. The current Vofy setup supports text-to-video and image-to-video in 16:9 or 9:16, with 720p, 1024p, and 1080p outputs from 4 to 20 seconds.

Text to videoImage to videoLonger clips

How to choose an AI model

Pick by output type, input support, and the quality-speed tradeoff your project needs.

Guide01

Start with the output

Choose image models for still visuals and video models for motion, animation, or clips.

Guide02

Check input support

Look for text prompts, reference images, image editing, image-to-video, or video edits.

Guide03

Balance quality and speed

Use lighter models for drafts and higher-quality models for final creative work.

FAQ

Which AI image model is best for product photos?+

GPT Image and Nano Banana Pro are strong defaults for photorealism, clean composition, and text-heavy layouts. Seedream is useful when you need fast lifestyle or product variations.

Which AI video model should I use for short films?+

Sora Pro is a strong choice for narrative scenes with audio. Kling is useful for multi-shot sequences, while Veo and Seedance are practical for polished image-to-video motion.

Can I switch models in the middle of a project?+

Yes. Vofy keeps your prompts, references, and project context together so you can compare models without rebuilding the brief.

Do all models support reference images?+

Support varies by model. Image models commonly support reference-based edits, and several video models support image-to-video workflows. Each model detail page lists the relevant inputs.

How are credits charged across models?+

Credits depend on model tier, resolution, duration, and generation mode. The pricing page explains the credit cost for common image and video workflows.

How quickly do new models arrive on Vofy?+

Vofy is designed to add new frontier model releases quickly, so teams can try new OpenAI, Google, xAI, ByteDance, and Kuaishou models from the same workspace.

Ready to create with these models?

Create with Vofy using the image or video model that fits your next project.