Vofy
Vofy
BlogModelsAppsCampaignImageVideoPricing
BlogModelsAppsCampaignImageVideoPricing
Vofy

Your ALL-IN-ONE AI Creative Studio

Status unavailableJoin Discord
© 2026 Vofy. All rights reserved.
Product
  • Image
  • Video
  • Models
  • Rankings
  • Apps
  • Pricing
Company
  • Blog
  • Contact
Legal
  • Privacy
  • Terms

Multi-Reference Image Generation Test: GPT Image 2 vs Wan 2.7 vs Seedream 5 Pro

Compare GPT Image 2, Wan 2.7 Image, and Seedream 5.0 Pro with a controlled multi-reference image generation test and reusable scoring method.

Try Vofy free →
Multi-Reference Image Generation Test: GPT Image 2 vs Wan 2.7 vs Seedream 5 Pro - Featured visual guide
Marcus Chen
Marcus Chen•Senior AI Researcher•Aug 2, 2026

Disclosure: Vofy is an all-in-one AI creative studio. This comparison uses image models available through Vofy, but the protocol is designed to expose reference drift rather than favor a platform or vendor. All nine retained first outputs and their measured scores are published below; the ranking applies only to this test set.

Multi-reference image generation is easy to demonstrate and difficult to evaluate. Uploading a portrait, a product photo, a moodboard, and a composition reference can produce an attractive image, yet the result may quietly change the face, redesign the product, ignore the layout, or average all four references into one vague style. A useful test therefore has to ask more than whether the output looks polished.

This comparison puts GPT Image 2, Wan 2.7 Image, and Seedream 5.0 Pro through the same three assignments. Each assignment gives every reference one explicit job and measures whether the protected details survive together. The protocol, prompts, ratios, scoring rules, retained outputs, and final scores are published below.

TL;DR

  • Within each assignment, all three models receive the same four-reference pack, prompt wording, 2K setting, target ratio, and first-output rule.
  • As configured in Vofy on August 7, 2026, GPT Image 2 accepts up to 16 input images, Wan 2.7 Image up to 9, and Seedream 5.0 Pro up to 10. Input capacity is not evidence of better reference fidelity.
  • Three assignments isolate different risks: identity plus wardrobe preservation, product geometry plus material, lighting, and layout control, and multi-reference campaign consistency.
  • Every output is scored on identity or object fidelity, reference-role separation, prompt compliance, composition, and artifact integrity. In this retained set, GPT Image 2 ranks first at 9.7, followed by Wan 2.7 Image at 7.3 and Seedream 5.0 Pro at 7.0.
  • The practical choice may be workflow-specific: GPT Image 2 exposes the largest input allowance and inpainting, Wan 2.7 Image includes a sequential-image option, and Seedream 5.0 Pro provides a focused one-output generation path.

1. The Decision You Are Actually Making

The real decision is not which model can accept several files. It is which model can keep the right information from each file without blending protected details. A creator may want Reference 1 to control a person's face, Reference 2 to control a jacket, Reference 3 to control the location, and Reference 4 to control the camera angle. If the result borrows the lighting from the portrait, the face from the pose reference, and a redesigned version of the jacket, the model technically used the references but failed the production brief.

Good multi-reference work depends on role separation. Identity references should not become style references. A product reference should control geometry and label placement, not the whole background. A moodboard should guide palette, texture, and light without contributing an extra subject. A composition reference should control framing and spacing without replacing the people or objects in the scene. The prompt must make these boundaries explicit, and the score must penalize cross-reference leakage.

This matters most when one wrong detail creates expensive repair. A portrait with the wrong person cannot be fixed by color grading. A bottle with a new cap shape is not a faithful product image. A campaign frame that loses its required negative space may force a complete rerun. In this retained set, GPT Image 2 protects the most important constraints with the least repair work, although its Test 1 portrait still shows visible identity drift.

Use only source images you own or have permission to edit. The identity test should use a consenting adult, and its outputs should not be presented as documentary evidence or used to impersonate another person.

Test 1 scoring reference pack with the identity portrait, cobalt technical jacket, blue-hour greenhouse, and seated three-quarter pose

Test 1 reference pack used for scoring: Reference 1 provides identity, Reference 2 the jacket, Reference 3 the greenhouse and lighting, and Reference 4 the seated pose.

2. Comparison at a Glance

The table describes controls exposed by the current Vofy configuration, not benchmark performance. These limits were checked on August 7, 2026. Provider updates or integration changes can alter them later, so a rerun should record its own date and visible settings.

Vofy capabilityGPT Image 2Wan 2.7 ImageSeedream 5.0 Pro
Image-to-image modeYesYesYes
Maximum input images16910
Shared test resolution2K2K2K
Shared test ratios1:1, 3:4, 16:91:1, 3:4, 16:91:1, 3:4, 16:9
Maximum outputs per requestUp to 10Up to 121
Distinct workflow controlQuality control and inpaintingSequential image generationPNG or JPEG output choice
What the test must verifyWhether more references remain distinctWhether related outputs preserve protected detailsWhether a single polished output preserves all assigned roles

Maximum input count measures capacity, not judgment. A model can accept ten references and still ignore the one that matters most. Likewise, a model with several output slots does not automatically provide a more consistent first image. To keep this comparison fair, every scored request asks for one result, uses four references, and sets the common 2K resolution rather than exploiting a model-specific maximum.

For vendor context, consult the official OpenAI image generation guide, ByteDance's Seedream 5.0 Pro launch post, and Alibaba Cloud's Wan2.7-Image announcement. Those pages describe intended capabilities. Only retained outputs from the locked protocol below should determine this article's scores.

3. How the Multi-Reference Test Works

Every model receives the same reference files in the same order. The prompt identifies each file by number and assigns it one role. We use the common 2K setting, request one output, disable any optional sequence behavior, and keep the first completed image. There are no hidden negative prompts, model-specific prompt rewrites, rerolls, or local edits before scoring. A technical error is recorded separately and does not become an aesthetic failure.

The three assignments cover different production pressures. Test 1 asks the model to preserve a consenting adult's identity while borrowing a separate garment, environment, and pose. Test 2 combines exact fictional product geometry with a surface, lighting direction, and advertising composition. Test 3 asks for a square campaign image that keeps one character, one prop family, one palette, and one layout system distinct. Together, they reveal whether a model follows reference roles or simply blends the visual average.

Each image receives zero, one, or two points in five categories. Protected fidelity measures the face, product, or hero object. Role separation checks whether each reference controls only its assigned property. Prompt compliance covers required and forbidden details. Composition evaluates framing, hierarchy, and usable negative space. Artifact integrity covers hands, faces, geometry, repeated elements, and unwanted text. One visible defect may lower more than one category only when it independently violates both requirements, and the scoring rationale must identify both effects. The maximum is ten points, but the article should also publish the decisive failures because a single identity or product error can outweigh a high average.

ScoreMeaning
0The requirement is missing, replaced, or unusable without generative repair
1The requirement is recognizable but visibly drifts or needs localized repair
2The requirement is preserved well enough for normal crop, color, and export work

4. Test 1: Identity, Wardrobe, Environment, and Pose

This is the highest-risk assignment because four references can all contain human cues. Reference 1 is the only identity source. Reference 2 shows the garment without a person. Reference 3 supplies the greenhouse location and light. Reference 4 supplies pose and framing through a silhouette or mannequin so it cannot compete with the face. That source design makes a failure easier to diagnose: if identity changes, the model cannot blame a second portrait.

Target aspect ratio: 3:4 Resolution: 2K Output rule: one first output per model

Use Reference 1 only for the adult subject's facial identity, skin tone, hairstyle, and age. Use Reference 2 only for the exact cobalt technical jacket, including its high collar, diagonal chest pocket, matte fabric, and silver zipper. Use Reference 3 only for the glass greenhouse setting, blue-hour light, wet floor reflections, and cool green palette. Use Reference 4 only for the seated three-quarter pose and camera framing. Create a photorealistic vertical editorial portrait of the subject seated on a simple metal bench inside the greenhouse, looking slightly past the camera, hands relaxed and visible, jacket fully readable, restrained magazine color grade, natural skin texture. Do not borrow identity from any reference except Reference 1. Do not add jewelry, text, logos, extra people, or a hat.

Pass checks: recognizable identity from Reference 1; jacket construction from Reference 2; greenhouse and lighting from Reference 3; seated three-quarter pose from Reference 4; plausible hands; no added person, text, logo, jewelry, or hat.

GPT Image 2 output for the identity, wardrobe, greenhouse, and seated-pose multi-reference test
GPT Image 2, Test 1 first output.
Wan 2.7 Image output for the identity, wardrobe, greenhouse, and seated-pose multi-reference test
Wan 2.7 Image, Test 1 first output.
Seedream 5.0 Pro output for the identity, wardrobe, greenhouse, and seated-pose multi-reference test
Seedream 5.0 Pro, Test 1 first output.

5. Test 2: Product Geometry, Material, Lighting, and Layout

Product work replaces identity risk with geometry risk. Reference 1 shows a fictional ceramic fragrance bottle named ARCA, including its cap, shoulder curve, label size, and pale gray glaze. Reference 2 supplies only the dark basalt plinth material. Reference 3 supplies a hard sunset light direction and long shadow. Reference 4 supplies a wide advertising composition with the product on the right and empty copy space on the left. The label is intentionally short so the test can separate typography drift from larger product redesign.

Target aspect ratio: 16:9 Resolution: 2K Output rule: one first output per model

Use Reference 1 only for the exact ARCA fragrance bottle geometry, pale gray ceramic glaze, black cylindrical cap, and front label placement. Use Reference 2 only for the rough black basalt plinth material. Use Reference 3 only for the low warm light coming from camera left and the long crisp shadow direction. Use Reference 4 only for the wide composition: generous empty space on the left, bottle group on the right, low horizon. Create a premium photorealistic fragrance campaign still with one bottle standing upright on the basalt plinth. Preserve the bottle proportions and cap design. The front label should read only "ARCA". No flowers, no hands, no extra bottles, no added copy, no logo, no watermark.

Pass checks: one upright bottle; preserved silhouette, cap, glaze, and label placement; basalt material; left-side warm light; negative space on the left; no added props or copy.

GPT Image 2 output for the ARCA bottle, basalt plinth, warm light, and wide-layout multi-reference test
GPT Image 2, Test 2 first output.
Wan 2.7 Image output for the ARCA bottle, basalt plinth, warm light, and wide-layout multi-reference test
Wan 2.7 Image, Test 2 first output.
Seedream 5.0 Pro output for the ARCA bottle, basalt plinth, warm light, and wide-layout multi-reference test
Seedream 5.0 Pro, Test 2 first output.

6. Test 3: Character, Props, Palette, and Social Layout

The third assignment tests whether several soft visual signals can remain distinct without turning into an incoherent collage. Reference 1 controls one illustrated courier character. Reference 2 controls three parcel props and their shapes. Reference 3 supplies a coral, cobalt, cream, and charcoal palette. Reference 4 supplies a square editorial layout with a strong central diagonal and a clean upper-right corner for later deterministic text. The output must feel like one campaign image while preserving the specific visual job of each source.

Target aspect ratio: 1:1 Resolution: 2K Output rule: one first output per model

Use Reference 1 only for the illustrated bicycle courier's face, short dark hair, yellow raincoat, and graphic ink-and-gouache rendering. Use Reference 2 only for the exact three parcel shapes: one tall blue tube, one flat coral box, and one small cream cube. Use Reference 3 only for the coral, cobalt, cream, and charcoal color system. Use Reference 4 only for the square composition, central diagonal movement, and empty upper-right area. Create a polished editorial campaign illustration of the courier riding uphill in light rain while carrying all three parcels securely. Keep the upper-right corner visually quiet for later copy. No readable text, no logo, no additional parcels, no extra person, no watermark.

Pass checks: character continuity; exactly three specified parcels; palette without source leakage; central diagonal movement; clear upper-right copy space; no extra text or people.

GPT Image 2 output for the courier, three parcels, palette, and square-layout multi-reference test
GPT Image 2, Test 3 first output.
Wan 2.7 Image output for the courier, three parcels, palette, and square-layout multi-reference test
Wan 2.7 Image, Test 3 first output.
Seedream 5.0 Pro output for the courier, three parcels, palette, and square-layout multi-reference test
Seedream 5.0 Pro, Test 3 first output.

7. Reading the Results Without Cherry-Picking

The first score to inspect is protected fidelity. If the person, product, or hero character drifts, a polished composition cannot rescue the assignment. The second is role separation. Look for lighting copied from the wrong reference, colors leaking from a portrait into the garment, background objects borrowed from a moodboard, or a composition reference replacing the intended subject. These failures often look plausible until the output is compared directly with every source.

Publish the full first-output set, not one favorite from each model. Insert every native output directly so readers can inspect faces, hands, bottle geometry, label placement, and parcel counts. Keep each scored output unchanged before publication. Repairability can be discussed later, but altered pixels cannot serve as benchmark evidence.

Test 1 is scored against the published four-reference pack. GPT Image 2 keeps the subject recognizable but shows visible facial and beard drift, so protected fidelity receives one point rather than two. Wan 2.7 Image replaces the identity and changes the jacket construction and blue-hour treatment. Seedream 5.0 Pro preserves the jacket, greenhouse, pose, and artifact quality but replaces the identity. A hard failure counts one completed test output that misses at least one explicit pass check; it does not count every deviation within the same image separately.

Test 1 scoreProtected fidelityRole separationPrompt complianceCompositionArtifact integrityTotal / 10
GPT Image 2122229
Wan 2.7 Image011226
Seedream 5.0 Pro012227

Test 2 rewards exact product control as well as a convincing campaign image. GPT Image 2 preserves the ARCA bottle, basalt surface, warm camera-left light, and left-side negative space. Wan 2.7 Image meets the broad scene requirements but shows enough bottle-silhouette drift to miss the preserved-silhouette pass check. Seedream 5.0 Pro retains the protected product more closely, but the outdoor landscape leaks into the assigned roles and the shadow treatment is softer than the requested crisp campaign lighting.

Test 2 scoreProtected fidelityRole separationPrompt complianceCompositionArtifact integrityTotal / 10
GPT Image 22222210
Wan 2.7 Image122229
Seedream 5.0 Pro211228

Test 3 separates overall illustration quality from count accuracy. GPT Image 2 preserves the courier, the three specified parcel types, the palette, uphill movement, and clear upper-right space. Wan 2.7 Image introduces an additional coral parcel. Seedream 5.0 Pro adds multiple coral parcels and shows more character drift. Those unwanted repetitions lower both prompt compliance and artifact integrity, and each output receives one hard failure for missing the exact-three-parcels check.

Test 3 scoreProtected fidelityRole separationPrompt complianceCompositionArtifact integrityTotal / 10
GPT Image 22222210
Wan 2.7 Image211217
Seedream 5.0 Pro111216
ModelTest 1 / 10Test 2 / 10Test 3 / 10Hard failuresAverage / 10
GPT Image 29101009.7
Wan 2.7 Image69737.3
Seedream 5.0 Pro78627.0

Across this retained three-test set, GPT Image 2 ranks first with a 9.7 average and no hard failures, followed by Wan 2.7 Image at 7.3 and Seedream 5.0 Pro at 7.0. This ordering is specific to these prompts, references, settings, and first outputs; it should not be treated as a universal model ranking.

8. Which Model Fits Which Multi-Reference Workflow?

Choose GPT Image 2 when protected-detail reliability is the priority. It ranks first in this retained set, scores 9.7 on average without a hard failure, and also fits projects that need more than ten source images, quality selection, or later masked repair. Its Test 1 identity score of one still shows that a leading average does not guarantee exact facial fidelity.

Choose Wan 2.7 Image when the project is likely to extend from one composite into a related image set and its sequential-image control matters. It produces the second-highest average here and performs strongly on the product assignment, but its identity replacement, bottle-silhouette drift, and extra parcel create hard failures in all three tests. Its Vofy configuration supports up to nine image inputs and 1K or 2K output; the first-output protocol keeps sequence mode disabled so the comparison stays level.

Choose Seedream 5.0 Pro when the brief uses ten or fewer inputs and the immediate goal is one finished candidate rather than several outputs from the same request. It preserves the Test 2 bottle better than Wan 2.7 Image, but role leakage, identity replacement, and parcel duplication reduce its retained-set average to 7.0. The current Vofy configuration supports 1K or 2K output, common square, portrait, and landscape ratios, and PNG or JPEG export.

The retained outputs confirm that workflow controls and visual fidelity answer different questions. GPT Image 2 leads this specific benchmark, while Wan 2.7 Image and Seedream 5.0 Pro remain relevant when their distinct controls better match the production workflow. Run a small brief-specific test before treating this ranking as transferable to another subject or reference pack.

9. Running the Same Test on Vofy

Open GPT Image 2, Wan 2.7 Image, and Seedream 5.0 Pro in Vofy Image Studio. For each route, upload the four Test 1 references in the stated order, choose 3:4 and 2K, and paste the prompt without rewriting it. Save the native output and generation record before moving to the next model. Repeat the process for the wide product brief and square illustration brief.

Use a simple naming scheme such as t1-gpt-image-2-first.webp, t1-wan2.7-image-first.webp, and t1-seedream-5.0-pro-first.webp. Keep failed or awkward outputs. The purpose is to measure production risk, and deleting the failures would remove the most useful evidence.

Do not normalize the outputs with generative edits before publishing the comparison. Ordinary color-managed export and lossless cropping for detail boards are acceptable. If a model returns an unexpected native ratio, publish that fact and keep the uncropped image available. If a request fails technically, retry only under a declared technical-retry rule and do not count the retry as the original first output.

10. Conclusion

Multi-reference generation should be judged by preservation, not by how many images fit into an upload control. In these nine retained first outputs, GPT Image 2 produces the strongest result with a 9.7 average and no hard failures. Wan 2.7 Image averages 7.3 with three hard failures, while Seedream 5.0 Pro averages 7.0 with two.

The protocol makes that conclusion auditable by fixing reference order, prompt text, shared resolution, target ratios, first-output rules, pass checks, and a ten-point rubric before scoring. The result is specific rather than universal: GPT Image 2 is the best fit for this retained set, but every new workflow should test its protected detail first and keep failed outputs visible.

FAQ

What is multi-reference image generation?

Multi-reference image generation uses several source images in one request, with each source potentially controlling a different part of the output. Common roles include subject identity, wardrobe, product geometry, environment, lighting, palette, visual style, and composition. A strong prompt assigns each source one job instead of asking the model to "use these images" without boundaries.

Which model accepts the most reference images on Vofy?

As configured on August 7, 2026, GPT Image 2 accepts up to 16 input images in image-to-image mode, Seedream 5.0 Pro accepts up to 10, and Wan 2.7 Image accepts up to 9. These are integration limits, not quality scores. A smaller, clearly assigned reference pack can outperform a larger ambiguous one.

Does the model with the highest score automatically win?

Not necessarily. Average score should be read alongside hard failures. An identity swap, product redesign, or missing required object can make an image unusable even when its lighting and composition score well. Choose the model whose failures are least damaging for your workload.

Why request only one output from each model?

Seedream 5.0 Pro currently returns one output per request in the Vofy configuration, so one retained first output creates a common baseline. A later repeatability test can run four or more independent requests per model, but it should be reported as a separate experiment with its own sample size.

Can I optimize the prompt for each model?

Yes, but that answers a different question. This test compares default behavior under one shared, explicit brief. A model-specific optimization test should disclose every prompt difference and should not combine those results with the fixed-prompt scores.

Should I use real people or real products in a public benchmark?

Use only images you own or have permission to edit. Obtain consent from any identifiable adult and avoid misleading impersonation. A fictional product is preferable when the purpose is to test geometry and label preservation without creating an unauthorized commercial association.

Related Reading

  • Best AI Image Generators: The Core Model Test
  • GPT Image 2: What It Is and When to Use It
  • Wan 2.7 Image and Video Models on Vofy

More Articles

  • Best AI Image Generators in 2026: The Core 6-Model Test
    Best AI Image Generators in 2026: The Core 6-Model TestJul 28, 2026
  • GPT Image 2 Guide: What It Is, Key Features, and Use Cases
    GPT Image 2 Guide: What It Is, Key Features, and Use CasesApr 27, 2026
  • Wan 2.7 Image and Video Models Are Now Live on Vofy
    Wan 2.7 Image and Video Models Are Now Live on VofyJul 9, 2026
  • Best AI Image Generators in 2026: Text & Illustration Test
    Best AI Image Generators in 2026: Text & Illustration TestAug 4, 2026
  • GPT Image 2
    GPT Image 2Apr 21, 2026
  • Seedream 5.0 Pro
    Seedream 5.0 ProJul 8, 2026

Try it yourself on Vofy

Generate AI images and videos with the best models - all in one studio.

Start for free →

Discover More

Best AI Image Generators in 2026: The Core 6-Model Test
Marcus ChenMarcus Chen•Jul 28, 2026

Best AI Image Generators in 2026: The Core 6-Model Test

We test six AI image generators with the same three production prompts, keep every failure, and compare their first-pass usability rates.

GPT Image 2
Apr 21, 2026

GPT Image 2

GPT Image 2 is OpenAI's state-of-the-art image generation model for fast, high-quality image generation and editing. OpenAI positions it as a major step forward in instruction following, dense text rendering, multilingual layouts, stylistic fidelity, flexible sizing, and stronger world knowledge.

Jul 8, 2026new

Seedream 5.0 Pro

Seedream 5.0 Pro is ByteDance Seed Team's multimodal image creation model for professional design work. It improves image-text alignment, structural coherence, text rendering, visual aesthetics, dense information layouts, precision editing, realistic textures, and multilingual generation.