Upload a portrait
Start with one clear selfie, headshot, or upper-body photo where the face is easy to read and identity preservation matters.
Upload one portrait and turn it into a street interview video
— Video gallery —
Three Shibuya-style interview examples show the app's generated TV frame, portrait animation, handheld motion, and preserved broadcast-caption feel.
Street Interview Video Generator turns one uploaded portrait into a short AI clip that looks like a Japanese late-night variety-show street interview. It creates a Shibuya or Shinjuku broadcast look with telop graphics, a question box, reaction inset, microphone, handheld texture, and a shy on-camera answer.
That makes it different from a generic image-to-video prompt box. The app is tuned for one very specific result: a portrait that keeps the person's likeness, gains a Japanese street-TV atmosphere, and animates with shy eye movement, a nervous smile, microphone motion, camera shake, and preserved on-screen graphics. For another motion direction, try 2026 World Cup Live Crowd Cam Video Generator when the project needs a separate video effect rather than this exact result.
Start with one clear selfie, headshot, or upper-body photo where the face is easy to read and identity preservation matters.
The app adds the Shibuya street-interview look with Japanese telop graphics, reaction box, question panel, and microphone cues.
Generate an 8-second interview clip with subtle facial motion, microphone movement, handheld camera feel, and stable TV graphics.
Tip: 01
Use an adult subject photo that you own or have permission to transform.
Tip: 02
Choose a clear face with enough shoulder or torso detail for a believable street-interview frame.
Tip: 03
Avoid source photos that already contain lots of text, because the effect adds broadcast graphics.
Tip: 04
Keep the default prompt fixed if you want the Japanese captions and TV layout to stay stable.
Make a social-ready interview gag, creator intro, or character-style clip with the familiar look of Japanese street variety TV.
Use the romance-question box, large telop caption, and shy answer motion for a playful late-night dating-interview feel.
Turn a normal selfie into a broadcast-style reaction moment for reels, fan edits, group chats, and lighthearted posts.
Test how one portrait behaves in a locked Japanese street-TV format without writing a fresh prompt from scratch.
Start with one portrait and generate an 8-second Shibuya interview clip with the app's default TV-show look.
Use a selfie, headshot, or waist-up photo where the face is visible and the subject can plausibly fit a night street-interview frame.
Tip: Photos without heavy filters, face obstruction, or existing text usually give cleaner TV graphics.
Click Generate to prepare the Japanese variety-show frame, then click Generate again to create the interview clip with Shibuya street atmosphere, telop captions, a reaction box, and a microphone in frame.
Tip: If identity changes too much, retry with a sharper source image.
Preview the shy interview answer, then download the result once the face, microphone, camera shake, and preserved graphics look stable.
Tip: Regenerate if the model adds unwanted text or loses the TV layout.
“The preset gets the TV frame, captions, and motion into one clean run without making me write a long prompt every time.”
“Uploading one selfie and getting a full variety-show interview moment is much faster than writing a fresh prompt every time.”
“The prompt focus on preserving existing text matters. Street-interview clips fall apart fast when captions keep changing.”
It creates a short Japanese street-interview style video from one uploaded portrait, with TV variety-show captions, microphone motion, and a Shibuya night interview feel.
No. The app is designed around a single portrait upload, and the interview look is already configured.
The TV graphics, Japanese captions, reaction box, and street-interview composition need to stay coherent, so the default effect keeps those details tightly guided.
The video prompt explicitly asks the model to preserve existing on-screen text and not add new subtitles, captions, logos, or text.
Clear adult portraits, selfies, or waist-up shots with visible facial details work best. Avoid blurry photos, covered faces, or images that already contain many text overlays.
AI-generated text can vary. The effect is tuned for the overall Japanese TV interview look and for keeping the on-screen graphics stable in the final clip.
New video models, motion prompts, and one practical generation idea worth testing - quietly delivered every Friday.