Ore wa Manabe Daito
An R&D project that generates a song and music video starring a person from one portrait and a short text prompt. It connects lyric generation, music generation, script generation, shot planning, video generation, and editing into one production pipeline, testing how minimal input can become a narrative video work.
Building an MV from a photo and a song
In this MV, the song lyrics and portrait were used as the starting point. Multiple short scenes were generated while keeping the character consistent, then edited into one video.
The video includes several times and places derived from the song: a father working alone late at night, family everyday life, a festival performance, and an ending at the daughter's wedding. Runway and Seedance were used for video generation, allowing the present, memory, and an imagined future to overlap inside the same song.
The project treats AI video generation not as a single-shot visual output, but as an MV production workflow shaped by the full timeline of a song.
YouTube release
The final music video is presented as the public YouTube release. An alternate face-error version, where the daughter and wife characters break visually, is also archived in this section.
By changing only the protagonist or character reference photos, the same music-video content and structure can be generated in a one-shot workflow. The versions below keep the same song, lyric timeline, and scene structure while swapping the reference photos.
Public pipeline
This public pipeline is organized so a user can provide a theme and a The first person image, then reproduce music generation through MV generation locally. YouTube publishing is outside the scope of the reproducible pipeline.
Success and failure examples for character references
In the step that generates mom and daughter references from the daito input image, the character transformation can work well or fail. Failed outputs can retain too many features from the source image. These examples show both outcomes side by side.
The input person's atmosphere is referenced while mom and daughter are transformed into distinct characters.
When generation fails, features from the input image, such as glasses, facial hair, or facial structure, remain too strongly in mom and daughter outputs.
Right-side vertical subtitles
The subtitles are placed vertically on the right side of the image rather than along the bottom, so the lyrics become part of the frame instead of an explanatory overlay.
The design lets the text enter the empty space of the image like film subtitles. It uses only white text and a black outline, without a translucent black rectangle.
Punctuation is kept minimal. Line breaks and fade timing carry the rhythm of the lyrics.
Music and lyrics
The song and lyrics were generated with ChatGPT and Suno. The theme, father-perspective story, family relationships, and word flow for the MV were organized first, then turned into music with Suno.
Frames from the finished video
The images on this page are web derivatives from the finished video, not copies of the input reference portrait.
Credits
Main production roles and generation tools.
Related links
Public links for the video, song, and generation tools referenced in production.
Related reference pages
Open the FAQ, glossary, authority, measurement, and AI index pages. Each link now states what it is for.