← Selected Works

Selected Work / Real-time Generative Mirror

Transformirror

Daito Manabe と Kyle McDonald が開発制作した、来場者自身の姿を SDXL Turbo のリアルタイム image-to-image 変換で大型LEDスクリーンに映し返す、世界初期の公共的AIミラー・インスタレーション。

One of the earliest public AI mirror installations, developed and created by Daito Manabe and Kyle McDonald, reflecting visitors on a large LED screen through real-time SDXL Turbo image-to-image transformation.

Transformirror installation at W Osaka
W OSAKA / TRANSFORMIRROR / REAL-TIME DIFFUSION

Installation Images

W Osaka での初期展示記録。

Documentation from the original W Osaka presentation.

Overview

SDXL Turbo のリアルタイム image-to-image と Stable Audio の text-to-music を組み合わせ、体験者がウェブカムに捉えられると、その画像が即座に変換され、大型LEDスクリーンに表示される。身体、プロンプト、イメージ、音が継続的に再解釈される公共的な鏡として構成されている。

The system combines real-time SDXL Turbo image-to-image transformation with Stable Audio text-to-music. When a visitor is captured by a webcam, their image is immediately transformed and displayed on a large LED screen, creating a public mirror where body, prompt, image, and sound are continuously reinterpreted.

  1. 1. ドキュメンテーション映像1. Documentation video
  2. 2. 展示・発表歴2. Exhibitions and presentations
  3. 3. 出発点、コンセプト、技術的動機3. Background, concept, and motivation
  4. 4. Krueger からカメラ入力型作品へ4. From Krueger to camera input
  5. 5. 重要な関連映像5. Key reference videos
  6. 6. リアルタイム拡散システムとしての Transformirror6. Transformirror as real-time diffusion
  7. 7. 関連作品・技術リスト7. Related works and technologies

History / Exhibitions and Presentations

Transformirror の展示・発表歴。

Exhibitions and public presentations of Transformirror.

W OSAKA WET DECK & WET BAR

大阪・心斎橋の W Osaka で初期展示。SDXL Turbo のリアルタイム image-to-image と Stable Audio の text-to-music を組み合わせた、Transformirror の原点となる公共展示。

The original public presentation at W Osaka in Shinsaibashi, Osaka, combining real-time SDXL Turbo image-to-image transformation with Stable Audio text-to-music.

Transformirror Osaka archive

KIKK Festival 2024, KIKK in Town

ベルギー・ナミュールの KIKK in Town / True-False 展示ルートで発表。ツアー時のGPU運用・低遅延配信の課題とも接続する展示機会。

Presented in Namur, Belgium as part of KIKK in Town / the True-False exhibition route, connecting the work to touring GPU logistics and low-latency delivery.

KIKK in Town feature / Realtime diffusion in the cloud

ORGANISM New Media Art Show

ロサンゼルスで開催された ORGANISM New Media Art Show で紹介。AI、計算、インタラクティブメディアを横断する文脈の中で、リアルタイム生成AIによる鏡像体験が提示された。

Presented in Los Angeles within a program bringing together artists working across AI, computation, and interactive media.

Event information

FILE SAO PAULO 2025: SYNTHETIKA

ブラジル・サンパウロの FILE SAO PAULO 2025 に出展。合成イメージ、生成システム、現代の電子芸術を扱う国際プログラムの中で紹介された。

Presented in Sao Paulo, Brazil as part of FILE SAO PAULO 2025, within an international program on synthetic images, generative systems, and contemporary electronic art.

FILE artwork page / Daito Archive

Cinekid Festival 2025 MediaLab

オランダ・アムステルダムの Cinekid Festival MediaLab で展示。子どもや家族がインタラクティブメディアと出会う文脈の中で、生成AIを身体的に体験できる作品として位置づけられた。

Shown in Amsterdam in Cinekid's MediaLab program, emphasizing direct, embodied encounters with generative AI for family and children's media contexts.

Cinekid project page

TOKYO NIGHTTIME PROJECT, Shinjuku Neon Walk

東京・新宿中央通りで開催された TOKYO NIGHTTIME PROJECT / Shinjuku Neon Walk にて屋外展示。夜の都市空間を歩く人々を、季節感のあるAI生成イメージへ変換した。

Presented outdoors on Shinjuku Chuo-dori, translating pedestrians into seasonal AI-generated images on a public nighttime street.

TOKYO NIGHTTIME PROJECT

The Tech Interactive, AI Literacy Weekend

米国サンノゼの The Tech Interactive で「Transformirror Trickery」として紹介。来場者が小道具を使って、AI が場面をどう解釈するかを試す教育的なポップアップとして展開された。

Presented at The Tech Interactive in San Jose as Transformirror Trickery, a pop-up activity where visitors used props to test how AI interprets a scene.

The Tech Interactive

nuitnumerique #23 MUTATIONS, Saint-Ex

フランス・ランスの Saint-Ex「nuitnumerique #23 MUTATIONS」に出展。人の動きを捉え、シルエットを別の風景へ置き換えるリアルタイムの幻覚的インスタレーションとして紹介された。

Shown at Saint-Ex in Reims, France for nuitnumerique #23 MUTATIONS, as a real-time generative AI hallucination transforming visitors' silhouettes into other scenes.

Saint-Ex event page

Introduction / Background, Concept, and Technical Motivation

出発点は、ライブカメラ入力を拡散モデルで変換できるのか、という問いでした。

The starting point was whether live camera input could be transformed by a diffusion model.

このプロジェクトの出発点の一つは、2022年12月に行われた Midjourney の David Holz との会話にあります。当時、Stable Diffusion をはじめとする拡散モデルは、テキストから高品質な画像を生成する技術として急速に注目を集めていました。Latent Diffusion Models は、拡散モデルをピクセル空間ではなく潜在空間で処理することで、高解像度画像生成の計算負荷を下げる重要な基盤を提示しました。また Stable Diffusion は、2022年8月に公開され、テキストから画像を生成する AI モデルを広く一般に普及させる大きな転換点となりました。

しかし、2022年末の時点では、これらの拡散モデルは主に「1枚の画像を生成する」ことを中心に使われていました。ライブカメラ入力をリアルタイムに取得し、それを拡散モデルで変換し、さらに連続した映像として安定して表示することは、まだ大きな技術的課題でした。

その会話の中で問われたのは、非常にシンプルですが本質的なことでした。カメラから入力された映像を、拡散モデルによってリアルタイムに別のイメージへ変換することは、いつ可能になるのか。たとえば、20fps 程度の速度で、観客の動きに応答しながら画像を生成し続けることは可能なのか。

David Holz は、そのようなリアルタイム拡散型の画像変換は、2023年中には実現されるのではないかと予想しました。そしてその予想は、結果的に的中しました。2023年後半には、高速な拡散モデル推論、少数ステップ生成、リアルタイム image-to-image 変換のためのパイプライン最適化が急速に進み、当初議論していた「ライブカメラ入力を拡散モデルでリアルタイムに変換する」という構想は、実際のインタラクティブ・インスタレーションとして実装可能な段階に到達しました。

One starting point for the project was a December 2022 conversation with David Holz of Midjourney. At the time, diffusion models such as Stable Diffusion were rapidly becoming central to high-quality text-to-image generation. Latent Diffusion Models had shown how operating in latent space could reduce the computational cost of high-resolution generation, and the public release of Stable Diffusion in August 2022 made text-to-image generation broadly accessible.

Yet at the end of 2022, diffusion models were still primarily used to generate single images. Capturing live camera input, transforming it with a diffusion model, and returning it as stable continuous video remained a major technical challenge.

The question was simple but fundamental: when would it become possible to transform camera input into another image in real time with a diffusion model? Could a system generate continuously, around 20fps, while responding to the movement of a viewer?

Holz predicted that this kind of real-time diffusion-based transformation could become feasible during 2023. By late 2023, fast inference, few-step generation, and real-time image-to-image optimization had advanced enough that the idea could be implemented as an interactive installation.

01

連続するフレーム間の時間的一貫性を、どのように保つか。

How can temporal consistency be maintained across sequential frames?

02

リアルタイム・インタラクションに十分な速度と低遅延を、どのように実現するか。

How can the system reach the speed and low latency required for real-time interaction?

このプロジェクトは、カメラを記録装置ではなく、リアルタイム生成インターフェースとして再定義します。

The project redefines the camera not as a recording device, but as a real-time generative interface.

この問いは、単に生成速度の問題ではありませんでした。最も重要な課題の一つは temporal consistency、すなわち時間的一貫性です。拡散モデルは、1枚の画像を高品質に生成することには優れています。しかし、ライブカメラ映像のように連続するフレームを扱う場合、各フレームが独立して生成されると、人物の輪郭、顔、服、背景、質感、色彩、細部の形状がフレームごとに揺らぎやすくなります。その結果、出力は映像というよりも、連続して表示された不安定な静止画のように見えてしまいます。

インタラクティブ・インスタレーションにおいて、この不安定さは大きな問題になります。観客は自分の身体を動かし、その動きがスクリーン上で変換されることを期待します。しかし、映像がフレームごとに大きく揺れたり、身体の形が崩れたりすると、観客は自分の動きと AI の出力との関係を感じにくくなります。したがって、このプロジェクトにおける最初の課題は、単に AI で美しい画像を作ることではなく、連続したカメラ入力を、破綻の少ない映像体験として成立させることでした。

もう一つの大きな課題は speed and latency、すなわち速度と遅延です。インタラクティブな映像環境では、観客がカメラの前で動き、その動きがスクリーン上の映像として返ってきます。もしシステム全体の遅延が大きければ、身体感覚と視覚応答の関係は失われてしまいます。カメラキャプチャ、画像解析、拡散モデルによる生成、後処理、LED スクリーンへの出力までを含め、パイプライン全体を低遅延で動かす必要がありました。

目指したのは、カメラ映像に単純なフィルターをかけることではありません。観客の身体、動き、空間を入力として、拡散モデルがそれをリアルタイムに再解釈するシステムです。観客の身体は入力となり、拡散モデルはその入力に応答する視覚環境となります。スクリーンは単なる表示面ではなく、AI によって変換された自己像を映し返す、新しいタイプの鏡として機能します。

The question was not only about generation speed. A central issue was temporal consistency. Diffusion models can produce a strong single image, but when live camera frames are generated independently, contours, faces, clothing, backgrounds, textures, colors, and small details can fluctuate from frame to frame. The result can feel like an unstable sequence of still images rather than video.

In an interactive installation, that instability matters. Viewers expect their bodily movement to return through the screen. If the output flickers heavily or the body collapses, it becomes difficult to feel the connection between bodily action and AI response.

The second challenge was speed and latency. Camera capture, image analysis, diffusion generation, post-processing, and LED output all needed to run with sufficiently low latency. The goal was not to put a filter on camera video, but to build a system where the viewer's body, movement, and space are reinterpreted by a diffusion model in real time.

1. From Krueger's Vision / Early Interactive Input

身体を入力として、映像環境はどのように応答できるのか。

How can a visual environment respond to the body as input?

Transformirror を歴史的に位置づけるためには、Myron Krueger の仕事から始める必要があります。Krueger は 1970年代から「responsive environment」という概念を展開しました。1977年の論文 “Responsive Environments” では、人間の行動を知覚し、それに対して音響的・視覚的なフィードバックを返す環境が論じられています。

代表的な作品である Videoplace は、参加者の身体を映像として取り込み、コンピュータ生成された環境の中で相互作用させるものです。ここで重要なのは、Krueger にとってカメラやセンサーは単なる記録装置ではなかったという点です。カメラは、身体を読み取るための入力装置であり、環境が人間の存在や動きに応答するための知覚器官でした。コンピュータと人間の関係は、キーボードやマウスを介した命令ではなく、身体的な動きそのものによって成立していました。

この考え方は、現在のリアルタイム生成 AI インスタレーションにも直接つながっています。Transformirror のようなシステムでも、観客はキーボードで命令を入力するのではありません。観客は自分の身体でシステムに関わります。動く、立つ、近づく、離れる、ポーズを取る。その身体的な入力がカメラによって取得され、AI モデルによって変換され、映像環境として返ってきます。

1990年代末から2000年代にかけて、カメラ入力を用いたインタラクティブ作品は、より詩的で空間的な表現へ発展しました。Camille Utterback と Romy Achituv による Text Rain は、参加者のシルエットが落下する文字と関わる作品です。Golan Levin、Zachary Lieberman らによる Messa di Voce は、声や歌をリアルタイムの合成グラフィックスとして可視化する作品です。Reface [Portrait Sequencer] は、観客の顔をカメラで取得し、目、口、眉などの映像断片を分割・再構成します。

これらの作品に共通しているのは、カメラ入力を使って人間の姿や動きをリアルタイムに解析し、それを別の視覚体験へ変換している点です。ただし、この時代の変換は、主にシルエット、輪郭、顔の部位、身体位置、音声解析などを使ったルールベース、あるいは比較的明示的な変換でした。それでも、この系譜は重要です。現在の AI によるリアルタイム映像変換も、根本的には同じ問いを引き継いでいるからです。

Kinect 以前の多くの作品では、コンピュータが見ていたのは厳密には「人間」ではなく、背景差分、シルエット抽出、blob tracking、optical flow によって得られる「動いているピクセルの塊」でした。肘や膝、骨格、姿勢、身体部位を安定してリアルタイムに扱うことは難しかったのです。

2010年の Kinect は、この流れを大きく変えました。重要だったのは RGB カメラだけではなく、民生化された深度センサーであり、身体を距離、骨格、部位を持つ構造として扱えるようになったことでした。さらに openFrameworks のようなクリエイティブコーディング環境があったことで、Kinect は研究室の技術ではなく、展示やパフォーマンスで使えるリアルタイム表現の素材になりました。

Transformirror は Kinect のように骨格推定を主役にする作品ではありません。しかし、観客の身体をただの映像ではなく、システムが読み取り、応答し、別の視覚環境へ返す入力として扱う点で、Kinect 以降の身体解析の歴史を引き継いでいます。

To position Transformirror historically, it is useful to begin with Myron Krueger. From the 1970s onward, Krueger developed the idea of the responsive environment: an environment that senses human behavior and returns audiovisual feedback.

In Videoplace, the participant's body is captured as an image and interacts inside a computer-generated environment. For Krueger, the camera was not merely a recording device. It was an input device for reading the body, a sensing organ through which the environment could respond to presence and movement.

From the late 1990s through the 2000s, works such as Text Rain, Messa di Voce, and Reface [Portrait Sequencer] transformed silhouettes, voices, and faces into live visual experiences. Before Kinect, many such systems relied on background subtraction, blob tracking, and optical flow; they often saw moving pixels rather than a structured body.

Kinect changed this trajectory in 2010 by making depth sensing and skeletal structure available as a practical creative material. Transformirror is not a Kinect-style skeleton-tracking work, but it inherits the shift toward treating the body as structured input that can be read, responded to, and returned as another visual environment.

2. Style Transfer and GANs / The Road to Real-Time

画像生成は、静止画からリアルタイム変換へ向かい始めました。

Image generation began moving from still images toward real-time transformation.

2010年代半ばになると、ディープラーニングによる画像生成・画像変換が急速に発展しました。2014年、Goodfellow らは Generative Adversarial Networks、すなわち GAN を提案しました。GAN は、生成モデルと識別モデルを競わせながらデータ分布を学習する枠組みであり、後の画像生成、顔生成、画像変換、映像合成に大きな影響を与えました。

2015年には、Google Research の DeepDream / Inceptionism が大きな注目を集めました。この技術は、ニューラルネットワークの内部表現を可視化し、ネットワークが画像の中に見出す特徴を増幅する方法として紹介されました。同時に、ニューラルネットワークがアーティストのための新しい視覚表現の道具になりうることも示していました。

同じ 2015年、Gatys、Ecker、Bethge による A Neural Algorithm of Artistic Style は、画像の content と style をニューラルネットワークの特徴表現として分離し、それらを再結合する方法を示しました。この研究は、いわゆるニューラル・スタイル転送の基礎となり、1枚の画像を別の絵画的スタイルへ変換する表現を広く普及させました。

ただし、初期のスタイル転送は計算負荷が高く、ライブカメラ入力に対して高いフレームレートで動かすことは容易ではありませんでした。2016年、Johnson、Alahi、Fei-Fei による Perceptual Losses for Real-Time Style Transfer and Super-Resolution は、フィードフォワードネットワークを用いることで、従来手法より大幅に高速なスタイル転送を可能にしました。

ここで重要な転換が起きました。画像生成・画像変換は、「1枚の画像を時間をかけて生成する」ものから、「入力画像を高速に別の見え方へ変換する」ものへ向かい始めたのです。ライブカメラ映像にスタイルを適用する実験も増え、人物や風景をリアルタイムに絵画的、抽象的、あるいは別の質感を持つ映像へ変換する表現が可能になっていきました。

その後、pix2pix、CycleGAN、vid2vid などの image-to-image / video-to-video synthesis が登場しました。pix2pix は、条件付き GAN を用いて入力画像から出力画像への対応関係を学習する汎用的な枠組みを提示しました。CycleGAN は、対応するペア画像が存在しない場合でも、ある画像ドメインから別の画像ドメインへ変換する方法を示しました。vid2vid は、入力映像から出力映像を生成する枠組みを提示し、時間的なダイナミクスを考慮しないまま画像合成手法を動画に適用すると、時間的に不安定な映像になりやすいことを指摘しました。

この流れの中で、Memo Akten の Learning to See は非常に重要な先行例です。複数のニューラルネットワークがライブカメラ映像を解析し、テーブル上の物体をリアルタイムに再解釈します。観客は手で物体を動かし、その入力がニューラルネットワークによって風景や別のイメージとして再構成される様子を見ることができます。この作品は、AI が世界をそのまま見るのではなく、学習したデータセットやモデルの構造を通じて世界を再解釈することを体験的に示していました。

In the mid-2010s, deep learning rapidly changed image generation and image transformation. GANs introduced a powerful framework for learning data distributions through competition between generator and discriminator. DeepDream showed that neural networks could become tools for visual expression by amplifying internal features.

Neural style transfer separated content and style as feature representations and recombined them, but early methods were not practical for high-frame-rate live camera input. Feed-forward approaches to real-time style transfer shifted the field toward faster image transformation.

pix2pix, CycleGAN, and vid2vid expanded image-to-image and video-to-video synthesis. In particular, vid2vid made temporal stability a central problem, directly anticipating the temporal consistency issues later faced by real-time diffusion systems.

Memo Akten's Learning to See is a crucial precedent. It used live camera input and neural networks to reinterpret objects on a table as other visual worlds, showing that AI does not simply see reality; it rereads it through training data, model structure, and learned associations.

問い 内容 Question Content
入力とは何か カメラ映像、顔、身体、シルエット、ポーズ、深度。 What is input? Camera image, face, body, silhouette, pose, depth.
変換とは何か スタイル転送、ドメイン変換、顔変換、映像合成。 What is transformation? Style transfer, domain translation, face transformation, video synthesis.
時間性とは何か 単一画像ではなく、連続フレームとして破綻しないこと。 What is temporality? Not a single image, but a stable sequence of frames.
体験とは何か 観客が自分の身体を通じて AI の出力に関わること。 What is experience? A viewer engaging AI output through their own body.

Reference Videos

重要な関連作品の映像。

Key related works in video.

Transformirror を理解する上で重要な、身体入力、カメラによる応答環境、音声・文字・生成モデルによるリアルタイム変換の系譜を映像でたどる。

These videos trace important precedents for Transformirror: body input, camera-responsive environments, and real-time transformations through text, voice, and learned visual models.

Myron Krueger - Videoplace, Responsive Environment, 1972-1990s

身体をカメラ入力として扱い、映像環境が人間の存在や動きに応答するという発想の原点。

An origin point for treating the body as camera input and making a visual environment respond to human presence and movement.

Camille Utterback & Romy Achituv - Text Rain, 1999

観客のシルエットが文字の落下に介入する、カメラ入力型インタラクションの代表的作品。

A canonical camera-based interaction in which the viewer's silhouette intervenes in falling text.

Messa di Voce

声、身体、映像をリアルタイムに接続し、音声を空間的な視覚表現へ変換する作品。

A work connecting voice, body, and image in real time, transforming vocal performance into spatial visual expression.

Learning to see: Gloomy Sunday, 2017

ライブカメラの現実を、学習済みの視覚世界を通してリアルタイムに読み替える重要な先行例。

An important precedent for rereading live camera reality through a learned visual world in real time.

Daito Manabe + Kyle McDonald + Rhizomatiks - Generative MV at ICC, 2023

坂本龍一「Perspective」の映像背景を、観客のテキスト入力に応じてリアルタイムに変化させた関連展開。

A related deployment in which the background of Ryuichi Sakamoto's "Perspective" changed in real time in response to audience text input.

3. Breaking New Ground / Our Diffusion-Based Interactive System

拡散モデルを、ライブカメラ入力と接続されたリアルタイム・システムとして扱います。

Diffusion is treated as a real-time system connected to live camera input.

2023年後半になると、拡散モデルをリアルタイム・インタラクションに接続するための技術的条件が急速に整い始めました。一つの重要な流れは、少数ステップでの高速生成です。Latent Consistency Models は、従来の拡散モデルが多数の反復ステップを必要とするため推論が遅いという問題に対して、少ないステップで高品質な画像を生成する方法を提示しました。

また、Stability AI は2023年11月に SDXL Turbo を発表しました。SDXL Turbo は Adversarial Diffusion Distillation に基づくモデルであり、1ステップで画像を合成し、リアルタイム text-to-image 出力を可能にする重要な進展でした。さらに、2023年12月には StreamDiffusion が発表されました。StreamDiffusion は、リアルタイム・インタラクティブ生成のための拡散パイプラインとして設計されており、連続入力を扱うライブビデオ、メタバース、放送などの状況で、既存の拡散モデルがリアルタイム性に課題を持つことを明確に指摘していました。

この技術的変化の中で、当初の問いを実際の作品として実装することが可能になりました。それが、Daito Manabe と Kyle McDonald による Transformirror です。Transformirror は、SDXL Turbo を用いたリアルタイム image-to-image 変換と Stable Audio の text-to-music 技術を統合し、体験者がウェブカムに捉えられると、その画像が即座に変換され、大型 LED スクリーンに表示される作品として公開されました。

Kyle McDonald も後に、2023年12月に Daito Manabe と Rhizomatiks とともにリアルタイム拡散ツールキットを実装したこと、そのツールキットが SDXL Turbo を基盤としていたことを記述しています。このページでは背景、文脈、コンセプトを扱い、ZeroMQ、WebRTC、WebSocket、GPU ワーカー、RunPod / Vast、展示ごとの差分などの実装詳細は Kyle McDonald の資料へ導線を置きます。

By late 2023, the technical conditions for connecting diffusion models to real-time interaction were rapidly emerging. Few-step generation, Latent Consistency Models, SDXL Turbo, and StreamDiffusion opened a practical path for connecting live input to diffusion systems.

Within that technical shift, the initial question could become an actual artwork: Transformirror by Daito Manabe and Kyle McDonald. The work integrated real-time image-to-image transformation with SDXL Turbo and text-to-music techniques from Stable Audio. When a participant was captured by webcam, the image was transformed immediately and displayed on a large LED screen.

Kyle McDonald's technical history documents the implementation path, including cloud deployment, low-latency transport, GPU workers, and exhibition-specific branches. This complete page focuses on the conceptual and historical background while directing readers to Kyle's material for engineering detail.

1

Input

カメラで観客や空間をリアルタイムに取得します。

Capture viewers and space in real time.

2

Analysis / Preprocess

人物領域、ポーズ、深度、構図などを抽出・安定化します。

Extract and stabilize figure, pose, depth, or composition when needed.

3

Diffusion Transformation

プロンプトや画像条件を使い、拡散モデルで入力映像を生成的に変換します。

Use prompts and image conditions to transform input through diffusion.

4

Temporal Stabilization

フレーム間の揺れや破綻を抑え、連続映像として成立させます。

Reduce flicker and failures so the result works as continuous video.

5

Output

大型 LED スクリーンや音響環境へ出力し、観客の身体体験へ戻します。

Return the result to the body through LED display and sound.

通常の鏡は、現実をそのまま反射します。Transformirror は、AI モデルの想像力を通して現実を再生成します。

An ordinary mirror reflects reality. Transformirror regenerates reality through the imagination of an AI model.

このプロジェクトにおいて重要だったのは、拡散モデルを単に画像生成装置として使うのではなく、ライブカメラ入力と接続されたリアルタイム・システムとして扱った点です。観客はカメラの前に立ち、システムはその姿を取得します。拡散モデルはその入力をプロンプトや条件に基づいて変換し、出力はスクリーン上に即座に返されます。

従来のスタイル転送では、入力画像の構図や形を保ちながら、別の絵画や画像の質感を重ねることが中心でした。一方、拡散モデルによる変換では、入力された身体や空間を保ちつつも、プロンプトによって世界観そのものを変えることができます。観客は自分の姿を見ていますが、その姿は AI モデルによって別の存在、別の空間、別の物語へと再生成されます。

そのため、Transformirror は単なる映像フィルターではありません。むしろそれは、AI によるリアルタイムの「鏡」です。観客は自分自身を見ます。しかし、その自己像は物理的な反射ではなく、拡散モデルによって解釈された自己像です。身体は入力であり、AI は変換装置であり、スクリーンは新しい種類の鏡として機能します。

The important point was not simply using diffusion as an image generator, but treating it as a real-time system connected to live camera input. The participant stands in front of a camera; the system captures the body; the diffusion model transforms the input through prompts and conditions; the result returns immediately on the screen.

Transformirror is not merely a filter. It is closer to a real-time AI mirror. The participant sees themselves, but the image is not a physical reflection; it is an image interpreted by a diffusion model. The body is input, AI is the transformation system, and the screen becomes a new kind of mirror.

This page does not claim Transformirror as the first in absolute terms. More precisely, it positions the work as one of the earliest public interactive installations, by late 2023, to transform live camera input through diffusion with low latency.

4. Future Horizons / From Pioneering to Accessible Tools

先駆的実装から、より開かれた制作環境へ向かいます。

From pioneering implementation toward more accessible creative tools.

Transformirror の意義は、作品そのものにとどまりません。このプロジェクトは、リアルタイム拡散モデルが、今後のインタラクティブ映像制作、ライブ演出、バーチャルプロダクション、放送、展示、教育、パフォーマンスへ拡張できることを示す初期の実践でもありました。

初期段階では、このようなシステムを実現するために、非常に複雑な実装が必要でした。カメラキャプチャ、GPU 推論、モデル最適化、並列処理、映像出力、LED 制御、プロンプト管理、音響生成など、多くの要素を低遅延で統合しなければなりませんでした。当時、AI によるコーディング支援はすでに登場し始めていましたが、実際の展示環境で安定して動くリアルタイム・システムを作るには、低レベルの映像処理や GPU 処理、パイプライン設計に関する高度な知識が必要でした。そのため、Kyle McDonald のような専門的なクリエイティブコーダー/エンジニアとの協働が不可欠でした。

一方で、2024年以降、このような技術はより多くのクリエイターに開かれつつあります。StreamDiffusion のようなリアルタイム image-to-image 変換のためのパイプライン、TouchDesigner のようなリアルタイム・ビジュアル開発環境、生成 AI 向けの最適化ライブラリが接続されることで、かつては高度な専門家だけが実装できたシステムが、より多くのアーティストやデザイナーにとって利用可能になっていくでしょう。

2023年時点のリアルタイム拡散映像は、常に完全に滑らかで安定していたわけではありません。生成結果にはちらつきやノイズがあり、身体の輪郭やディテールが揺らぐこともありました。しかし、その不完全さ自体が、生成 AI がリアルタイム映像環境へ入り始めた初期段階の特徴でもありました。重要なのは、その時点で「これは将来的に必ず改善される」と確信できたことです。

The significance of Transformirror extends beyond the work itself. It was also an early demonstration that real-time diffusion models could expand into interactive image production, live performance, virtual production, broadcasting, exhibitions, education, and public-space experiences.

At the beginning, realizing such a system required a complex implementation. Camera capture, GPU inference, model optimization, parallel processing, video output, LED control, prompt management, and sound generation all needed to be integrated with low latency. AI coding assistance was beginning to appear, but building a real-time system stable enough for exhibition still required deep knowledge of low-level video processing, GPU processing, and pipeline design.

Since 2024, these techniques have become more accessible. Real-time image-to-image pipelines, visual development environments such as TouchDesigner, and optimization libraries for generative AI point toward a future where systems once limited to specialists become available to more artists and designers.

Installation

観客の身体をリアルタイムに生成映像へ変換します。

Transform viewers' bodies into generated moving images in real time.

Performance

ダンサー、演奏者、観客の動きに応答する AI 映像演出です。

AI visuals responding to dancers, performers, and audiences.

Virtual Production

グリーンスクリーンやライブ撮影背景を生成 AI で動的に変換します。

Dynamically transform green-screen or live-shot backgrounds with generative AI.

Broadcast / Music Video

生放送や音楽映像でリアルタイム AI 変換を使います。

Use real-time AI transformation in broadcasts and music videos.

Education / Research

AI が画像をどう解釈するかを身体的に理解する体験です。

A physical way to understand how AI interprets images.

Public Space

都市スクリーンや建築ファサードと身体入力を接続します。

Connect bodily input with urban screens and architectural facades.

今後の発展には倫理的な課題もあります。観客の身体や顔をカメラで取得し、それを生成モデルで変換する場合、同意、プライバシー、データ保存、肖像、モデルの偏り、表象の問題を慎重に扱う必要があります。

さらに、AI が「何を見ているのか」という問題も重要です。Memo Akten の Learning to See が示したように、AI は世界をそのまま見るのではありません。AI は、学習データ、モデル構造、与えられた条件を通じて世界を再解釈します。Transformirror もまた、観客をそのまま映す鏡ではなく、モデルが学習した視覚世界を通して観客を再生成する鏡です。

したがって、リアルタイム拡散インスタレーションは、AI の能力を見せるだけのものではありません。それは、AI の想像力、偏り、限界、そして人間の身体との関係を体験させる装置でもあります。

Future development raises ethical questions. When a system captures viewers' bodies or faces and transforms them with a generative model, consent, privacy, data storage, likeness, model bias, and representation need to be handled carefully.

AI does not see the world as it is. It reinterprets the world through training data, model structure, and given conditions. Transformirror is not a mirror that simply reflects the viewer, but a mirror that regenerates the viewer through the visual world learned by a model.

Conclusion

身体と映像環境の系譜を、拡散モデルによって新しい段階へ進める試みです。

A lineage of body and image environments moves into a new phase through diffusion models.

このプロジェクトは、カメラ入力によるインタラクティブ・アートの長い系譜の中に位置づけられます。Myron Krueger は、身体の動きに応答する環境を構想しました。Text Rain や Messa di Voce、Reface などの作品は、カメラ、シルエット、顔、声、身体をリアルタイムに別の視覚表現へ変換しました。2010年代には、DeepDream、スタイル転送、GAN、pix2pix、CycleGAN、vid2vid、Learning to See などを通じて、機械学習による画像変換と映像変換の表現が発展しました。

そして2020年代前半、拡散モデルは画像生成の中心的技術となりました。しかし、1枚の画像を生成することと、ライブカメラ入力をリアルタイムに変換することの間には、大きな技術的隔たりがありました。その隔たりを埋めるためには、時間的一貫性、低遅延、高速推論、パイプライン最適化、そして身体的インタラクションの設計が必要でした。

2022年12月の David Holz との会話で立てられた問いは、まさにこの隔たりに関するものでした。カメラ入力を拡散モデルでリアルタイムに変換することは、いつ可能になるのか。David Holz は、それが2023年中には実現されるのではないかと予想しました。そして2023年後半、その予想は現実になりました。高速推論、少数ステップ生成、SDXL Turbo、StreamDiffusion などの技術的進展によって、リアルタイム拡散変換は実装可能な段階に入りました。

Transformirror は、その転換点に位置するプロジェクトです。これは、AI で画像を生成する作品であると同時に、カメラ、身体、拡散モデル、LED スクリーン、音響、プロンプト、リアルタイム処理を統合したインタラクティブな視覚環境です。このプロジェクトの本質は、カメラを記録装置から生成インターフェースへと変えることにあります。

それは、単に AI が画像を生成する時代から、AI が身体的・空間的・リアルタイムな環境として体験される時代への移行を示しています。

Transformirror belongs to a long lineage of camera-input interactive art. Krueger imagined environments responsive to bodily movement; later works transformed silhouettes, faces, voices, and bodies into live visual forms; machine-learning works then expanded image and video transformation through DeepDream, style transfer, GANs, pix2pix, CycleGAN, vid2vid, and Learning to See.

In the early 2020s, diffusion models became central to image generation. But there was still a large gap between generating a single image and transforming live camera input in real time. Closing that gap required temporal consistency, low latency, fast inference, pipeline optimization, and the design of bodily interaction.

Transformirror sits at that turning point. It is both an AI image-generating artwork and an interactive visual environment that connects camera, body, diffusion model, LED display, sound, prompts, and real-time processing.

Related Works and Technologies

関連作品・技術の完全リスト。

Complete related works and technologies list.

Period Work / Technology 位置づけ Position
1970s- Myron Krueger, Videoplace responsive environment、身体入力型インタラクションの原点。 Origin point for responsive environments and body-input interaction.
1999 Camille Utterback & Romy Achituv, Text Rain 身体シルエットと文字のインタラクション。 Interaction between body silhouettes and falling text.
2003 Golan Levin et al., Messa di Voce 声・身体・リアルタイム映像生成の統合。 Integration of voice, body, and real-time graphics.
2007 Golan Levin & Zachary Lieberman, Reface [Portrait Sequencer] 顔画像の取得・分割・再構成。 Capturing, segmenting, and recomposing facial video fragments.
2010s Kinect / Depth Camera Systems 深度情報と骨格認識の普及。 Widespread use of depth information and skeleton recognition.
2018 OpenPose body、foot、hand、face keypoints を含むリアルタイム複数人物姿勢推定のオープンソースシステム。 Open-source real-time multi-person pose estimation including body, foot, hand, and face keypoints.
Period Work / Technology 位置づけ Position
2014 Generative Adversarial Networks 生成モデルの大きな転換点。 A major turning point in generative modeling.
2015 Google DeepDream / Inceptionism ニューラルネットワークの視覚特徴を増幅。 Amplifying visual features found by neural networks.
2015 Gatys et al., Neural Style Transfer content と style の分離・再結合。 Separating and recombining content and style.
2016 Johnson et al., Real-Time Style Transfer 高速スタイル転送。 Fast feed-forward style transfer.
2016/2017 pix2pix 条件付き GAN による image-to-image translation。 Image-to-image translation with conditional GANs.
2017 CycleGAN ペアなし画像変換。 Unpaired image-to-image translation.
2017- Memo Akten, Learning to See ライブカメラ入力をニューラルネットワークで再解釈。 Reinterpreting live camera input through neural networks.
2018 vid2vid 時間的一貫性を意識した video-to-video synthesis。 Video-to-video synthesis with attention to temporal consistency.
Period Technology / Work 位置づけ Position
2021/2022 Latent Diffusion Models / Stable Diffusion 高品質な画像生成を広く普及させた基盤。 Foundation that made high-quality image generation broadly accessible.
2023 ControlNet pose、depth、edge、segmentation などによる条件制御を拡散モデルに追加。 Adds conditional control to diffusion models using pose, depth, edge, segmentation, and other signals.
2023 Latent Consistency Models 少数ステップでの高速生成。 Fast few-step generation.
2023 SDXL Turbo 1ステップ生成・リアルタイム text-to-image への重要な進展。 Important advance toward one-step generation and real-time text-to-image.
2023 StreamDiffusion リアルタイム・インタラクティブ生成のための拡散パイプライン。 A diffusion pipeline for real-time interactive generation.
2023 Transformirror ライブカメラ入力をリアルタイム拡散変換するインタラクティブ・インスタレーション。 An interactive installation transforming live camera input through real-time diffusion.

Core Statement and Key Terms

公開文で使う中心表現。

Core language for public descriptions.

English

Transformirror redefines the camera as a real-time generative interface: the viewer's body becomes the input, and the diffusion model becomes a responsive visual environment.

Japanese

Transformirror は、カメラをリアルタイム生成インターフェースとして再定義します。観客の身体は入力となり、拡散モデルはそれに応答する視覚環境となります。

日本語で言いたいこと 自然な英語表現 Meaning Natural English expression
このプロジェクトの出発点The starting point of this projectStarting pointThe starting point of this project
David Holz との会話a conversation with David HolzConversationa conversation with David Holz
2023年中には実現すると予想したhe predicted that it could become feasible within 2023Predictionhe predicted that it could become feasible within 2023
リアルタイム画像変換real-time image-to-image transformationTransformationreal-time image-to-image transformation
連続フレームの一貫性temporal consistency across sequential framesTemporal consistencytemporal consistency across sequential frames
身体を入力として使うuse the body as a real-time input interfaceBody as inputuse the body as a real-time input interface
先行作品への敬意を払うacknowledge and respect prior workPrior workacknowledge and respect prior work
最初期の実践の一つone of the earliest practical examplesPositioningone of the earliest practical examples
記録装置ではなく生成インターフェースnot as a recording device, but as a generative interfaceCamera conceptnot as a recording device, but as a generative interface
AI Reference Surface

関連する参照ページ

FAQ、用語集、外部参照、計測、AI用索引へ移動できます。それぞれのリンク先で何を確認できるかを明記しました。