Filmtools
Filmmakers go-to destination for pre-production, production & post production equipment!
Shop Now
Picture a Korean period drama. Recognizable actors, a director, a DP, a gaffer, hair, makeup, costume. The camera is handheld. The performances are real.
The set is an office space.
At SIGGRAPH, Beeble CEO and co-founder Hoon Kim showed me footage from a Korean production experimenting with a very different way of making films. The crew photographed real actors, real costumes and real camera movement, then used Beeble’s SwitchX model to rebuild the world around them.
Shots came back while the crew was still standing on set. Hoon puts the turnaround at about five minutes—fast enough to look at the result, change something and shoot it again.
Let’s get one thing out of the way: this isn’t virtual production. There’s no LED wall showing the actors the final environment. It also isn’t the fully synthetic text-to-video AI slop currently filling everyone’s feeds with strangely elastic people walking through their own legs.
It sits somewhere in the gap between the two. The performance and photography are captured the old-fashioned way. Almost everything else in the frame remains up for grabs until later.
The workflow is real enough—and interesting enough—to raise the hard question: Is this where filmmaking is going?
The accidental VFX company
Beeble didn’t set out to become a film company. Its five co-founders were AI researchers with zero production experience.
“None,” Hoon said when I asked whether any of them had worked in the industry. “We’re just completely nerd AI researchers.”
Hoon’s own background ran through self-driving cars, medical imaging and text-to-speech. What pulled him toward film was Google’s Total Relighting research, which demonstrated that the illumination on a photographed subject could be changed after capture.
The first product Beeble built around that idea was a consumer selfie app. Upload a photograph, pick a template, change the lighting and background. It found an audience in India.
It did not find much revenue. Then VFX artists started calling.

Green-screen work has always demanded a certain amount of faith. The cinematographer lights a subject for an environment that may not exist yet. If the eventual background’s lighting sources don’t match the lighting on the greenscreen set—or the director changes their mind in week three of post—the foreground and background stop looking as if they were ever in the same space.
Relighting after capture offered a way out. The artists saw a fix and Beeble found a market.
Four years on, the company has assembled a broad set of AI-assisted VFX tools: relighting, background replacement, roto, upscaling, and HDR conversion.
The original technology became SwitchLight, a video-to-PBR system that extracts physically based rendering information from ordinary footage. Beeble built its own light stage to capture training data and published the SwitchLight research at CVPR 2024. The current model generates base-color, normal, roughness, metallic and specular passes, along with depth and alpha.
In practical terms, it attempts to make photographed footage behave a little more like a 3D render. Rather than being permanently stuck with whatever light was baked in on the day, a compositor can bring those generated passes into Beeble’s Nuke tools and add point lights, directional lights or HDRI illumination.
Switch anything?
SwitchLight breaks a shot into usable passes so an artist can relight it. SwitchX is doing something considerably more aggressive.
It’s a video-to-video diffusion model. It can replace a background, change clothing, add or remove objects and substantially rework the lighting and atmosphere. What separates it from the usual generative-video workflow is that it is built around holding onto selected parts of the source footage.
Two controls matter: a mask and a reference image.
The mask decides what survives and what gets replaced. Depending on the mode, SwitchX can find a foreground subject automatically, let you select specific objects, process the whole frame or accept a custom alpha matte created in Nuke, After Effects or whatever other application you persuaded into producing it.
This mask is what separates SwitchX from other commercial and open-source products on the market: while it regenerates the foreground with new lighting, the software is careful to full preserve the actor’s performance, and facial identity.

The reference image decides what the transformed shot should look like: environment, lighting, color, wardrobe and general visual intent. It can be AI-generated. It can also be a photograph, a concept frame or a Photoshop composite somebody spent a morning art-directing.
So the prompt isn’t the center of the process. For anyone who works in VFX and has tried to wrangle AI into their pipeline, that alone is something of a relief.
Beeble describes the reference as a visual blueprint and the alpha mask as the control over what changes. The SwitchX documentation is unusually detailed about how those two inputs interact, including where camera tracking can fail when the retained foreground contains too little information for the model to infer the move.
None of this makes the generative risk disappear. Any time footage passes through a model like this, details can shift that you didn’t ask to shift. But the source photography acts as a much harder constraint than it does in a conventional text-to-video workflow.
The objective isn’t to invent a fresh performance that bears a passing resemblance to the original actor. It’s to carry the actual performance, timing and camera movement into a different environment.
That distinction matters a lot when your cast expects their faces to remain attached to their own heads.
Welcome to the hybrid film set
So how do you photograph a scene when you know it will eventually pass through SwitchX?
One answer is to light everything flat and neutral and make the creative decisions later. Hoon Kim has seen productions do exactly that, and it does buy considerable freedom in post.
“We’ll fix all of it later” has a long history in this industry, though, and not a particularly distinguished one.

Instead, getting lighting close on set seems to be a better strategy. Hoon’s strategy of onset previews using SwitchX really pays off here.
SwitchLight on its own doesn’t replace good photography. Hoon was candid that aggressive relighting can look plastic, particularly on skin. PBR maps can get you a long way, but subsurface scattering and the small, high-frequency details that read as skin rather than a rendering of skin remain a challenge even in big ILM/Weta pipelines.
SwitchX will make larger changes, but it still benefits from useful information in the source plate.
In one example Hoon showed me, the filmmakers wanted flickering light playing across the actors. SwitchX couldn’t invent temporally changing illumination with sufficient control, so the crew created the flicker practically during photography. The model then amplified it as it rebuilt the scene.
That’s what a hybrid set actually looks like in practice. You aren’t necessarily building or lighting the entire finished world. You are photographing the pieces the model cannot reliably infer—and the pieces that change what the actors are doing.
Camera movement tells the same story. The Korean test footage was handheld, and that physical movement carried through into the generated environment. Nobody asked the model how the camera should move. An operator had already made that decision.
The traditional set isn’t disappearing here so much as changing jobs. Instead of constructing every visible element of the final frame, the crew concentrates on capturing the performance, interaction, motion and lighting cues that anchor everything else.
Location scouting comes into its own

One of the louder claims around generative AI is that it does away with whole departments. Why scout a location when you can generate one? Why build a costume when you can prompt one?
The production Hoon Kim described ran into reality fairly quickly.
They needed imagery from a specific period of Korean history. A general-purpose image generator would happily produce something that looked broadly old and broadly Korean, but it had no dependable grasp of the architectural distinctions the production required.
So the location team went to the real place and photographed it from the angles the filmmakers planned to use. Those photographs became the source for the generated backgrounds. The model could shift the angle, change the light and fill in surrounding detail, but the historical foundation came from someone who knew what the building was supposed to be.

The rapidly emerging field of gaussian splatting leans even heavier into this workflow: fully capture an environment, then synthesize the exact camera angles needed for the final production shots.
Costumes went the same way. Asking a general model for a specific traditional Korean garment didn’t reliably produce the right thing. Renting the real costume for a few hours was cheap, fast and correct.
This is where the more confident predictions about production crews disappearing get murky. The tool reduces the amount of physical material a shot requires and, in the same motion, makes good source material more valuable than it was before.
You may be asking: what’s the point then? If you’re going to use real locations and real costumes, why bother with the AI process at all? At the end of the day it comes down to production expense. Renting an historical Korean palace grounds for a week with an entire film crew would be a major production expense. Depending on clearances, a location scout gaussian splat capture could cost two hours of the scout’s time and maybe a $15 admission to the space.
Or think of a shoot in Paris in front of the Eiffel tower: it’s orders of magnitude cheaper to fly a single scout to capture the environment than an entire crew. And golden hour? Golden hour lasts forever with AI relighting.
Crews will shrink, but content production will expand
It’s hard to look at a workflow like this and not recognize that there will absolutely be less need for large crews. I can see a world where a new genre of filmmaking with synthetic actors opens up (I mean, we’re already OK with watching Pixar movies with non-human actors), but I find it hard to believe audiences will move away from wanting to see real humans perform in most of our on-screen content. So actors are safe, but what happens to all the gaffers, grips, and schleppers on a set?
The argument can be made that with the accessibility AI brings, there will be an expansion of content being made. The indie filmmaker who dreams of creating the next space opera trilogy all of a sudden has a viable production pathway without a $200,000,000 budget. So more content, more—albeit smaller—crews.
Does that amount to financially viable careers? That’s harder to prognosticate. We will have to wait and see whether market forces squeeze wages, or make hard-won expertise more valuable.
AI is very good at producing plausible detail. Plausible and correct are different things—a distinction the film industry was arguing about long before anyone started prefixing everything with “AI.”

Hoon’s framing is that SwitchX and SwitchLight amplify what the crew does. That’s a considerably more believable near-term future than a director sitting alone in a dark room, typing increasingly desperate prompts into a laptop.
The shoot Hoon described wasn’t crewless. Director, actors, DP, gaffer, camera operators, hair, makeup, costume and location work were all present. What changed was that each department produced the essential information for the shot rather than every final pixel.
Then there’s the audience. Hoon noted that Korean viewers have been no more uniformly enthusiastic about visibly AI-generated content than Western audiences. A badly received, fully generated sequence in a television drama drew exactly the comments you would expect.
Audiences may not care which tools were used. They notice when the result looks bad. That particular quality-control mechanism has proven remarkably resistant to disruption.
The AI result is a pass, not a result
“The AI result is a pass, not a result.”
That was probably the most useful thing Hoon Kim said during our conversation, and it’s where Beeble’s approach starts to look less like consumer generative video and more like an emerging VFX pipeline.
The SwitchX output doesn’t have to be the finished shot. Bring it into Nuke and combine it with the original photography. Pull high-frequency detail from the source and layer it back over the generated material. Use SwitchLight’s PBR passes for another relighting stage. Mask the areas that still need conventional compositing work.
Beeble’s Canvas pushes further in that direction: a node-based workspace where masks, reference images, image generators, video generators and SwitchX processing connect in a visible graph. That matters, because one-click tools tend to work beautifully right up until they don’t. And the shot that doesn’t work is invariably the one with an immovable delivery date.

A visible graph at least gives you somewhere to intervene. Swap the mask, change the reference, branch the workflow, compare versions and work out which step sent the result wandering off into generative never-never land.
There are still real pipeline limits. Beeble’s web application isn’t intended to replace a proper high-bit-depth finishing environment, and Hoon recommends exporting to Nuke or DaVinci Resolve for final compositing and color management.
SwitchHDR generates HDR material from SDR source, while the VFX Pass Generator can deliver EXR sequences containing SwitchLight’s passes. The distinction is important: HDR detail created from an SDR source is generative reconstruction, not sensor data Beeble has somehow traveled backward through time to photograph. But it addresses one of the biggest issues professionals have with generative video output: trained on pre-graded data, the output’s often leave little room for creative grades by a colorist.
Fast versus controllable
At the moment, generative AI and traditional VFX sit at opposite ends of the same problem. Generative tools are fast and difficult to control. Traditional compositing is controllable and slow.
Hoon doesn’t think filmmakers should accept that as a permanent trade-off.
“It’s a problem that they have to solve,” he said. “It has to be fast and controllable.”
That is the actual challenge facing Beeble and everyone else trying to move generative imagery into professional production. One spectacular result isn’t enough. The tool has to hold up across a sequence, preserve the actor, respect the edit and accept a precise note after the client sees version 27 and asks whether the background could be “a little more premium.”
Mostly, it has to fit the way people already work.
Getting started with Beeble
Beeble Cloud is the easiest way in and requires no local GPU. SwitchX runs in the cloud; Beeble Studio handles local PBR-pass generation and 3D relighting (not SwitchX use). At the time of writing, the published Studio minimum is Windows 10 or 11, or Rocky Linux 8 or 9, with an Nvidia Ampere-generation GPU or newer and 12GB of VRAM. Beeble recommends 24GB (think RTX 4090 or RTX 5090). Mac users are, for now, out of luck; Beeble Cloud is the only option there.
There is also a Beeble API for studios wanting to put SwitchX inside their own pipelines. That may ultimately matter more than the standalone interface. Facilities rarely want another isolated application. They want a service they can wrap in their existing asset management, review, versioning and render systems.
So, is this the future?
Maybe.
It is certainly a more credible version of AI filmmaking than replacing the entire production with a text box.
Actors remain actors. The camera still has to go somewhere. Somebody still decides where the light comes from, what the costume means and whether the architecture belongs to the correct century. The model gives you flexibility. It does not give you taste.
What changes is how permanently those decisions are baked into the photographed frame.
A basement becomes a palace. A practical flicker motivates lighting that was never there. A photographed location is extended, altered or seen from an angle nobody captured. A small crew attempts imagery that would previously have required an LED volume, substantial set construction or a much larger VFX budget.

Emerging tech like AI-based reframing of the camera may cause further disruption. But modern cinematographers’ passion for lens aberrations seems to point more toward filmmakers wanting to capture the “look” on a set and not synthesize it in post.
The hybrid approach makes the most sense when a production can photograph the elements humans still do best—performance, camera, physical interaction and specific visual reference—then use AI to extend those decisions. It makes considerably less sense when “we’ll generate it later” becomes a substitute for deciding what the shot is supposed to be.
This future won’t arrive through speed alone. The tools have to become predictable, controllable and integrated enough that artists can treat their output as material rather than magic.
For now, the hybrid set offers a fascinating preview: real actors, real performances, real filmmaking—and almost everything else up for grabs.
