Premiere is turning the timeline into a place where media can be created
Adobe released Premiere 26.5 in September 2026 with a workflow change that moves generative AI closer to the edit itself.
The new Generative Media Tool lets an editor define a range directly on the timeline, describe the video or sound needed, choose an AI model and generate the result without leaving the sequence.
The generated media arrives as an editable clip inside the project, so it can be trimmed, moved, layered and refined like other timeline media.
That changes the role of the timeline. It is no longer only the place where existing footage is assembled. It can also become the place where new material is created for a specific moment in the edit.
Generate Video can use the surrounding edit as visual context
The video workflow is designed around the sequence that is already being edited.
Adobe says editors can generate video to fill a selected range in the timeline and use reference frames from the project to guide the result. A first frame can define the starting visual, while first and last frames can guide a generated transition.
Editors can also add multiple reference frames to give the model more visual context.
The generation controls extend beyond the prompt. Depending on the selected model, Premiere can expose settings such as resolution, aspect ratio, frame rate, seed and duration.
That makes the feature feel closer to an editing tool than a separate prompt box. The editor starts from a specific point in the sequence, defines what belongs there and keeps working from the generated clip.
Premiere can switch between Adobe and partner AI models
Adobe is also making model choice part of the workflow.
The company says the Generative Media Tool can work with Adobe Firefly as well as partner models including Google Veo, Kling, Runway and Luma.
The editor chooses the model from inside Premiere, then configures the generation settings supported by that model.
This creates a different relationship between creative software and generative AI. Instead of opening a separate service for every model, the editing application becomes the layer that organizes where the generated result belongs.
The timeline stays constant while the model can change underneath it.
Sound effects can be generated at the exact moment they are needed
The same tool also works with audio.
Generate Sound Effects lets an editor create a sound directly for a selected part of the sequence. Adobe says its audio model is commercially safe, and the tool can use text descriptions to create the requested effect.
Editors can also use their voice to guide qualities such as rhythm, timing and intensity.
That is useful because sound design is closely tied to timing. A sound effect is not only about what it sounds like; it has to land at the right moment against the picture.
Putting sound generation inside the timeline keeps the visual timing visible while the audio is being created.
Soundscape and Music add broader generative audio layers
Adobe is extending the same idea beyond individual effects.
Generate Soundscape, currently labeled beta by Adobe, can analyze up to 15 seconds of video and create a context-informed audio starting point with ambience and sound-effect layers timed to the edit.
Generate Music, also labeled beta, can create original music with multiple variations and controls for tempo and looping. Adobe says the generated music is commercially safe and includes synchronization rights across media types.
Together, the tools create several levels of generated audio inside Premiere: one sound effect, a broader soundscape, or a music layer.
The common idea is that the edit itself provides the context for what gets created.
Paper Edit can build a rough cut from transcript selections
Premiere 26.5 also adds a different kind of AI-assisted workflow through Paper Edit.
For dialogue-heavy footage, an editor can work from the transcript instead of assembling the first sequence directly on the timeline.
Open Paper Edit in the Text panel, select the sentences or parts of sentences that should remain, review those selections and create a new sequence from them.
Premiere then assembles the selected dialogue into a rough cut.
This creates a useful split between editorial decisions and timeline refinement. The first question can be, “Which parts of the conversation belong in the story?” Once that structure is created, the editor can return to the timeline for pacing, visuals, sound and finishing.
The release also pushes more audio cleanup into the editing workspace
Adobe is pairing generative creation with more AI-assisted audio tools.
The September update introduces a new Enhance Audio workflow, Separate Crosstalk and Dynamic Auto Ducking in public beta.
Enhance Audio can work with dialogue, music, ambience and sound effects inside a clip. Separate Crosstalk can split overlapping speakers into independent tracks for individual mixing. Dynamic Auto Ducking can automatically balance dialogue, music and other elements.
These tools sit on a different side of the workflow from generation, but they point in the same direction: more of the audio process is becoming available directly inside Premiere.
The editing application is becoming the orchestration layer
The most interesting part of the update is not one individual AI feature.
Premiere is starting to act as the orchestration layer around several kinds of media intelligence.
A transcript can become a sequence. A selected timeline range can become generated video. A reference frame can guide the visual direction. A sound effect can be created against the exact timing of the picture. A soundscape can be generated from the scene itself. Different AI models can be selected without changing the editing environment.
The creative decision remains attached to the project.
That is an important shift. Generative AI becomes less of a destination and more of a capability the editor can call when the timeline needs it.

The Upgrade Feeling
Premiere 26.5 shows what happens when generative AI stops living beside the creative tool and starts becoming part of the tool.
The editor does not have to begin every AI task from a blank prompt. The timeline already contains context: timing, reference frames, dialogue, picture and the shape of the story.
Adobe is using that context as the starting point.
Generate Video fills a visual need inside the sequence. Generate Sound Effects works against the timing of the edit. Paper Edit turns transcript decisions into a rough cut. Soundscape and Music build broader audio layers. Model selection happens inside the same workspace.
The upgrade is not simply that Premiere has more AI.
It is that the editing timeline is becoming a workspace where existing media, generated media and editorial decisions can meet in one place.
Your Camera Doesn’t Have to Follow the Subject Anymore — AI Can Reframe the Video After You Shoot It
For most of the history of video, framing happened while you were recording.
If a person moved left, you moved left. If they crossed the frame, you followed them. If you lost them behind someone else, the shot changed with them.
Samsung’s My FanCam on the Galaxy Z Fold8 series moves part of that job into editing.
The feature lets you open an existing video, choose a person and have the phone automatically track that person through the clip. It then reframes the footage around the selected subject and lets you choose an aspect ratio for the finished result.
Samsung explained the engineering behind My FanCam on September 1, 2026. The interesting part is not that AI can crop a video.
It is that the phone is separating two decisions that used to happen at the same time: capture the whole scene first, then decide who the camera should have been following.
That matters because a wide recording can preserve more than one possible story. A concert clip can become focused on one performer. A school event can be reframed around one child. A group video can be turned into several different subject-centered edits without recording the moment again.
The physical camera still points where you pointed it. The final moving frame no longer has to be decided there and then.
My FanCam Starts After the Video Already Exists
My FanCam is not primarily an automatic camera operator working during capture.
Samsung describes the workflow inside Gallery.
You select a video that already exists, launch My FanCam and choose the person you want to follow. The system analyzes the footage, tracks the selected person across frames and creates a reframed result centered on that subject.
That distinction matters.
The original recording can remain wide enough to contain several people. The decision about who becomes the focus is made later.
Samsung also says the feature works with older footage, including videos that were not recorded on a Galaxy device. That means My FanCam is not dependent on special tracking metadata captured only by the Fold8 camera. It can analyze ordinary video as input.
This makes the Fold8 behave as both a camera and an AI editing workstation. The source file is treated as visual evidence that can be reinterpreted after the event, rather than as a composition that became permanently fixed the moment the user stopped recording.
For users, the practical idea is simple: the raw footage preserves the scene, while the AI helps choose the final composition afterward.
The Camera Can Stop Chasing the Subject
A traditional fan-cam style video asks the person holding the phone to do several things at once.
Watch the event. Find the subject. Keep that person in frame. Adjust the phone when they move. Avoid sudden pans. Choose a crop that works for the final platform.
My FanCam moves several of those editing decisions to the phone after capture.
Samsung’s own advice is to shoot at high resolution and leave some extra space around the subject. That wider recording gives the software room to follow movement later.
This is an important inversion.
Instead of trying to create the final composition at the moment of capture, the user can preserve more of the scene and let the final framing happen afterward.
That can also change how a person behaves while recording. Rather than staring at the screen and constantly correcting the frame, they can leave more breathing room and concentrate on keeping the important action somewhere inside the source image.
The camera still records the pixels. The AI decides which part of those pixels should become the moving frame. It is less like replacing the camera operator and more like giving the editor a virtual camera that can move inside the footage later.
Tracking a Person Is More Than Tracking a Position
The central technical problem is identity.
A simple tracker could follow coordinates: the person is here in one frame, then slightly farther right in the next.
Real footage is harder. People cross in front of each other. They turn around. They move quickly. Several people can wear similar colors. A selected subject can disappear behind someone else and return later.
Samsung Research says My FanCam uses an AI model that recognizes both a person’s location and visual characteristics. The model is designed to remember the selected individual across the video so it can continue identifying the same person while the scene changes.
That means the system is solving two connected problems: Where is the person now? Is this still the same person?
The second question is what makes long-range tracking useful for reframing. If the tracker only followed the nearest body-shaped region, it could easily jump to someone else when two people crossed paths.
Identity-aware tracking is therefore the bridge between raw detection and an edit that feels intentional. The system needs continuity, not just repeated detections.
Occlusion Is Where Real Video Becomes Difficult
Tracking looks easy when one person walks across an empty background.
Concerts, school performances, sports days and group videos are not like that.
Bodies overlap. The subject may move behind someone else. They may leave the visible area. They may not even appear until halfway through the recording.
Samsung says these real-world situations were a major part of development. The teams found that models that behaved well in research conditions could respond differently when tested on footage with fast movement and overlapping people.
That led to repeated testing between Samsung Research and the mobile development team.
My FanCam also lets the user move through the video and select the subject at a moment when that person is clearly visible. So the system does not force the user to identify someone in the opening frame. The useful frame can be anywhere in the clip.
That small interface decision is technically meaningful. It lets the human provide a cleaner identity example when the beginning of the video is visually ambiguous. Instead of pretending the AI can infer everything from a difficult first frame, the workflow gives the user a practical way to help the tracker start from better evidence.
The Phone Analyzes More Than One Person at a Time
There is another design decision hidden behind the editing interface.
A group video may contain several people the user wants to try. One implementation could analyze the clip again every time the user selects someone new.
Samsung chose a different approach.
After the initial analysis, My FanCam can let the user switch between detected people without repeating the full wait each time. Samsung says the feature analyzes multiple people together so that changing the selected subject can show a result immediately after the first analysis completes.
This is a user-experience decision built on computation.
The system spends time understanding the clip once, then reuses that analysis for multiple editing choices. That matters in a real Gallery workflow because experimentation is part of editing. A user may not know which person produces the best result until they preview several options.
By doing more of the expensive analysis upfront, the feature turns later subject changes into interactive choices rather than repeated batch jobs.
The AI model becomes part of the editing timeline rather than a one-shot effect applied to a single person.
Samsung’s Official Camera Video Shows the Broader Editing Direction
Samsung published an official Galaxy Z Series video in July titled “AI-Powered Camera: Effortless Shooting and Editing.”
The video predates the September 1 engineering interview and is not a recording of that interview.
It is useful here because it shows the broader camera-and-editing direction Samsung introduced with the 2026 Galaxy Z lineup, including AI-assisted ways to turn captured footage into a more deliberate final result.
My FanCam fits directly into that direction.
The phone is not only trying to improve the image at the moment the shutter is pressed. It is becoming an editing system that can reinterpret what was already captured.
That distinction is important for the article because the September interview explains the engineering, while the official video provides visual product context. Together they show the same shift from two sides: one explains how the tracking pipeline was built, and the other shows how Samsung is packaging AI-assisted camera editing as part of the wider Galaxy Z experience.
All of the Person Tracking Runs On-Device
Samsung says the My FanCam AI workload runs entirely on the device without requiring a separate server for the tracking process.
That creates a very different engineering constraint from running a large video model in a data center.
Video is computationally expensive. A longer clip contains more frames. Each frame has to be decoded, analyzed and connected to what the model already knows about the people in previous frames.
Doing that locally also means working inside the phone’s limits for processing speed, power consumption and heat.
Samsung Research says it optimized the AI model to reduce the required computation. The mobile team also optimized the surrounding pipeline, including video decoding, data transfer and analysis.
On-device processing also changes the interaction model. The feature can work as a local editing function instead of requiring the user to send every source video to a remote processing service before seeing a result.
The model is only one part of the feature. The complete path from compressed video file to tracking result has to fit inside a phone and still feel reasonable enough for ordinary Gallery editing.
The Video Pipeline Matters as Much as the AI Model
AI features are often described as if the model receives perfect input instantly.
A real mobile video workflow has more stages.
The file has to be opened. Compressed frames have to be decoded. Image data has to move through memory. The tracking model has to process the frames. The selected crop has to be previewed. The final video has to be rendered.
Samsung specifically points to optimization across video decoding, data transfer and analysis because every stage can add compute, latency and energy use.
This is why My FanCam is an interesting mobile-AI example.
The user sees one button and one subject selection. Behind that interaction is a pipeline that has to keep the AI workload practical enough for a consumer phone.
A fast model can still feel slow if decoding or memory movement becomes the bottleneck. Likewise, efficient video handling cannot help if the tracker itself is too computationally heavy.
The feature only works as a product when all of those components are engineered together. My FanCam is not just a tracking model. It is a tracking model integrated into a complete video-processing system.

Reframing Is a Crop That Moves Through Time
A still-image crop chooses one rectangle.
Video reframing chooses a rectangle again and again as time moves forward.
If the subject walks across the original frame, the crop has to move with them. If the output is vertical, the available horizontal space changes. If the output is wide, the system has more room around the person.
My FanCam lets the user adjust the aspect ratio after the subject has been selected.
That means the same original video can support more than one final composition. The AI tracking provides the subject path. The aspect ratio defines the shape of the output window. The editing system combines the two.
This makes reframing a temporal problem rather than a single crop decision. The system has to preserve a usable composition across a sequence of changing positions instead of simply centering one frame.
It is especially useful for footage that may later be shared in different formats, because the capture does not have to commit to one social-video shape at the moment it is recorded. One wide source can become different subject-centered versions later.
Old Footage Becomes New Input
One of the most useful details in Samsung’s September interview is that My FanCam works with existing videos.
Samsung’s developers mention old concert clips, trips, children’s activities and other previously recorded footage. They also say the source video does not have to come from a Galaxy device.
That widens the feature beyond the camera hardware that ships with the Fold8 series.
The phone can act as an AI editing device for footage that already exists elsewhere.
This is a different kind of upgrade from improving the next photo you take. It can change how you use videos you already own.
A wide group recording from years ago can potentially become a subject-focused edit today, provided the footage contains enough usable visual information for the tracker.
That gives AI editing a backward-looking value. New hardware is often sold around the quality of future captures. My FanCam can also create a reason to revisit an archive, because the source footage becomes raw material for an editing capability that did not exist when the clip was recorded.
Wider Capture Gives the AI More Room to Work
Post-capture reframing still depends on what was captured in the first place.
Samsung recommends recording at a high resolution and framing a little wider when possible.
The reason is geometric.
If the subject reaches the edge of the original frame, there is no image outside that boundary for the software to recover. If the original video is already tightly cropped, a moving digital crop has less room to follow the person.
A wider source frame preserves more spatial margin. Higher resolution also gives the final crop more pixels to work with.
That creates a practical tradeoff. Shooting wider may look less finished as raw footage, but it can preserve more editing freedom. Shooting tighter can produce a stronger composition immediately, but it leaves less room for a virtual camera to move later.
This does not mean every wide video will produce the same result, and Samsung does not publish a universal quality threshold for all footage.
AI can move the frame inside the captured image. It cannot invent an unlimited camera view beyond it.
Capture and Composition Are Becoming Separate Stages
This is the larger camera shift.
Smartphone photography has already separated capture from many traditional camera decisions.
HDR can combine exposures after the shutter press. Portrait modes can estimate depth and change background rendering. Computational zoom can combine sensor data and processing.
My FanCam applies a similar separation to moving composition.
The original camera records a scene. The final virtual camera can be decided later.
That does not make physical camera movement irrelevant. Optical perspective, motion blur, exposure, focus and what actually enters the frame are still determined during capture.
But composition is becoming less final.
The recorded frame can become a larger canvas from which a second, moving frame is generated during editing. In that sense, mobile video is borrowing a workflow that is familiar in high-resolution production: capture more image area than the final output needs, then use the extra pixels to create movement and alternative framing in post.
What changes on the phone is accessibility. The tracking and crop path can be generated automatically instead of being built manually frame by frame.
This Is Not the Same as Generating Missing Video
It is useful to keep My FanCam separate from generative-video systems.
Samsung describes the feature as tracking and reframing a selected person in existing footage.
The core idea is not to generate a new performance or synthesize a person who was not present. The system is following visual information that already exists in the source video and selecting a moving region around it.
That distinction matters because “AI video” now covers several very different technologies.
One system may generate frames from text. Another may remove an object. Another may interpolate motion. Another may estimate depth or reconstruct detail.
My FanCam is primarily a person-tracking and reframing workflow.
Its value comes from changing the edit around existing footage rather than creating an entirely new scene.
Keeping those categories separate also makes the technical achievement easier to understand. The challenge here is persistent identity tracking, mobile computation and a usable moving crop. It is not open-ended scene synthesis.
What Samsung Has Confirmed — and What It Has Not
Samsung has confirmed several important details.
My FanCam can automatically track a selected person through a video. The underlying AI model uses location and visual characteristics to recognize individuals. Multiple people can be analyzed during the initial pass so the user can switch subjects afterward. The workload is processed on-device. The feature can work with existing videos, including footage not captured on a Galaxy phone. Users can adjust the output aspect ratio.
Samsung has also described the engineering challenges around occlusion, fast movement, processing speed, power and heat.
There are limits to what we should infer.
Samsung has not published a universal tracking-accuracy percentage for every scene. It has not said every person in every video can always be recovered after being fully obscured. It has not claimed that post-capture reframing replaces careful shooting. And it has not described the system as generating unseen parts of the scene outside the source frame.
The confirmed feature is narrower and more useful: the final subject can be chosen after the scene has already been recorded, with the phone building the crop path around that person.
The Bigger Upgrade Is a Camera You Can Aim After Recording
The most interesting thing about My FanCam is not the name or the fan-cam use case.
It is the timing of the creative decision.
A traditional camera asks you to decide where to point before the moment passes. Samsung’s feature lets the original recording hold a wider version of the moment, then uses on-device AI to decide which person the final frame should follow.
The camera still matters. The source resolution still matters. Where you stand still matters. What enters the original frame still matters.
But one part of directing the shot has moved into editing.
That gives smartphone video a new workflow: record the scene, select the subject, let the phone reconstruct the framing path, then choose the output shape.
It also suggests a broader direction for computational cameras. AI does not have to change what the sensor captured to change how the moment is presented. Sometimes the useful upgrade is simply being able to make a better decision later.
The camera does not literally move after recording.
The frame can.
And for everyday video, that may be the more useful change.
Generative Video Is Starting to Behave More Like an Editing Workflow
Generative video began with a simple interaction.
Write a prompt.
Generate a clip.
Choose whether to keep it.
That model is still useful, but it is no longer the whole workflow.
Gemini Omni 1.1 Flash adds controls that move generation closer to the structure of video editing.
A creator can define the first frame.
Define the last frame.
Generate the movement between them.
Extend an existing scene.
Reuse part of an earlier video as reference context.
Create a low-resolution draft.
Then render or upscale the chosen result at a higher resolution.
Those steps look less like one isolated generation and more like a sequence of editorial decisions.
Google released Omni 1.1 Flash on August 27, 2026 and made the new capabilities available through the Gemini API, Google AI Studio, Gemini Enterprise Agent Platform and Google Flow.
The model still generates video.
The difference is how much of the generation process can now be shaped around a planned sequence.
The creator is beginning to specify not only what should appear.
They can also specify where the shot begins, where it ends and how the sequence continues.
First and Last Frames Create a Keyframe-Like Control Point
One of the clearest additions is first-and-last-frame interpolation.
Gemini Omni 1.1 Flash can take one image as the starting frame and another image as the final frame.
The model then generates continuous video between them.
That changes the creative problem.
Instead of asking the model to invent both the start and the end of a shot, the creator can define those endpoints directly.
The generated part becomes the transition.
Google gives examples including camera orbits, zoom transitions and seamless loops.
The API documentation exposes the same structure with media-role tags such as FIRST_FRAME and LAST_FRAME.
That makes the control explicit.
The first image is not only a reference.
It is the opening state of the shot.
The last image is not only style guidance.
It is the target state the generated sequence should reach.
This resembles a basic editing or animation idea: establish key points in time, then create the movement between them.
The model is still doing the generation.
The creator has more control over the temporal boundaries of the shot.
A Camera Move Can Now Be Defined by Its Destination
Camera direction becomes more specific when the destination is known.
A text prompt such as “zoom in” describes motion.
A first-and-last-frame pair can describe the visual result that motion should connect.
That gives the model more information about the intended path.
Google demonstrates this with continuous camera moves such as whip pans, orbits and zooms.
The creator can choose the opening composition.
Choose the final composition.
Then describe how the camera should travel between them.
This is different from asking for a complete shot from text alone.
The final composition is already fixed by the input.
The generation problem is narrowed to continuity, motion and transition.
For production workflows, that can be useful when the next shot needs to land on a particular framing.
A video can begin on a wide composition and end on a close-up.
It can start outside a room and end inside.
It can begin with one subject position and end with another.
The model is not only deciding what a camera move might look like.
It is generating toward a defined visual destination.
The Same Frame Can Define Both Ends of a Seamless Loop
The first-and-last-frame system also creates a direct looping workflow.
Google’s developer documentation shows that the same image can be used as both the first frame and the last frame.
The generated motion then begins from that image and returns to it.
That gives creators a structured way to build looping clips.
A product can rotate and return to its initial orientation.
A scene can move through an environmental animation and come back to the same composition.
A background can cycle while preserving a continuous boundary between the end and the beginning.
Looping video has traditionally required careful alignment between the last frame and the first frame.
Here, the endpoint is defined before generation.
The model receives the loop condition as part of the input structure.
That is another example of generative video becoming more timeline-aware.
The creator is not asking only for content.
They are defining a temporal relationship between the beginning and the end of the clip.
Scene Extension Turns One Clip Into the Beginning of a Longer Sequence
Omni 1.1 Flash also adds scene extension.
The model can take an existing video and continue generating from where it ends.
Google’s current documentation says extensions are generated in 10-second increments and can continue to a cumulative length of up to 40 seconds.
The model uses the preceding video as context for the continuation.
That changes how longer sequences can be built.
A creator can generate an initial scene.
Review it.
Then decide what should happen next.
The next prompt can continue the same location, introduce another action, change the camera movement or move the story into another connected scene.
The previous clip becomes input to the next generation step.
This creates a sequential workflow.
Generate.
Inspect.
Extend.
Inspect again.
Continue.
The model is no longer limited to producing one independent short clip at a time.
The previous shot can become the context for the next part of the sequence.
Ten Seconds of Prior Context Helps the Extension See More of the Scene
The amount of prior video available to the model matters when a scene continues.
Google says Omni 1.1 Flash can analyze up to 10 seconds of previous video context when generating an extension.
The company contrasts that with earlier models that referenced only the final second.
A larger context window gives the extension access to more of what happened before the cut point.
That can include movement direction.
Character position.
Camera motion.
The visual layout of the environment.
Objects introduced earlier in the shot.
Audio or dialogue already present in the sequence.
The extension can therefore be conditioned on a longer slice of the previous scene.
This is a timeline concept again.
The next generated segment is not built from one frozen endpoint alone.
It can look back across several seconds of preceding motion before deciding how the continuation should unfold.
For longer-form generation, that creates a stronger link between one segment and the next.
The clip becomes history for the following generation step.
Extension Prompts Can Direct Both Continuity and Change
Scene extension is not limited to “continue exactly as before.”
The prompt can describe what should remain and what should change.
Google’s documentation gives examples of continuing the same characters, changing the music, introducing a scene cut or directing a new camera move.
That gives the creator a way to edit the future of the clip.
The existing video defines the past.
The prompt defines the next event.
The model generates the transition between those two states.
This structure fits narrative work especially well.
A conversation can continue.
A camera can pull back to reveal more of the location.
A character can move into a new part of the scene.
The audio can continue or change.
The next extension can then build on that result.
The process becomes iterative rather than one-shot.
The creator can make decisions at each stage of the sequence instead of having to specify the entire finished video in one prompt.
360p Drafts Add a Preview Stage Before the Final Render
Another production-oriented change is the 360p draft mode.
Google says Omni 1.1 Flash can generate 360p previews up to 60 percent faster than its standard 720p output based on system throughput.
Google also says the 360p option costs one third of the standard 720p generation price in the current API pricing structure.
The purpose is iteration.
A creator may need several versions before choosing a shot.

One version changes the camera path.
Another changes timing.
Another changes the subject position.
Another changes the ending frame.
Generating every experiment at final resolution uses more time and compute than the creative decision requires.
A draft stage separates composition from finishing.
Generate a lightweight preview.
Compare alternatives.
Select the version that works.
Then move the selected shot into the higher-resolution stage.
That is a familiar pattern in editing, visual effects and 3D production.
The early decision does not require the final render.
Omni 1.1 Flash now gives generative video the same kind of two-stage workflow.
Drafting Makes Variation Testing More Structured
Low-resolution previews become more useful when they are treated as controlled experiments.
Google describes a “Draft Room” concept where creators can generate several 360p versions and change one variable at a time.
That might be camera direction.
Timing.
Lighting.
Character position.
The final frame.
The extension prompt.
The creator can then compare the variations side by side.
This changes prompting from repeated guessing into a more structured process.
Keep most of the shot fixed.
Change one parameter.
Observe the result.
Then decide which direction to continue.
That resembles ordinary creative iteration.
An editor tests several cuts.
A designer compares versions.
A photographer changes one exposure variable.
A 3D artist renders a low-quality preview before committing to final settings.
Generative video can now support a similar loop.
The model remains probabilistic, but the workflow around it can become more systematic.
The draft is not the final asset.
It is evidence for the next creative decision.
1080p and 4K Create a Separate Finishing Stage
Once the shot has been selected, Omni 1.1 Flash supports higher-resolution output.
Google lists 1080p and 4K as available output resolutions for the current model.
That creates a clear separation between preview and final delivery.
The 360p generation can be used for fast iteration.
The selected result can then move to a higher-resolution output for production use.
This matters because resolution becomes a workflow variable rather than something fixed at the beginning.
The creator can spend compute where it matters.
Early experimentation can stay lightweight.
The chosen clip can receive the higher-resolution treatment.
Google Flow uses the same general idea in its creative interface, allowing users to draft at lower resolution and move selected work into higher-quality output.
The result is a more familiar production sequence.
Plan.
Preview.
Choose.
Finish.
The generation model is becoming one stage inside that sequence rather than the entire sequence by itself.
Video References Add Motion and Character Context to the Prompt
Omni 1.1 Flash can also accept short video references.
Google says the model can use up to three seconds of reference video when building a new scene.
That gives the prompt another type of context.
An image reference can show how a subject looks.
A video reference can also show movement.
A dance.
A gesture.
A camera behavior.
A character performance.
The model can then use that reference while generating a different scene.
Google demonstrates a workflow where several characters are assigned movements derived from reference videos.
The reference is not necessarily the video being edited.
It can be guidance.
That expands the role of input media.
Text describes intention.
Images can define appearance or keyframes.
Video can provide motion reference.
An existing generated clip can provide extension context.
The creator can combine those media types to define different parts of the shot.
The prompt becomes a composition of roles rather than one block of text.
Media Roles Make Multimodal Prompting More Explicit
Google’s API documentation now gives uploaded media explicit roles.
FIRST_FRAME.
LAST_FRAME.
IMAGE_REF.
VIDEO_REF.
PREVIOUS_VIDEO.
Those labels matter because the same media file can be used in different ways.
An image used as FIRST_FRAME is the literal opening frame.
The same image used as IMAGE_REF provides guidance without becoming the opening frame.
A video used as PREVIOUS_VIDEO becomes the sequence being extended.
A video used as VIDEO_REF becomes reference context.
That distinction makes multimodal prompting more like a production specification.
The creator does not only upload assets.
They define what job each asset performs.
This is similar to the way an editing project separates source media, timeline clips, reference footage and output targets.
The model receives both the media and the role.
That reduces ambiguity at the interface level.
The creative system knows which asset is supposed to be a boundary, which is supposed to be reference material and which is supposed to be the existing sequence being continued.
Conversational Editing Keeps the Same Video Inside an Ongoing Interaction
Gemini Omni was designed around conversational video generation and editing.
The Interactions API allows a developer to keep working with the same creative sequence across turns.
A user can generate a clip.
Then ask for an edit.
Then extend it.
Then change another part of the scene.
The interaction history becomes part of the workflow.
This is different from exporting a clip and starting a completely new generation every time.
The model can remain inside an ongoing creative session.
Google introduced Omni around the idea of combining text, images, audio and video as input and refining video through conversation.
Omni 1.1 Flash adds more structured controls to that conversation.
The prompt still matters.
Now the creator can combine the prompt with keyframes, reference media, previous video and resolution choices.
Conversation becomes the control layer around those assets.
The editing session is therefore both natural-language driven and media-structured.
Google Flow Turns the New Controls Into a Creator Interface
The same features are appearing in Google Flow.
Google announced the Omni 1.1 Flash controls in Flow on August 27, 2026.
Creators can define start and end frames.
Generate 360p drafts.
Move selected work into higher-resolution output.
Continue shaping video inside the Flow environment.
This matters because the capabilities are not limited to API developers.
The same model concepts are being translated into a visual creative tool.
A creator can think in terms of shots and transitions instead of request objects and API fields.
The developer sees FIRST_FRAME and LAST_FRAME.
The creator sees start frame and end frame.
The underlying idea is the same.
The model is receiving timeline boundaries.
This is how a generative capability becomes part of an editing workflow.
The technical control appears first as model functionality.
The creative application turns that control into an interface a person can use repeatedly.
The Model Can Sit Inside Existing Creative Tools
Google is also positioning Omni Flash as a model that can be integrated into other creative software.
Google’s August announcement cites Adobe Firefly, Figma Weave and Runway among tools using or integrating Gemini Omni Flash.
That matters for workflow design because the generative model does not need to become the entire editing application.
It can become one capability inside a larger creative environment.
The host application can manage projects.
Assets.
Versions.
Canvas organization.
Collaboration.
Export.
Omni can provide generation and editing operations inside that system.
This follows the same architectural pattern seen across AI software.
The model supplies a capability.
The application supplies the workflow around it.
For video production, that means generative controls can live beside conventional editing, design and asset-management tools rather than requiring a separate creative process.
Generative Video Is Becoming a Sequence of Controlled Decisions
The important change in Omni 1.1 Flash is not one isolated feature.
It is the combination.
First frame.
Last frame.
Scene extension.
Ten seconds of prior video context.
Reference video.
Media roles.
360p previews.
1080p and 4K output.
Conversational editing.
Each feature controls a different part of the process.
Where the shot starts.
Where it ends.
What came before it.
What movement or character should be referenced.
How quickly a draft should be generated.
When the output should move to finishing resolution.
Together, those controls create a sequence of decisions around the generated video.
That is what makes the workflow feel closer to editing.
The creator does not surrender the entire timeline to one prompt.
The model generates inside boundaries the creator can define at several stages.
Generative video remains generative.
The workflow around it is becoming more deliberate.
That is the upgrade.