Google introduces Gemini Omni 1.1 Flash: greater control over video generation

Google has announced the launch of Gemini Omni 1.1 Flash, a significant update to its video generation model that offers creators and developers new tools for more precise control over generated sequences. Presented on August 27, this update aims to bridge the gap between impressive AI video demonstrations and practical tools that creators can use to refine their projects.

New features for more precise control

Gemini Omni 1.1 Flash introduces several innovations that allow better management of generated video sequences. Among the main novelties, the ability to extend existing scenes and generate specific movements between initial and final frames. This allows creating camera movements, zooms, and transitions without having to start from scratch every time.

Another relevant aspect is flexibility in production. Users can generate previews in 360p format faster and at lower costs, ideal for experimentation, while final clips can be produced in 1080p or 4K. This feature is particularly useful for teams that need to test different versions of a scene before choosing the definitive one.

Limits and technical challenges

Despite the progress, Gemini Omni 1.1 still has some limitations. Video generation is limited to segments of 3-10 seconds at a time, and the maximum duration of 40 seconds must be achieved through repeated extensions. Additionally, the model can only use 3 seconds of footage as a reference, meaning it is not yet possible to freely transform a long clip from start to finish.

Regarding resolution, Omni 1.1 can generate videos natively at 360p and 720p, while 1080p and 4K are only available through upscaling. This indicates that, although Google has improved the consistency and control of sequences, generating long-form videos in a single pass remains a challenge.

Availability and pricing

Google is making Gemini Omni 1.1 Flash available through various platforms, including Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform. The new features are also integrated into Google Flow.

For API users, Google adopts a token-based pricing system, with costs varying based on resolution:

  • 360p videos: $0.03 per second

The 360p option is designed for experimentation, allowing creators to test prompts, camera movements, and transitions before moving on to higher-quality versions. This approach could make repeated iteration significantly less costly, especially for teams that need to test different versions of a scene.

Implications for creators

For those experienced with AI video generation, the biggest challenge is not just creating visually impressive clips but achieving the desired movement and maintaining scene consistency. Gemini Omni 1.1 could be particularly useful in this regard, offering greater control over where a shot starts and ends and allowing the extension of already appreciated scenes without having to regenerate them completely.

For example, a marketing team creating a product video might want to extend an initial scene while maintaining consistency in characters and visual details. A social media content creator might want to specify a camera movement between two shots. Even those working on storyboarding could generate different low-resolution versions before choosing the definitive one.

Integration with other Google tools

These new features fit into Google's broader context of making AI video creation more accessible. The company is expanding Google Vids with AI-generated presenters, music tools, and other features to reduce the manual production work required.

Despite current limitations, Gemini Omni 1.1 represents a significant step toward integrating AI video tools into creators' daily workflows. The ability to maintain and control the parts of a video that work well is a substantial improvement over the impressive but isolated demonstrations we have seen so far.

For further insights into Google's tools for building and testing generative AI applications, you can consult the TechRepublic guide to Google AI Studio.

The future of AI-assisted video production

Gemini Omni 1.1 Flash represents a significant step forward in AI-assisted video production, but its true potential will manifest within a broader ecosystem of tools. While the model is currently optimized for creating clips up to 40 seconds long, integration with other Google technologies could revolutionize creative workflows.

One of the most interesting aspects is the ability to maintain visual consistency across different generations. When a video is extended in 10-second segments, the model uses up to 10 seconds of previous footage as context for the next generation. This approach allows preserving crucial details such as lighting, object positions, and character expressions, solving one of the main problems of current AI video generators.

Comparison with other market solutions

In the competitive landscape of AI video generation tools, Gemini Omni 1.1 positions itself as a versatile solution for those needing precise control over sequences. Unlike some competitors that focus exclusively on the visual quality of individual clips, Google is developing an integrated ecosystem that combines video generation, virtual presenters, and music tools.

For example, while platforms like Runway ML offer powerful AI-based video editing tools, Gemini Omni 1.1 stands out for its ability to maintain temporal consistency across different segments of a video. This feature is particularly useful for producing advertising content, tutorials, and documentaries where narrative continuity is fundamental.

Technical challenges and current limitations

Despite the progress, generating long videos remains a complex challenge. The limitation of 3-10 seconds per generation and the need for repeated extensions introduce potential points of discontinuity. Additionally, the 1080p and 4K quality obtained through upscaling may not be sufficient for professional projects requiring maximum visual detail.

Another challenge is represented by the management of complex transitions between different scenes.

Ethical and legal considerations

With the increasing capability of AI video generation, important ethical and legal questions arise. The ability to create realistic content could raise concerns about disinformation and the improper use of technology.

Google will need to address these challenges by implementing measures for authenticating and tracking generated content. Additionally, it will be necessary to develop clear guidelines for the ethical use of these technologies, especially in sensitive sectors such as politics and journalism.

Future perspectives

While Gemini Omni 1.1 represents an important step forward, the future of AI video generation is still evolving. Possible developments include:

  • Integration with motion capture technologies for greater precision in movements
  • Expansion of generation capabilities to higher resolutions
  • Development of algorithms capable of managing complex transitions between different scenes
  • Integration with traditional video editing platforms for a hybrid workflow

For those interested in exploring the potential of Google AI Studio further, you can consult the TechRepublic guide that provides a comprehensive overview of the available tools.

Editorial Note and Disclaimer

The guides and content published on GoYou are the result of independent research and analysis activities, for informational, educational, and in-depth purposes.

GoYou does not constitute a journalistic publication or an editorial product pursuant to Law No. 62/2001 and does not perform real-time information activities.

The GoYou project does not provide professional, technical, legal, or financial advice and disclaims any responsibility for the improper use of the information published.

In the Crypto sector, every investment involves risks: the reader is invited to always inform themselves autonomously before making any decision.