On August 27, 2026 (local time), Google released "Gemini Omni 1.1 Flash," a new multimodal model capable of video generation and editing. In addition to text-to-video generation, this model supports generation based on specified start and end frames provided as images, as well as generation using reference videos of up to 3 seconds.
Furthermore, it features a function that analyzes existing videos of up to 10 seconds to generate a continuation of the video while maintaining character identity, lighting, and context. To improve workflow efficiency, the model allows for stepped output, ranging from low-cost, high-speed preview generation at 360p to upscaling up to 4K resolution.
In tests conducted by the third-party organization Arena, the model ranked first in the text-to-video generation category. The model is available via Google AI Studio and API. API pricing is set based on the number of seconds generated, at $0.03 per second for 360p and $0.30 per second for 4K.
Source: Googleが動画生成AI「Gemini Omni 1.1 Flash」をリリース、「最大4Kアップスケール」「動画の続きを生成」「低解像度で高速プレビュー」など - GIGAZINE (Google News: Gemini)