7The Batch (DeepLearning.AI)·18d ago

Grok Imagine 1.0 Sharply Cuts Costs for High-Quality Video Generation

xAI launched Grok Imagine 1.0, a text-and-image-to-video model that topped the Artificial Analysis Video Arena leaderboard in both text-to-video and image-to-video categories at launch. The model generates up to 15-second clips with audio at $4.20 per minute of output, significantly undercutting Google Veo 3.1 ($12/min) and OpenAI Sora 2 Pro ($30/min). It is integrated with the X social network, enabling direct generation and sharing, though xAI disclosed no technical details about the model's architecture. The launch highlights continued rapid cost compression in video generation, with a seven-fold price gap between Grok Imagine 1.0 and Sora 2 Pro.

Frontier Model Releases Evaluation and Benchmarking Inference Economics Multimodal Progress Artificial Analysis Grok Imagine Google Veo 3.1 X LM Arena IVEBench xAI Kling 2.5 Turbo OpenAI Sora Runway Gen-4.5 Andrew Ng

Related guides (4)

Frontier Model ReleasesTopic guide

Frontier Model Releases: The Race From Language to Action

Read asBeginner In-depth

Multimodal ProgressTopic guide

Multimodal Progress: How AI Learned to See, Hear, and Act

Read asBeginner In-depth

Inference EconomicsTopic guide

Inference Economics: The Cost Structure of Running AI Models in Production

Read asIn-depth

Evaluation and BenchmarkingTopic guide

Evaluation and Benchmarking: The Shifting Yardstick of AI Capability

Read asIn-depth

Related events (8)

5Latent Space·19d ago·source ↗

Why Video Agent Models Are Next — Ethan He, xAI Grok Imagine

Latent Space interviews Ethan He, the lead behind xAI's Grok Imagine video generation product, covering its development in roughly three months. The discussion explores the distinction between video generation models and world models, and positions video agents as a significant near-term frontier. He argues Grok Imagine is underrated relative to its capabilities.

Frontier Model Releases Agent and Tool Ecosystem Grok Imagine world model video agents +4 more

8Openai Blog·1mo ago·source ↗

Sora Video Generation Model Launches at sora.com

OpenAI has publicly launched Sora, its video generation model, available at sora.com. The model supports video generation up to 1080p resolution and 20 seconds in length, with widescreen, vertical, and square aspect ratios. Users can generate content from text prompts or bring existing assets to extend, remix, and blend.

Frontier Model Releases Enterprise Deployment Patterns sora.com OpenAI Sora +1 more

6Google Deepmind Blog·1mo ago·source ↗

Introducing Veo 3.1 and Advanced Creative Capabilities

Google DeepMind has announced Veo 3.1, an updated version of its video generation model, with significant enhancements to creative control features. The announcement comes from DeepMind's official blog, indicating a formal product update rather than a research preview. Specific capability details are not provided in the body text, but the framing suggests improvements to user-facing generation controls.

Frontier Model Releases Multimodal Progress Veo Veo 3.1 Google DeepMind

6Google Deepmind Blog·1mo ago·source ↗

Veo 3.1 Ingredients to Video: More consistency, creativity and control

Google DeepMind has released Veo 3.1, an updated video generation model that improves consistency, creativity, and control in generated clips. The update produces more natural and dynamic video content and adds support for vertical video generation. The announcement comes from DeepMind's official blog as a tier-1 source.

Frontier Model Releases Multimodal Progress Veo Veo 3.1 Google DeepMind

6The Batch·19d ago·source ↗

ByteDance Launches Seedance 2.0 Video Generation Model Globally via CapCut

ByteDance has deployed Seedance 2.0, a multimodal video generation model, to hundreds of millions of CapCut users across multiple global regions. The model supports text, image, audio, and video inputs with synchronized audio-video output, lip-synced dialogue, and camera control via prompts. It ranks within the top two on Arena AI and Artificial Analysis video leaderboards, and is available via API at $0.30 per second of output. The issue also features Andrew Ng's editorial arguing against the 'AI jobpocalypse' narrative, attributing it to incentive structures at labs and companies.

Frontier Model Releases Inference Economics Seedance 2.0 Artificial Analysis CapCut +8 more

7The Batch·19d ago·source ↗

ByteDance Deploys Seedance 2.0 Video Model to CapCut's 736M Users as OpenAI Shutters Sora

ByteDance has integrated Seedance 2.0, its multimodal video generation model, into CapCut for paying users across multiple global regions, reaching a platform with approximately 736 million monthly active users. The model supports text, image, audio, and video inputs, generates synchronized audio-video output in a single pass including multi-shot sequences, and ranks in the top two on Arena AI and Artificial Analysis video leaderboards, with Alibaba's HappyHorse-1.0 as its closest competitor. Simultaneously, OpenAI is discontinuing the Sora app and API after daily active users fell below 500,000 and operating costs reached an estimated $1 million per day. The contrast illustrates a broader market shift where Chinese developers are accelerating video model releases while U.S. consumer video products retreat.

Frontier Model Releases Evaluation and Benchmarking Seedance 2.0 Artificial Analysis CapCut +15 more

7Google Deepmind Blog·1mo ago·source ↗

Veo 2 Video Generation Launches in Gemini Advanced and Whisk Animate

Google DeepMind is rolling out Veo 2 video generation capabilities to Gemini Advanced and Whisk, enabling users to create high-resolution eight-second videos from text prompts or animate still images. Gemini Advanced subscribers can generate videos directly from text, while Whisk Animate converts input images into short animated clips. This marks a consumer-facing deployment of Veo 2, DeepMind's second-generation video generation model.

Frontier Model Releases Enterprise Deployment Patterns Gemini Advanced Veo 2 Whisk +3 more

8Google Deepmind Blog·1mo ago·source ↗

Google DeepMind Introduces Veo 3, Imagen 4, and Flow Filmmaking Tool

Google DeepMind has announced Veo 3 and Imagen 4, new generative video and image models respectively, alongside a filmmaking tool called Flow. The announcement comes from DeepMind's official blog and represents the next generation of their generative media capabilities. These releases expand Google's multimodal generative AI portfolio targeting creative and professional media production use cases.

Frontier Model Releases Agent and Tool Ecosystem Imagen 4 Veo 3.1 Flow +2 more