AI Roundup February 2024: Sora Stuns, Gemini 1.5 Goes 1M Tokens, Gemma Opens Up
TL;DR
February was the month video models stopped being a joke. OpenAI dropped Sora and instantly reset everyone's expectations for text-to-video, then Google answered the same day with Gemini 1.5 Pro and a genuinely useful 1 million token context window. Google also open-sourced Gemma, Stability previewed Stable Diffusion 3, and Microsoft put money into Mistral while shipping Mistral Large on Azure. Short month, absurd density.
OpenAI's Sora Made Every Other Video Model Look Like a Tech Demo
On February 15, OpenAI previewed Sora, a text-to-video model capable of generating up to a minute of coherent, high-definition footage from a prompt. The sample clips (snowy Tokyo streets, a woolly mammoth herd, fake gold-rush footage) were not the usual flickering two-second loops we had grudgingly accepted. They held temporal consistency, plausible physics, and camera motion that actually tracked a scene.
Sora was not released to the public this month. It went to red teamers and a handful of artists, and OpenAI was upfront about the failure modes: physics that quietly breaks, objects that pop in and out of existence, cause and effect that drifts. But the architecture matters here. Sora is a diffusion transformer that treats video as patches of spacetime, which is the same kind of scalable recipe that made LLMs take off.
For builders, the takeaway was less go use this tomorrow and more recalibrate your roadmap. The bar for generative video just jumped, and anyone betting on the old limitations as a moat needed a new plan.
Gemini 1.5 Pro Shipped a 1 Million Token Context Window
Also on February 15 (Google clearly did not want Sora to own the day), Google announced Gemini 1.5 Pro. The headline was the context window: up to 1 million tokens in private preview, enough to stuff in roughly an hour of video, 11 hours of audio, or codebases north of 30,000 lines, all in a single prompt.
The more interesting part was that it held up. Google reported near-perfect needle-in-a-haystack retrieval across that full window, finding embedded facts about 99% of the time. It also leaned on a mixture-of-experts design to hit quality comparable to the older 1.0 Ultra while using less compute.
For anyone building retrieval pipelines, this was a real shot across the bow. A lot of RAG plumbing exists purely to work around small context windows. When the window gets this big and stays accurate, the question becomes whether you chunk and retrieve at all, or just hand the model the whole document and let it sort it out.
Google Open-Sourced Gemma, Its Lightweight Open Models
On February 21, Google released Gemma, a family of open-weight models in 2B and 7B sizes, each with pretrained and instruction-tuned variants. They were built from the same research lineage as Gemini, text-in and text-out only, and explicitly aimed at developers who want to run models on their own hardware.
This was Google planting a flag in open-weights territory that Meta's Llama had largely defined. For the local-AI and homelab crowd, a 7B that runs comfortably on a single consumer GPU, ships with permissive usage terms, and arrives with day-one support across the usual tooling is exactly the kind of thing worth pulling down and benchmarking yourself.
Stability Previewed Stable Diffusion 3
On February 22, Stability AI announced an early preview of Stable Diffusion 3, opening a waitlist rather than a full release. The model family ranged from 800M to 8B parameters, and Stability pitched big gains in multi-subject prompts, overall image quality, and the long-standing pain point of rendering legible text inside images.
Architecturally it is notable for the same reason Sora is: SD3 uses a diffusion transformer plus a technique called flow matching for more efficient training. The convergence is hard to miss. The frontier of generative media, image and video alike, is consolidating around transformer-based diffusion. For builders who care about open image models they can actually fine-tune and self-host, SD3 was the most important roadmap update of the month.
Microsoft Backed Mistral and Put Mistral Large on Azure
Late in the month, Microsoft and Mistral AI announced a partnership: Mistral's new flagship, Mistral Large, would be available first on Azure, alongside Azure supercomputing for training and a Models-as-a-Service offering. Microsoft also made a reported 15 million euro investment in the French startup.
Mistral Large itself was a capable closed model with strong reasoning, solid code and math, and native handling of English, French, German, Spanish, and Italian. The strategic read was sharper than the benchmarks: Microsoft, already deep with OpenAI, was visibly hedging by getting close to Europe's leading lab. For builders, more credible frontier-class options behind a standard API is simply good news.
Key Takeaways
- Generative video crossed a threshold: Sora made minute-long, coherent video a real research artifact, not a gimmick, and reset everyone's expectations overnight.
- Massive context is here and it actually works: Gemini 1.5 Pro's 1M-token window with near-perfect retrieval forces a rethink of how much RAG plumbing you really need.
- Open weights kept getting better: Gemma gave self-hosters another strong small-model option and pushed the open ecosystem past Llama's monopoly on the conversation.
- Transformer diffusion is the new consensus: Sora and Stable Diffusion 3 both bet on diffusion transformers, signaling where open and closed media models are converging.
- The labs are pairing off: Microsoft backing Mistral showed the big platforms hedging across multiple frontier labs, which means more credible model choices for the rest of us.