← Back to all posts
News

AI Roundup December 2024: OpenAI o3, Gemini 2.0 Flash, and Sora Goes Public

December 31, 2024 · News
AI Roundup December 2024: OpenAI o3, Gemini 2.0 Flash, and Sora Goes Public

TL;DR

December was the loudest month of the year. OpenAI ran a 12-day shipping marathon that ended with o3, a reasoning model that smashed the ARC-AGI benchmark and reignited the AGI argument. Google answered with Gemini 2.0 Flash plus the genuinely impressive Veo 2 video model, Sora finally opened to the public, and Meta quietly shipped Llama 3.3 70B, a 70B open-weight model punching at 405B level. If you build with this stuff, your stack just changed.


OpenAI Announces o3 and Posts an 87.5% on ARC-AGI

On December 20, OpenAI capped its 12 days of shipping by announcing o3 (skipping o2 to dodge a trademark clash with the carrier). It is not a normal model bump. On the ARC-AGI benchmark, the test built specifically to resist memorization and reward genuine novel reasoning, o3 hit 75.7% at the standard compute limit and 87.5% in a high-compute configuration. For reference, GPT-4o scored about 5% on the same test earlier in the year. It also posted 96.7% on the 2024 AIME and 87.7% on GPQA Diamond.

The honest take: this is a real capability jump, and also the high-compute run reportedly cost a fortune per task, so nobody is running 87.5% o3 in production tomorrow. ARC Prize themselves were careful to say this is not AGI. But the trend line is what matters. Test-time compute, letting the model think longer, is now a lever you can pull for accuracy, and o3 is the clearest proof yet that it scales. o3 was an announcement, not a release, with safety testing first and a smaller o3-mini promised early in the new year.


Sora Turbo Goes Public and ChatGPT Pro Lands at $200/Month

The same shipping marathon (kicked off December 5) put two other things in builders' hands. On December 9, OpenAI released Sora Turbo to the public at sora.com, ten months after the original demo melted everyone's brain. It generates up to 1080p clips as long as 20 seconds in widescreen, vertical, or square. Access rides on a ChatGPT subscription, US-first, with the EU and UK locked out at launch.

OpenAI also launched ChatGPT Pro at $200 a month, a tier that includes the full o1 model and an o1 pro mode. That is a 10x jump over Plus, and it signals where the frontier is heading: the best reasoning is going to cost real money, and OpenAI is testing how much power users will pay for it.


Google Ships Gemini 2.0 Flash and Veo 2

Google did not sit still. On December 11 it unveiled Gemini 2.0 Flash, the first model in the 2.0 generation, pitched as the start of an "agentic era." Google claims it runs roughly twice as fast as 1.5 Pro while scoring better, keeps the 1 million token context window, and adds native multimodal output, meaning it can emit generated images and text-to-speech audio directly, not just consume them.

Alongside it came Veo 2, Google's text-to-video model, supporting 4K generation and noticeably better physics than the competition. The head-to-head clips against Sora made the rounds for a reason: Veo 2 frequently looked more coherent. The video model race is now genuinely contested, which is good news for anyone who wants these tools to get cheaper and better fast.


Meta Ships Llama 3.3 70B, the Quiet Win for Local AI

On December 6, with far less fanfare, Meta released Llama 3.3 70B Instruct. The pitch is efficiency: similar real-world performance to the 405B Llama 3.1 at a fraction of the serving cost, tuned hard for instruction following, coding, and multilingual work across eight languages.

For the homelab and local-AI crowd, this is arguably the most useful drop of the month. A 70B open-weight model that approaches 405B quality is something you can actually quantize and run on a serious workstation. While the frontier labs fight over $200 tiers and high-compute benchmark runs, Meta keeps handing capable weights to anyone who wants to own their stack. That matters.


Key Takeaways

  • Test-time compute is the new scaling axis. o3 proves you can buy accuracy with more thinking time, which reshapes how you'll budget for hard tasks in 2025.
  • Video generation is now a real race. Sora Turbo shipping and Veo 2 landing in the same week means competitive pressure, and that usually means faster, cheaper tools for builders.
  • The frontier is getting expensive on purpose. ChatGPT Pro at $200/month is a signal that top-tier reasoning will be priced as a premium product, not a commodity.
  • Open weights kept pace. Llama 3.3 70B gives self-hosters near-405B quality at a size they can actually run, keeping local AI viable as the gap to closed models narrows.
  • Agentic framing is now the marketing default. Both OpenAI and Google leaned into multi-step autonomy this month, so expect agent tooling, not just chat, to dominate the roadmap conversations next year.
openaio3gemini-2.0soraveo-2llama-3.3open-weightsvideo-modelsagents
CONSOLE
$