Skip to content

AI Agents

2 posts with the tag “AI Agents”

GPT-5.2, Runway 4.5, and Image AI: A Release Roundup

Three releases landed this week. OpenAI shipped GPT-5.2, Runway deployed Gen-4.5, and the industry formed a standards body for AI agents. OpenAI also announced a $1 billion investment from Disney. The announcements are below, with the numbers as reported.

GPT-5.2: Specs and First Benchmark Results

Section titled “GPT-5.2: Specs and First Benchmark Results”

OpenAI launched GPT-5.2 after a short delay. The release follows complaints that GPT-5.1 was unreliable on accuracy. The model ships with a 400,000-token context window, about 300,000 words, and a 128,000-token output limit. API pricing is $1.75 per million input tokens and $14 per million output tokens.

On SWE-bench Pro, GPT-5.2 scores 55.6%. That is up from 50.8% for GPT-5.1. Claude Opus 4.5 sits at 52%, and Gemini 3 Pro at 43.3%. These are vendor-reported figures on one benchmark. Independent comparisons are still thin, and accuracy tests in production settings are pending.

OpenAI announced a $1 billion investment from Disney. The deal gives OpenAI access to Disney’s IP library for Sora video generation and the native image tools. Possible products include personalized Disney+ shorts, such as AI-generated clips of Disney characters.

GPT-5.2 ships with native image generation. In testing, the model renders photoreal portraits, readable text, and code overlays. Examples include whiteboard slogans and JSON overlays on product shots. It shows fewer proportion errors than earlier GPT image models. Subtle artifacts remain in eyes and skin, and results vary on recognizable faces.

Agentic AI Foundation: A Standards Body for Agents

Section titled “Agentic AI Foundation: A Standards Body for Agents”

OpenAI, Anthropic, and Block launched the Agentic AI Foundation under the Linux Foundation. Google, Microsoft, Amazon, Bloomberg, and Cloudflare back the group. The goal is a common standard so agents from different vendors operate across apps under the same safety rules. Without such a standard, agents that handle email, bookings, and troubleshooting risk locking users into one vendor.

Runway started deploying Gen-4.5 this week. Runway calls the results state-of-the-art for motion, physics, and prompt adherence, and the model leads its internal text-to-video charts. It simulates weight, fluid dynamics, and consistent faces. It does not generate audio.

Hands-on tests of the deployed model:

  • Glass sphere on marble stairs: realistic bounces, water splashes, and refractions. The prompt match is close.
  • Rainy street walker: umbrella physics, a subtle smile, and handheld camera jitter read correctly.
  • Anime explorer: foreground consistency holds. The background is unstable.
  • Barista latte pour: swirling milk, steam, and blurred patrons look correct.
  • Neon alley chase: reflections are accurate. Minor physics and camera errors appear in the 5-second clip.

Prompt fidelity is the model’s main advantage. Veo 3.1 still leads on realism and sound integration.

  • Mistral released Devstral 2, a coding model with public weights. It scores 72.2% on internal benchmarks, close to DeepSeek v3.2.
  • Zhipu AI released GLM-4.6V, a vision model for tool calling. Qwen updated Omni Flash with more lifelike voices.
  • OpenAI paused shopping suggestions that looked like ads and added user controls.
  • ChatGPT gained Adobe connectors for Acrobat, Express, and Photoshop. Early tests show actual limits.
  • Meta took over the Limitless pendant, an always-on audio recorder. Privacy questions remain unanswered.
  • Alibaba released Image2LoRA, which builds style and character LoRAs from a single image.

At Rivian’s AI and Autonomy Day, the company showed custom silicon built with Nvidia and integrated LiDAR. Its roadmap targets hands-free driving and unsupervised Level 4 operation by 2027-28. A voice assistant handles calendar, messages, and car controls.

McDonald’s released a fully AI-generated holiday ad. It drew criticism for looking low-budget beside the company’s production spend. Commenters asked for work by people, with AI used in limited roles.

The week’s releases show a maturing market: specialized models, a standards body, and clearer pricing. The figures above come from the vendors. Independent testing will decide which claims hold.

DeepMind's AGI Claims: What the Announcement Actually Says

Google DeepMind published a podcast episode titled “The Arrival of AGI” with co-founder Shane Legg and host Hannah Fry. Around the same time, OpenAI’s Sam Altman posted a decade retrospective that predicts superintelligence within about ten years. Neither item is a technical result. Both are claims. This post separates the claims from the evidence cited.

Legg’s central claim is economic. He argues that AI will replace the exchange of mental and physical labor for resources, the arrangement that underpins hunter-gatherer tribes, medieval serfdom, and modern jobs. He compares a post-labor society to house cats, which are sustained without contributing and sleep about 18 hours a day. Education, he argues, would need to stop training people for economic roles that may not exist.

Altman’s claim is a timeline. He writes: “In 10 more years, we are almost certain to build superintelligence.” He also defends iterative deployment, releasing models in stages so society adapts as capabilities change. The retrospective reviews a decade of releases, from the 2017 Dota reinforcement learning work and the unsupervised sentiment neuron to ChatGPT in 2022.

A chart from the Federal Reserve Bank of Dallas circulated with the discussion. It plots US GDP per capita over 150 years and forks after 2035 into two paths: a benign singularity with steep growth, and an extinction path at zero. It is a scenario illustration from a bank research department, not a forecast with probabilities. Legg and Altman cite it to frame the stakes.

Epoch AI’s capability indexes show no plateau in measured benchmark trends, which supports the claim that scaling continues. Independent evals such as AI Village run top models on tasks with internet and tool access. Current agent tools are the concrete part. At AWS re:Invent 2025, Frontier Agents such as Kirao triaged bugs and handled developer backlogs. Amazon’s Nova 2 family covers voice (Sonic), multimedia (Omni), and UI automation (Act). Bedrock Agent Core adds policy controls, and Trainium 3 Ultra scales inference at lower cost. China’s pilot programs license robotaxis in stages to pace job displacement. These are working systems, but they perform narrow tasks.

None of this establishes that AGI has arrived. There is no agreed definition of AGI, so the episode title is a position, not a measurement. The evidence is a mix of scenario charts, extrapolated trend lines, and speaker opinion. The Dallas Fed chart describes possible futures, not observed outcomes. Altman’s ten-year window is a prediction. Legg’s labor argument assumes current scaling continues without interruption. The systems in operation handle bounded tasks with tool access. General reasoning across the full range of paid work remains unmeasured.

Treat AGI announcements as claims with attached evidence, and grade each piece of evidence on its own. A scenario chart is not a prediction. A benchmark trend is not a capability. Until a system demonstrates broad competence across the economy without hand-holding, the arrival of AGI is a thesis, not a fact.