Skip to content

Blog

Hands-On with GPT-5.2: What It Actually Delivers on Real Projects

GPT-5.2 in extended thinking mode can produce a complete project in one session. In testing, it built a complete 3D game as a downloadable zip file, with no code snippets to assemble by hand. This post reviews what the model delivered and what the published benchmark numbers do and do not show.

Start with the city destruction demo. Prompted to build a game where players fly through skyscrapers and fire miniguns and rockets, GPT-5.2 returned a full Three.js project folder. It included destructible environments, physics, a scoring system, and interactive controls. The zip file ran directly in a browser.

A second demo generated a 3D planet running Conway’s Game of Life, with asteroid impacts, bloom effects, meteor intervals, and pause controls. A third produced a tour of sci-fi megastructures, including Dyson spheres and orbital elevators, with autopilot fly-throughs and adjustable field of view. These builds took 20-55 minutes of extended reasoning each.

The GDP-Val benchmark tries to measure project-level work. It assigns tasks that mimic actual jobs. A manufacturing engineer designs a 3D cable reel stand with exploded views. A financial analyst maps the last-mile delivery market. A nurse analyzes skin lesion images and drafts a consultation report. An event planner optimizes vendor fair layouts or builds a luxury itinerary.

Human experts with an average of 14 years of experience judge the outputs blind. The judges come from firms including Goldman Sachs, Boeing, Google, and the US Department of Defense. They rate quality, completeness, and adherence to the spec.

OpenAI reports that GPT-5.2 Pro won or tied 74% of its matchups against the experts, with 60% outright wins. For comparison, GPT-5 High scored 38.8% and Claude 4.1 Opus 47.6% on the same benchmark in September 2025.

ModelWin/Tie RateWin Rate
Claude 4.1 Opus (Sept 2025)47.6%~35%
GPT-5 High (Sept 2025)38.8%~25%
GPT-5.2 Pro74%60%

One judge’s feedback read: “Exciting and noticeable, appears done by a professional company with staff, surprisingly well-designed layout.” That is a subjective read, and it is the read the benchmark is built on.

What the numbers do not show: judges pick the better deliverable, which is a preference call, not a measure of correctness. The benchmark comes from OpenAI, and the judging methodology is not fully public. Treat 74% as a reported result, not a settled fact.

GPT-5.2 also reports 100% on AIME 2025 and over 90% on ARC-AGI in extended mode, with gains on SWE-Bench Verified. The cost figure is more concrete. OpenAI says the price of a complex task dropped by a factor of 390 in one year. A task that cost $45,000 a year ago runs about $115 at the reported rate.

The useful frame is a rapid contractor, not a labor replacement. GPT-5.2 produces a deliverable in 20-55 minutes and takes another 20-30 minutes per revision. Early glitches appeared in testing, such as overexposed lighting, and prompts like “single-file output” fixed them. Output still needs review. A higher benchmark score does not remove hallucination risk.

The GDP-Val result is the most concrete evidence yet that model output can match experienced professionals on broad project tasks. It is not evidence that jobs disappear. Adoption depends on review workflows, liability, and trust, none of which benchmarks measure. Put plainly: GPT-5.2 is a capable project generator with high reported benchmark scores, and the labor-replacement claim remains unproven.

GPT-5.2, Runway 4.5, and Image AI: A Release Roundup

Three releases landed this week. OpenAI shipped GPT-5.2, Runway deployed Gen-4.5, and the industry formed a standards body for AI agents. OpenAI also announced a $1 billion investment from Disney. The announcements are below, with the numbers as reported.

GPT-5.2: Specs and First Benchmark Results

Section titled “GPT-5.2: Specs and First Benchmark Results”

OpenAI launched GPT-5.2 after a short delay. The release follows complaints that GPT-5.1 was unreliable on accuracy. The model ships with a 400,000-token context window, about 300,000 words, and a 128,000-token output limit. API pricing is $1.75 per million input tokens and $14 per million output tokens.

On SWE-bench Pro, GPT-5.2 scores 55.6%. That is up from 50.8% for GPT-5.1. Claude Opus 4.5 sits at 52%, and Gemini 3 Pro at 43.3%. These are vendor-reported figures on one benchmark. Independent comparisons are still thin, and accuracy tests in production settings are pending.

OpenAI announced a $1 billion investment from Disney. The deal gives OpenAI access to Disney’s IP library for Sora video generation and the native image tools. Possible products include personalized Disney+ shorts, such as AI-generated clips of Disney characters.

GPT-5.2 ships with native image generation. In testing, the model renders photoreal portraits, readable text, and code overlays. Examples include whiteboard slogans and JSON overlays on product shots. It shows fewer proportion errors than earlier GPT image models. Subtle artifacts remain in eyes and skin, and results vary on recognizable faces.

Agentic AI Foundation: A Standards Body for Agents

Section titled “Agentic AI Foundation: A Standards Body for Agents”

OpenAI, Anthropic, and Block launched the Agentic AI Foundation under the Linux Foundation. Google, Microsoft, Amazon, Bloomberg, and Cloudflare back the group. The goal is a common standard so agents from different vendors operate across apps under the same safety rules. Without such a standard, agents that handle email, bookings, and troubleshooting risk locking users into one vendor.

Runway started deploying Gen-4.5 this week. Runway calls the results state-of-the-art for motion, physics, and prompt adherence, and the model leads its internal text-to-video charts. It simulates weight, fluid dynamics, and consistent faces. It does not generate audio.

Hands-on tests of the deployed model:

  • Glass sphere on marble stairs: realistic bounces, water splashes, and refractions. The prompt match is close.
  • Rainy street walker: umbrella physics, a subtle smile, and handheld camera jitter read correctly.
  • Anime explorer: foreground consistency holds. The background is unstable.
  • Barista latte pour: swirling milk, steam, and blurred patrons look correct.
  • Neon alley chase: reflections are accurate. Minor physics and camera errors appear in the 5-second clip.

Prompt fidelity is the model’s main advantage. Veo 3.1 still leads on realism and sound integration.

  • Mistral released Devstral 2, a coding model with public weights. It scores 72.2% on internal benchmarks, close to DeepSeek v3.2.
  • Zhipu AI released GLM-4.6V, a vision model for tool calling. Qwen updated Omni Flash with more lifelike voices.
  • OpenAI paused shopping suggestions that looked like ads and added user controls.
  • ChatGPT gained Adobe connectors for Acrobat, Express, and Photoshop. Early tests show actual limits.
  • Meta took over the Limitless pendant, an always-on audio recorder. Privacy questions remain unanswered.
  • Alibaba released Image2LoRA, which builds style and character LoRAs from a single image.

At Rivian’s AI and Autonomy Day, the company showed custom silicon built with Nvidia and integrated LiDAR. Its roadmap targets hands-free driving and unsupervised Level 4 operation by 2027-28. A voice assistant handles calendar, messages, and car controls.

McDonald’s released a fully AI-generated holiday ad. It drew criticism for looking low-budget beside the company’s production spend. Commenters asked for work by people, with AI used in limited roles.

The week’s releases show a maturing market: specialized models, a standards body, and clearer pricing. The figures above come from the vendors. Independent testing will decide which claims hold.

Google Coral Edge TPU on a Raspberry Pi: An AI Accelerator Overview

Imagine taking the pocket-sized Raspberry Pi—a board beloved by hobbyists for its affordability and versatility—and transforming it into a beast capable of real-time video object recognition, one of the most demanding tasks in computer science. That’s exactly what Google’s latest Coral AI Edge TPU promises, and recent hands-on tests confirm it’s no hype.

At the heart of this upgrade is the Coral AI Edge TPU, a compact accelerator designed exclusively for machine learning inference. It’s not about raw CPU power; this USB stick-sized device offloads neural network computations from the Pi’s general-purpose processor, delivering speeds that make high-end GPUs blush on low-power setups. Priced accessibly and built for edge devices, it bridges the gap between cloud AI and on-device processing, enabling applications from smart cameras to autonomous drones without internet dependency.

Getting started is deceptively simple. Attach a compatible camera module to your Raspberry Pi, plug the Edge TPU into a USB port, and power up. Head to coral.ai for the essential packages—PyCoral libraries and model zoos—which install via a few terminal commands. No PhD required; even if the code looks like ancient runes at first glance, it’s plug-and-play for most.

Pre-built models are ready to roll. Point the setup at a snapshot of a bird, and in a blink—faster than you can say “neural net”—it classifies the feathered friend with pinpoint accuracy. The TPU’s magic shines here: inference times plummet from seconds on the Pi alone to mere milliseconds.

Real-Time Video: Where the Rubber Meets the Road

Section titled “Real-Time Video: Where the Rubber Meets the Road”

Static images are child’s play. The real test? Live video detection. Fire up the video object detection script from Coral’s repo, and you’re off to the races. In a demo, the rig effortlessly tracked a person striding into frame, guitar in hand, tagging it with a staggering 91% confidence score. No lag, no dropped frames—just smooth, responsive AI on hardware that costs less than a decent dinner out.

This isn’t throttled lab performance; it’s sustained operation on a device sipping power like a miser. The Pi’s CPU idles while the TPU crunches tensors, freeing resources for other tasks.

For tinkerers, it’s a game-changer: home security cams that spot intruders, wildlife monitors identifying species, or robotic arms sorting recyclables—all running locally with privacy intact. Developers gain a scalable path to production edge AI, unburdened by cloud costs or latency.

Google’s Coral ecosystem keeps expanding, with dev boards, PCIe cards, and more models incoming. Pair this with the Pi’s GPIO pins, and the possibilities explode—IoT gateways, portable analyzers, you name it.

The verdict? Yes, the Raspberry Pi can handle “supercomputer” workloads for AI inference. Grab a Coral Edge TPU, and watch your projects soar from toy to titan.

A word of caution for the eager maker: “Supercomputer” power generates supercomputer heat. The Coral USB Accelerator can get very hot—often exceeding 60°C (140°F) under load. If it overheats, it throttles performance to protect itself, killing that “real-time” responsiveness. Don’t just plug it in and bury it in an enclosure. Use a USB extension cable to keep it away from the Pi’s own heat, and consider a small heatsink or fan if you’re planning 24/7 inference. It sips power, but it spits fire—plan accordingly.

DeepMind's AGI Claims: What the Announcement Actually Says

Google DeepMind published a podcast episode titled “The Arrival of AGI” with co-founder Shane Legg and host Hannah Fry. Around the same time, OpenAI’s Sam Altman posted a decade retrospective that predicts superintelligence within about ten years. Neither item is a technical result. Both are claims. This post separates the claims from the evidence cited.

Legg’s central claim is economic. He argues that AI will replace the exchange of mental and physical labor for resources, the arrangement that underpins hunter-gatherer tribes, medieval serfdom, and modern jobs. He compares a post-labor society to house cats, which are sustained without contributing and sleep about 18 hours a day. Education, he argues, would need to stop training people for economic roles that may not exist.

Altman’s claim is a timeline. He writes: “In 10 more years, we are almost certain to build superintelligence.” He also defends iterative deployment, releasing models in stages so society adapts as capabilities change. The retrospective reviews a decade of releases, from the 2017 Dota reinforcement learning work and the unsupervised sentiment neuron to ChatGPT in 2022.

A chart from the Federal Reserve Bank of Dallas circulated with the discussion. It plots US GDP per capita over 150 years and forks after 2035 into two paths: a benign singularity with steep growth, and an extinction path at zero. It is a scenario illustration from a bank research department, not a forecast with probabilities. Legg and Altman cite it to frame the stakes.

Epoch AI’s capability indexes show no plateau in measured benchmark trends, which supports the claim that scaling continues. Independent evals such as AI Village run top models on tasks with internet and tool access. Current agent tools are the concrete part. At AWS re:Invent 2025, Frontier Agents such as Kirao triaged bugs and handled developer backlogs. Amazon’s Nova 2 family covers voice (Sonic), multimedia (Omni), and UI automation (Act). Bedrock Agent Core adds policy controls, and Trainium 3 Ultra scales inference at lower cost. China’s pilot programs license robotaxis in stages to pace job displacement. These are working systems, but they perform narrow tasks.

None of this establishes that AGI has arrived. There is no agreed definition of AGI, so the episode title is a position, not a measurement. The evidence is a mix of scenario charts, extrapolated trend lines, and speaker opinion. The Dallas Fed chart describes possible futures, not observed outcomes. Altman’s ten-year window is a prediction. Legg’s labor argument assumes current scaling continues without interruption. The systems in operation handle bounded tasks with tool access. General reasoning across the full range of paid work remains unmeasured.

Treat AGI announcements as claims with attached evidence, and grade each piece of evidence on its own. A scenario chart is not a prediction. A benchmark trend is not a capability. Until a system demonstrates broad competence across the economy without hand-holding, the arrival of AGI is a thesis, not a fact.

GitHub Actions Sleep-Loop Bug: Years of Failed Runs and the Cost to Developers

A four-line Bash function in the GitHub Actions runner has caused failed runs and idle machines for years. The function was meant to pause execution briefly. On busy runners it can spin forever, hold a CPU core at 100%, and leave a job running long after it should have ended. One developer reported 5,135 hours of billed idle time on a single runner.

The runner codebase started around 2016. Early commits show Windows developers borrowing a Stack Overflow trick from 16 years earlier: use ping to simulate a delay where sleep is not available. That pattern became a top-level function called safe_sleep:

Terminal window
if [ $? -eq 4 ]; then
sleep 5 || ping -n 6 127.0.0.1 > nul || (for i in `seq 1 5000`; do echo >&5; done)
fi

The chain tries sleep first, then ping for about 4 seconds, then 5,000 echo writes to /dev/null. It worked, at the cost of CPU time.

In 2022 the code changed to a tighter loop:

Terminal window
start=$SECONDS
while [ $((SECONDS - start)) -ne ${1?} ]; do :; done

The intent: read Bash’s SECONDS variable, which increments every second, and loop until the target is reached. The flaw shows up under load. If scheduling makes SECONDS jump from 4 to 6, the comparison -ne 5 is never false, so the loop never exits. With no sleep inside the loop, it pegs one CPU core at 100%, half of a standard 2-vCPU runner, and starves other tasks on the machine.

One developer measured a single runner idling for 5,135 hours. At GitHub’s $0.08 per vCPU-minute rate, that is about $2,400 in billed time for one machine. Projects felt the effect too. Zigg moved to Codeberg, citing “inexcusable bugs” and what it called “vibe scheduling” after Microsoft’s AI pivot, with job prioritization that stalled even main-branch commits.

A fix surfaced in 2024: use -le instead of -ne, so the loop stops once the target is reached or passed:

Terminal window
while [ $((SECONDS - start)) -le ${1?} ]; do :; done

A pull request with this change was proposed in February 2022. It was auto-closed after a month and merged about 1.5 years later, after public complaints. Other regressions followed the same pattern. A later refactor replaced a plain Object.getOwnPropertyNames call with nested loops and redundant if statements, and file hashing broke as a result:

function getKeys(obj) {
return Object.getOwnPropertyNames(obj);
}

The fix was small and had been available for years.

The runner is a shared codebase with a long review backlog, and the sleep function sits in a busy path where changes carry risk. Matt Lad of Antithesis summed up the engineering assessment: this is not peak engineering. The platform bills per minute, so a looping sleep is not harmless. Teams that depend on Actions should audit runner logs for jobs that run far past their expected time, and pin runner versions that contain the fix.