Skip to content

Blog

The 40-Minute Supply Chain Attack That Exposed 434,000 CI/CD Pipelines

The malicious packages behind the largest AI supply chain breach of 2026 survived on PyPI for only 40 minutes (TechJuice). The fallout is still being counted. On August 11, threat intelligence firm CloudSEK published a report linking more than 2,500 organizations and roughly 434,000 software pipelines to the compromise of LiteLLM (CloudSEK via PR Newswire). Independent analysis from Hudson Rock confirmed the scale the next day (Hudson Rock).

LiteLLM is an open-source proxy that gives applications a single API for many large language model providers (CyberInsider). Teams run it as the gateway between their code and models from OpenAI, Anthropic, and others. The library is downloaded more than 95 million times per month (CyberInsider). That reach is why it became a target. An environment running LiteLLM holds API keys, cloud credentials, and configuration files by design.

The attack did not start with LiteLLM. It started with Trivy, the open-source vulnerability scanner (Hudson Rock). TeamPCP, the group behind the campaign, first compromised Trivy’s GitHub Actions pipeline (CyberInsider). The group used an automation token that was rotated but never fully revoked. That gap gave them a 20-day window to force-push malicious code over Trivy’s version tags (TechJuice).

LiteLLM’s own CI pipeline used Trivy to scan its builds. The poisoned scanner had legitimate read access to the build runner. The attackers used that access to exfiltrate LiteLLM’s PyPI publishing tokens (Hudson Rock). With those tokens they published two poisoned releases, versions 1.82.7 and 1.82.8, to PyPI (CyberInsider). The malicious packages were pulled after about 40 minutes (TechJuice). Version 1.82.6 was the last clean release (Endor Labs).

The injection was small and surgical. Twelve lines of obfuscated code were added to a single file, litellm/proxy/proxy_server.py, during the wheel build (CyberInsider). The code decoded a base64 payload and launched it through a Python subprocess when the module was imported. Version 1.82.8 escalated the attack. It added a .pth startup file that runs the payload every time Python starts, even when LiteLLM is never imported (CyberInsider).

The payload harvests a wide credential set. It grabs SSH keys, AWS, GCP and Azure credentials, Kubernetes secrets, environment files, database configurations, and cryptocurrency wallets (CyberInsider). Stolen data is encrypted, packed into a file named tpcp.tar.gz, and exfiltrated to an attacker-controlled domain (CyberInsider). When that path fails, the malware creates a public repository in the victim’s own GitHub account. It uploads the stolen data as a release asset (TechJuice). The payload also moves laterally in Kubernetes. It deploys privileged pods that mount the host filesystem and install a persistent backdoor registered as a systemd service named “System Telemetry Service” (CyberInsider).

CloudSEK identified more than 2,500 organizations potentially impacted. The list spans technology, finance, telecom, cybersecurity, manufacturing, and logistics (CloudSEK via PR Newswire). Hudson Rock obtained a 153GB archive of the stolen data containing 433,909 files. It attributed 118,829 CI runner dumps to 2,488 corporate domains (Hudson Rock). Named victims include NVIDIA, Samsung Electronics, Cisco Systems, Siemens, S&P Global, ServiceNow, and Deloitte (Unite.AI). The trace also surfaced Boeing, Orange, and Roku (TechJuice). The exposed material covers AWS secrets, GitLab identities, Salesforce credentials, Slack tokens, Azure secrets, SSH keys, and AI provider API keys (TechJuice).

An AI gateway is the richest credential store in a modern stack. Every LLM provider key, cloud secret, and pipeline token flows through it. A single poisoned release in that position turns months of build history into an attacker’s keychain. The 40-minute window on PyPI is the core lesson: exposure time no longer measures damage. The packages were published in March 2026, yet organizations are only learning of their exposure in August (Unite.AI).

  1. Revoke, do not just rotate. The entry token was rotated but never revoked (TechJuice). Rotation leaves the old credential alive. Revocation kills it.
  2. Pin with hashes. A lockfile with integrity hashes blocks a malicious release from installing, even when it reaches the index. This is the single cheapest control in the chain.
  3. Separate publish access from build access. The scanner that reads your repo should not also hold your package-publishing tokens (Hudson Rock).
  4. Audit secrets continuously. Environment variables leak into runner dumps and public repos (TechJuice). Scan for them on every run, not once a quarter.
  5. If you ran LiteLLM 1.82.7 or 1.82.8, act now. Treat every credential in that environment as compromised and rotate them. The malware targeted .aws/credentials and .kube/config specifically (TechJuice).

The pattern is familiar to anyone who read our breakdown of credential theft through AI developer tools. The tool that has access becomes the target. The LiteLLM breach just proved it at the scale of the entire AI build ecosystem.

Power Is the New Cloud: Inside Anthropic's $9.1B Data Center Deal with a Bitcoin Miner

On August 10, bitcoin miner Riot Platforms disclosed a 20-year data center lease with a leading frontier AI lab (Riot Platforms, 2026). The deal covers 191 megawatts of critical IT capacity at Riot’s Rockdale, Texas campus (CNBC, 2026). Bloomberg identified the tenant as Anthropic, citing people familiar with the matter (The Decoder, 2026). Neither company confirmed the name publicly. Riot declined to comment, and Anthropic did not respond (crypto.news, 2026).

The contract is expected to generate roughly $9.1 billion in revenue over the initial term, which runs through June 2048 (Riot Platforms, 2026). Two five-year extension options could push the total value to about $16.1 billion (CNBC, 2026). Riot shares jumped roughly 25% in after-hours trading once the deal’s size became public (Quartz, 2026).

This is a colocation agreement, not a cloud contract. Riot builds the data center to the tenant’s specifications and provides the building, power connections, cooling, and operations (The Decoder, 2026). The tenant brings its own servers and AI chips (MLQ, 2026). The 191 MW is enough power for roughly 143,000 homes, per Bloomberg (The Decoder, 2026).

Delivery is phased. The first 96 megawatts go live in December 2027. The full 191 megawatts arrive by June 2028 (Riot Platforms, 2026).

TermDetail
Term length20 years, through June 2048
Capacity191 MW critical IT at Rockdale, Texas
Phase 196 MW by December 2027
Phase 2Full 191 MW by June 2028
Base contract value~$9.1 billion
With both extensions~$16.1 billion
Interim financing$573 million from Morgan Stanley
Riot providesBuilding, power connections, cooling, operations
Tenant providesServers and AI chips

Terms via Riot’s Q2 2026 release. Morgan Stanley’s $573 million interim facility funds initial development while an investment-grade credit backstop is finalized (Quartz, 2026).

Why a bitcoin miner is suddenly a data center developer

Section titled “Why a bitcoin miner is suddenly a data center developer”

Riot is one of the world’s largest bitcoin miners and has owned its power assets for years (Data Center Dynamics, 2026). It controls more than 1,100 acres and 1.7 GW of power capacity across two Texas facilities (Data Center Dynamics, 2026).

The pivot began in January 2026 with Advanced Micro Devices. Riot signed a lease for an initial 25 MW, which it delivered on time and on budget, and a second 25 MW expansion is under construction (Riot Platforms, 2026). That AMD agreement can expand to a total of 200 MW at the campus (Data Center Dynamics, 2026).

The two leases give Riot 241 MW of contracted capacity and about $9.8 billion in long-term contracted revenue (Riot Platforms, 2026). CEO Jason Les called the lease “a defining moment in our evolution into a leading developer of large-scale data centers” (Riot Platforms, 2026).

The financials show the transition in progress. Q2 2026 revenue was $174.2 million, up 14% year over year, with data center revenue of $23.2 million (Riot Platforms, 2026). Riot still posted a net loss of $237.2 million for the quarter (Quartz, 2026). Miners across the sector are chasing the same pivot. Shares of peers IREN, Applied Digital, and TeraWulf moved higher on the news (Yahoo Finance, 2026).

The deal gives Anthropic access to scarce, grid-connected power (CNBC, 2026). Miner campuses already own the hard part of the stack: land, substations, and interconnection rights. A frontier lab cannot wait years for a utility build.

Rockdale is one piece of a much larger Anthropic portfolio. The company is paying SpaceX an estimated $1.25 billion per month through May 2029 for the Colossus 1 data center and plans to deploy two gigawatts of AMD GPUs (The Decoder, 2026). Amazon is investing up to $25 billion toward up to five gigawatts of Trainium capacity (The Decoder, 2026). Gigawatts of Google and Broadcom TPU capacity come online starting in 2027, and a six-year, $10 billion contract with Volta Infra rounds out the portfolio (The Decoder, 2026). Bloomberg also reported a nearly $45 billion compute commitment to xAI in May (Quartz, 2026).

  • Power, not chips, now gates AI capacity. The scarce resource in this deal is 191 MW of interconnected electricity, not GPUs (CNBC, 2026). Capacity planning starts at the substation, not the rack.
  • Colo economics are the new frontier. The landlord supplies shell, power, cooling, and operations. The tenant owns the compute (MLQ, 2026). Budget for hardware separately from facilities.
  • Capacity lands in waves. 96 MW arrives in December 2027, and the rest lands six months later (Riot Platforms, 2026). Plan deployment as two campaigns, not one.
  • Watch the miners. Bitcoin miners hold the interconnected power the AI buildout needs, and they are monetizing it as landlords (Data Center Dynamics, 2026).

You Don't Need Cloud AI for Everything. Route It Instead of Burning It.

Your company is paying $25 per million output tokens for a model that answers the same question 10,000 times a month. The answer never changes. The bill does.

Wendell of Level1Techs made the full case in You Don’t need to use Cloud AI! Switchyard and Nemotron 3.5 Lightning. The argument is blunt: frontier inference for every task is setting money on fire. The fix is a local model router with a feedback loop. NVIDIA shipped both halves this month.

The trap is not just the token price. The trap is turning judgment over to the model.

When a company buys a frontier subscription and turns it loose, it loses the ability to answer three questions. What did the AI do? Why did it take that step? What information did it use? Without those answers, a failure teaches nobody anything (Level1Techs, 2026).

Software engineers want labor augmentation, not delegation. They want supervision. They want structure. They want to know what problem they are solving. A router that logs every decision gives them that. A blank chat window does not.

Switchyard is NVIDIA’s open-source supervision architecture for routing user requests across AI models (NVIDIA Developer Blog).

A request arrives. The router sends it to a specialized local worker, a small model, a customized model, or a tool-calling agent. Each step produces an observable trace. The router accepts the result or escalates it to a stronger model. Human supervision is a first-class component, not an afterthought.

The design gives you three concrete advantages:

AdvantageWhat it buys you
CostA $5-per-million-token local model handles what Opus-class models were doing at $25
ObservabilityYou know which component did what, where, in your organization
Organizational learningEvery routing decision and escalation becomes a dataset

That dataset is the real prize. It shows how your people actually use AI, where they get stuck, and which workflows repeat. It is institutional knowledge, not telemetry (Level1Techs, 2026).

Nemotron 3.5 Lightning: built to be customized

Section titled “Nemotron 3.5 Lightning: built to be customized”

Nemotron 3.5 Lightning is a 30-billion-parameter mixture-of-experts model with 3 billion active parameters (NVIDIA NIM model card). It is fast, it is local, and it is explicitly pitched as a model you are supposed to customize (NVIDIA Developer Blog).

Customization does not mean full retraining. NVIDIA ships the LoRA recipes, the supervised fine-tuning setup, the reinforcement learning config, and the training data. You freeze the base model and train a small set of additional weights for your specific job.

NIM takes it one step further. It keeps one base model resident and dynamically loads and unloads LoRA adapters while serving (NVIDIA Developer Blog).

Accounting gets the accounting adapter. Software engineering gets the code-review adapter. Support gets the support adapter. That undocumented internal product from 2017 gets the adapter containing the dark knowledge known only to Gary. Gary can finally take a vacation.

One base model, many specialists, no retraining, no cloud round-trips.

NVIDIA’s data flywheel blueprint makes the loop explicit (NVIDIA Developer Blog):

  1. Instrument the AI application and log production traffic.
  2. Build evaluation and fine-tuning datasets from those logs.
  3. Evaluate smaller models against the data.
  4. Customize the ones that work.
  5. Promote them and measure again.

You cannot improve what you do not measure. The flywheel saves tokens because the small local model is cheaper on every request. It saves sanity because every escalation is a recorded decision, not a guess.

The benchmark backs it up. On Humanity’s Last Exam, an NVIDIA-orchestrated system scored 37.1% versus 35.1% for GPT-5, at 30% of the cost and 2.5 times faster (Artificial Analysis, Wikipedia). The small model beat the frontier model at the task because it was inside a tool-calling ecosystem, not because it tried to know everything.

Wendell’s sharpest observation is the dark pattern. Recent Codex CLI updates surface less reasoning detail in the UI, and there are open issues about reasoning summaries missing from the session log (Level1Techs, 2026).

High-quality outputs and task traces are exactly the raw material you need to train a cheaper specialized model. OpenAI even sells model distillation around that idea. There is an economic incentive for frontier providers to keep useful internal signals from being trivially exportable.

That is the argument for owning your loop. The company that routes and logs its own AI traffic stops renting intelligence and starts accumulating it.

  1. Route by task class. Send simple, repeated tasks to a small local model. Send only the hard edge cases to the frontier.
  2. Capture every decision. Log what the model did, why, and what the human corrected. That log is your training data.
  3. Customize what repeats. If users ask the same question 10,000 times, give the small model a LoRA adapter that answers it from your own data.

The era of using a frontier model for every little task is ending. The high water mark for cloud token spend is here. The machines that replace it are smaller than you think, and they sit on a desk.

Muse Glimmer: Meta's 30B Open-Weight Agent Runs on One GPU

On August 10, Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter model built for always-on local agent workflows (Phoronix, 2026). The weights ship under Apache 2.0, and the model runs on a Mac or PC with a single consumer GPU (TechCrunch, 2026).

The same day, Mark Zuckerberg published a 6,500-word essay, “The Future is for Everyone,” on Meta’s site (Meta, 2026). He argues that AI concentrated in a few hands leads to worse outcomes for everyone else (AP News, 2026). Meta also promised to open the weights of Muse Spark 1.2, its most capable foundation model, within weeks (Ars Technica, 2026).

This is Meta’s first fully open release since the proprietary Muse Spark replaced the open-weight Llama family in April (VentureBeat, 2026). For DevOps and AI teams the change is practical. Capable agents no longer need a cloud API call.

Muse Glimmer is a dense causal transformer with a dedicated vision encoder (Hugging Face model card).

SpecValue
Parameters~30B total, 29.6B across 52 layers
Vision encoder~1.8B ViT-G/14
LicenseApache 2.0
Context window131,072+ tokens
Languages100+
Input / outputText and image in, text out
Knowledge cutoffJanuary 4, 2026
Target hardwareOne consumer GPU or a Mac

Specs via VentureBeat and the official page at developer.meta.com.

Glimmer is distilled from Muse Spark, the larger closed model Meta launched in April 2026 (TechCrunch, 2026). Training used logit distillation from Spark outputs, then agent-focused mid-training, supervised fine-tuning, and reinforcement learning (Neowin, 2026).

The hardware math is the interesting part. A 30B model needs over 55 GB of memory at full precision. Meta compresses it to roughly 4-bit and adds block-level speculative decoding with a DFlash drafter head, so it answers fast enough for a real agent loop (MarkTechPost, 2026). Quantized variants target 24/32 GB consumer cards (Hugging Face, 2026).

The benchmarks hold up against local rivals. Glimmer beats Gemma4-31B and Qwen3.6-27B on several popular LLM benchmarks (Neowin, 2026). Meta publishes IFBench 77.0, AIME 2026 94.7, and GPQA Diamond 83.5 (developer.meta.com, 2026).

The tooling is ready on day one. The weights are on Hugging Face now, and Ollama 0.32.7 added support the same day (Phoronix, 2026). Ollama, LM Studio, vLLM, SGLang, Together AI, Fireworks AI, and OpenRouter support is rolling out this week (VentureBeat, 2026).

Agents go local. Always-on agents currently mean per-token cloud bills and prompts crossing a network. Glimmer runs with no network call (MarkTechPost, 2026). Your code, logs, and prompts stay on your machine.

The frontier follows. Muse Spark 1.2 is the model behind Muse Code, the terminal coding agent Meta shipped on August 5 (VentureBeat, 2026). Opening those weights puts a shipped coding agent’s brain in your hands.

Policy is part of the pitch. Zuckerberg says US labs face extra restrictions on training data while Chinese open-weight models from DeepSeek, Alibaba, and Z.ai gain US traction at lower cost. He calls on Washington to lower the barriers (NY Post, 2026).

Governance is promised. Meta says its board will approve the safety criteria for model releases and will review each release against them (Forbes, 2026). A $1 billion “Future is for Everyone” fund targets communities hosting Meta data centers (Axios, 2026).

  • Pull the weights from meta-models/Muse-Glimmer-30B on Hugging Face.
  • Serve it with Ollama or vLLM on a 24/32 GB GPU, or a Mac with enough unified memory.
  • Point an agent scaffold such as OpenClaw at the local endpoint (Neowin, 2026).
  • Benchmark it against Gemma4-31B and Qwen3.6-27B on your own tasks before migrating. The 4-bit path trades some quality for the single-GPU fit.
  • Watch for Muse Spark 1.2 weights in the coming weeks (Ars Technica, 2026).

Meta’s open-source return is a bet. Distribute the models, keep the ecosystem, and let anyone run agents on hardware they own. The next few weeks will show whether Spark 1.2 follows through.

AMD Buys Taalas: What Happens When a Model Is Etched Into Silicon

AMD Buys Taalas: What Happens When a Model Is Etched Into Silicon

Section titled “AMD Buys Taalas: What Happens When a Model Is Etched Into Silicon”

On August 6, AMD announced a definitive agreement to acquire Taalas, a Toronto-based startup that bakes model weights directly into silicon (AMD press release, 2026). The deal targets the fastest-growing segment of the AI market: inference (CNBC, 2026).

Taalas calls its approach “the model is the computer” (Taalas, 2026). Instead of loading weights from memory, the chip etches them into the silicon itself. The result is a fixed-function ASIC that runs one model, and only that model (Anurag Kushwaha, 2026).

A GPU spends most of its time moving weights from HBM into compute cores. Every token re-reads the model from memory. That constant traffic is the memory wall, and it is the main cost driver for inference at scale (Anurag Kushwaha, 2026).

Taalas removes the wall. The weights live in the silicon as physical transistors, so data flows through the layers as a continuous electrical signal. No HBM, no repeated fetches (The Register, 2026).

The first test chip, HC1, was fabbed on TSMC’s 6nm process. It serves Meta’s Llama 3.1 8B at roughly 17,000 tokens per second (Taalas, 2026). When announced, that was about 48x faster than Nvidia GPUs and 8.5x faster than Cerebras accelerators at the same task (The Register, 2026).

The chip is married to its model. A bigger change than a LoRA adapter means a re-spin of the silicon (The Register, 2026). Taalas says a re-spin touches only two metal layers, which is cheaper than a full redesign, but it still takes time and money (The Register, 2026).

AMD plans to pair the technology with its Helios rack-scale systems and Instinct GPUs (AMD press release, 2026). That suggests a split workload: GPUs handle prompt processing, and Taalas chips generate tokens (The Register, 2026).

The model release cycle is now the hardware refresh cycle. A model-specific chip is only worth deploying when you are confident the model will stay in production long enough to pay for the silicon. That fits stable, high-volume workloads like code assistants and chat at massive scale.

For everyone else, the practical takeaway is simpler: inference cost is now a hardware design problem, not just a software one. When a vendor locks a model into a chip, the economics flip. The fast path and the flexible path are no longer the same path.