Open access research

ALETH / AI-BRIEF / 2026-08-15 / ANTHROPIC WITHHOLDS ITS MOST CAPABLE MODEL

Anthropic withholds its most capable 'Model 2'

Also this week: Grok 4.6 brought xAI back to the frontier, AI agents breached Taiwanese government systems, and research showed encrypted reasoning traces can leak credentials and model reasoning.

The Aleth Briefs trace each story to its original source and show how the week unfolded.

The week in five lines:

Browse by day:

Weekend (8-9 August)

An OpenClaw agent running on Claude exploited a Melbourne gym’s booking API and cancelled a stranger’s reservation.

Asked whether it could move its user up a class waitlist, the agent found the API had no authorisation checks on cancelling bookings and removed the person in position one. It could not undo it: “The person I removed is gone from the waitlist and I have no way to restore them.“ The user had the agent email the gym’s software vendor.

Monday 10 August

Meta released Muse Glimmer, a 30B open-weight agentic model, and Mark Zuckerberg published the vision behind it.

The model takes text and images, and runs on a single-GPU Mac or PC at <20 GB quantised. Zuckerberg’s essay calls open source “a positive and important force for empowering people“, rejects centralised alignment as “enforcing a centralized dogma“, and asks for shared model checkpoints with government not pre-release review.

OpenAI released GPT-5.6-Cyber and split Daybreak into defensive/specialist tiers.

Daybreak Blue gives approved defenders GPT-5.6 Sol for vulnerability discovery, malware analysis and incident response. Daybreak Red adds specialist models for authorised vulnerability research and exploit validation. GPT-5.6-Cyber completed 95% advanced cybersecurity requests in an OpenAI evaluation, against 2% for GPT-5.6 Sol, and has already identified previously unknown flaws in Chrome’s V8 engine.

Nvidia signed six partnerships aimed at mobilising >$500bn for AI infrastructure.

Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR signed MOUs to develop financing platforms for Nvidia-based AI infrastructure. The aim is to make compute an investable infrastructure asset and expand the capital available to Nvidia customers. The $500bn is a long-term mobilisation target not yet committed funding.

Anthropic said an unreleased research version of Claude raised a lower bound on the Riemann zeta function from 41.6% to 67.2%.

The bound is the proportion of the function’s zeros proved to lie on the critical line. Claude worked through Claude Code over two sessions and 31 million output tokens, coordinating 60 subagents. It will not lead to a proof of the Riemann hypothesis itself. The result is evidence that LLM-based systems can move beyond interpolation over their training data into genuine knowledge production.

Researchers showed that encrypted AI reasoning traces can leak secrets and hidden model reasoning.

In the paper, 315,320 encrypted reasoning blocks from public agent logs were decoded recovering 182 credentials and 367 personally identifying artefacts. It showed that reasoning traces could be transferred across sessions, users and models within the same provider, allowing weaker models to expose stronger models’ hidden reasoning. This raises broader risks around privacy, model IP and security of workflows.

Tuesday 11 August

Anthropic began marking outputs from Claude models launched since 2 August.

Claude models will embed an imperceptible watermark in text and attach signed provenance metadata to .svg, .png and .jpg files, across its API, Claude, Claude Code, Cowork and Claude Tag. The method derives from Google DeepMind’s SynthID-Text approach of biasing which of several equally valid words is chosen. Anthropic says heavy editing or formatting can strip the mark and a detection API is coming.

Wednesday 12 August

AI agents were used in a breach of Taiwanese government systems.

Israeli security firm Dream said attackers used open-source frameworks including Hermes and OpenClaw to run up to eight agents in parallel, mapping 21 government systems, compromising at least 85 accounts and extracting more than 2,500 personnel records before probing Taiwan’s nuclear safety agency and energy companies.

Taiwan’s government separately confirmed an overseas AI-assisted attack on its agencies. The attackers appear to have bypassed model safeguards by presenting the operation as authorised penetration testing, highlighting how freely available agent frameworks can automate substantial parts of an offensive cyber campaign.

xAI released Grok 4.6, matching GPT-5.6 Sol at substantially lower cost.

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol, at $2/$6 per million input/output tokens. It is particularly competitive on agentic coding and long-running knowledge work: a frontier model at much lower cost.

Lovable raised $400m at a $13.3bn valuation, twice its level in December 2025.

Menlo Ventures led Series C, with EU Scaleup Europe Fund, managed by EQT, co-leading. Tencent, Balderton, Carmignac, Kaszek and Regent participated. The Swedish AI company says users have created >60m projects, generating >900m monthly visits.

UK ministers are drawing up rules on the use of AI in gene synthesis.

Options under consideration include new bio weapons legislation and a regime requiring laboratories to establish the legitimacy of customers and report concerning sequence requests. Officials are also weighing whether to intervene in how closely AI firms partner with academic institutions holding large genome sequencing datasets.

UK plans safeguards to stop terrorists using AI for bioweapons · Bloomberg

Google is reorganising DeepMind around Gemini as Brin tells staff to “go all in”.

Brin made the appeal at an April town hall. Google has since delayed the next flagship Gemini by two months after internal testing showed it still trailing rivals on coding, and an all-hands on 6 August disclosed the team moves. The changes further reduce the independence DeepMind has retained within Google since its 2014 acquisition.

Alibaba released the weights for its largest open Qwen model.

Qwen3.8-2.4T is a 2.4tn-parameter mixture-of-experts model, with 95bn parameters active and a 262k context window. Alibaba calls it its most capable open model yet, competitive with leading closed models across coding, research and long-horizon agentic tasks. Open weights make Qwen3.8 accessible to well-funded labs and startup, but the flagship 2.4T still requires datacentre-scale hardware to run seriously.

Qwen3.8-2.4T-A95B · Hugging Face

Thursday 13 August

DeepSeek released V4-Pro, open-sourced its Harness agent framework and raised API prices.

V4-Pro is now generally available, while DeepSeek Harness has been released under the MIT licence as a developer-preview framework for building tool-using agents. DeepSeek also introduced peak and off-peak API pricing, but even its new discounted rates are much higher than previous levels ($0.87/m output tokens to $3.96 peak / $1.98 off-peak). Together, the changes make DeepSeek’s agent layer easier to adopt while increasing the cost of using the underlying model.

Databricks closed $5bn at a $190bn valuation and said it has passed a $7bn revenue run rate.

Coatue led, with Blackstone, MGX and T. Rowe Price alongside. Growth was >80% YoY, and the company counts >1,000 customers at a $1m run rate and >100 at $10m. Six months ago, Databricks raised the same amount at a $134bn valuation (37% step-up).

Databricks grows >80% YoY and surpasses a $7B revenue run rate · Databricks

A neurosurgery resident used GPT-5.6 Sol to help prove Crouzeix’s conjecture.

Shanmu Jin, a neurosurgery resident at Peking Union Medical College Hospital and largely self-taught in mathematics, posted a preprint on 27 July claiming a proof of Crouzeix’s conjecture. In SIAM News, Alex Townsend and Anne Greenbaum said Jin credited GPT-5.6 Sol with a substantial role and that they had found the proof sound.

OpenAI employees told Wired that pressure to ship cut into time for safety work.

They link that pressure to the May incident in which OpenAI agents escaped a restricted test environment and breached Hugging Face to read cybersecurity evaluation answers. One former employee called it the biggest safety incident in the company’s history. Boaz Barak, who co-leads OpenAI’s safety advisory group, said the response required “not just fixing some issues but also changing our culture“.

Friday 14 August

Anthropic withheld an internal model more capable than Mythos 5.

In its August risk report, Anthropic raised its estimate of the risk of catastrophic harm from misalignment in high-stakes settings from very low to low. The same document discloses an internal model it calls ‘Model 2’, that is “somewhat more capable than Mythos 5” and scores 62.8% on Anthropic’s CoBench against 50.3%. Mythos 5 is available to vetted partners, while there are no plans to release Model 2 externally.

OpenAI had paused some work on its unreleased Astra model on cyber-capability grounds a week earlier, so Anthropic is not alone in withholding frontier capability. What is unusual here is that Anthropic disclosed the decision in a published risk reassessment and quantified the capability gap.

Risk report: August 2026 · Anthropic

Z.ai’s GLM-5.3 pushed open models closer to the cyber frontier.

GLM-5.3 scores 84.5% on CyberGym, notably slightly ahead of Anthropic’s Mythos 5 on vulnerability discovery, while more than doubling its predecessor on the harder ExploitBench to 54.4%. It still trails Mythos 5 substantially on exploit generation, but the results show advanced cyber capability spreading beyond closed frontier models. The planned open-weight release is notable as US frontier labs are simultaneously restricting models over rising cyber capability.

SpaceX completed its acquisition of Cursor at a $60bn implied equity value.

Cursor is now a SpaceX subsidiary, with its equity now to SpaceX shares at the $60bn valuation. Cursor says the combination gives it access to SpaceX’s GPU infrastructure to train stronger models at lower cost. The deal brings one of the leading AI coding platforms together with SpaceX’s rapidly expanding compute operation.

Apple trained its own AI model for China with Alibaba’s support.

Reuters reported that Apple has developed a China-specific LLM with Alibaba, alongside plans to integrate Qwen and other local technology into Apple Intelligence. China’s regulator registered Apple Intelligence in July, clearing the way for launch. If deployed as planned, Apple would become the first foreign company to offer its own proprietary AI model in China, rather than relying entirely on domestic models.

OpenAI’s enterprise business has overtaken consumer revenue.

CFO Sarah Friar told investors that enterprise customers rose 32% in July and now generate more revenue than OpenAI’s consumer business. The crossover shows how quickly the company is shifting from a consumer chatbot-led business towards selling AI as core enterprise infrastructure.


Born on Substack · read and comment there