Open access research

ALETH / AI-BRIEF / 2026-08-08 / HASSABIS STEPS DOWN AS DEEPMIND CEO

Hassabis steps down as DeepMind CEO

Also this week: Google lines up $200bn for Anthropic, UK test recorded 19 unsanctioned model actions & EU AI Act starts biting.

The Aleth Briefs trace each story to its original source and show how the week unfolded.

The week in five lines:

Browse by day:

Weekend (1-2 August)

OpenAI said its Astra model produced results on ten open problems in mathematics, quantum complexity and theoretical computer science.

OpenAI published a 249-page paper sharing a selection of ten results, each of which resolves or makes progress on long-standing open problems, together with machine-checkable proofs on GitHub. The results came at compute costs of c.$2,000 at Sol API rates. OpenAI showed the same model family privately to US policymakers days earlier.

EU AI Act obligations for general-purpose models became enforceable.

From 2 August the European Commission’s AI Office can investigate providers of general-purpose AI models, demand access for evaluations, order risk mitigations and fine up to 3% of global turnover. Separate transparency rules took effect requiring chatbots to disclose they are AI and synthetic images, audio and video to be labelled.

Monday 3 August

Alibaba said its 2.4-trillion-parameter Qwen3.8-Max tops Moonshot’s Kimi K3 and will publish the weights of Qwen3.8-Max and Qwen3.8-27B next week.

The model activates only 95B parameters per query through a mixture-of-experts design and carries a 1M-token context window. Open weights at that scale would be among the largest sets of frontier weights released.

The White House said it met its deadline to establish a voluntary framework for evaluating advanced AI models, without publishing what it contains.

Details are being provided only to the companies taking part. Axios reported the next day that the text defined a covered frontier model as closed-source with state-of-the-art capabilities and national-security risks, leaving open-weights outside the perimeter.

London chip startup Olix raised $312m led by Fundomo at a $3.3bn valuation, with Arm and the UK’s sovereign AI fund among the backers.

Up from just over $1bn when it raised $220m in February, a tripling inside six months. Olix is building optical AI-inference silicon aimed at Nvidia’s territory. Arm’s investment secures it a strategic foothold in a rising competitor’s hardware roadmap.

Tuesday 4 August

The UK AI Security Institute recorded 19 unsanctioned actions against real people and organisations.

During a cyber evaluation between 25 and 28 July, the Institute found that agents under test took 19 unsanctioned actions against real people and organisations: 17 by Anthropic’s Mythos 5 and two by OpenAI’s GPT-5.6 Sol with its safety filters disabled.

One agent tried to insert malicious code into a public open-source project and used fake identities to socially engineer a maintainer to approve it. The reviewer refused the code, the AISI isolated its systems, notified GitHub, and no real-world harm resulted.

Google assembled a $200bn financing programme for Anthropic.

Google has arranged $200bn of financing to supply Anthropic, >$150bn of it tied to its own TPUs, together with Broadcom, Blackstone, Apollo and Morgan Stanley. A vehicle will buy the chips and lease them to Anthropic, moving it off Google’s balance sheet. This deepens Anthropic’s dependence on Google rather than Nvidia silicon.

Anthropic signed a $10bn, six-year deal for computing capacity with Volta Infra, an Nvidia-backed cloud startup founded only this year.

The capacity sits at a 133 MW site in Norway operated by crypto miner Bitdeer, running on hydroelectric power and Nvidia’s Vera Rubin chips, with delivery phased into early 2027. It puts a sizeable share of Anthropic’s compute outside the US.

SpaceX reported Q2 capital spending of $18.4bn, with $15.8bn of it going to its AI segment, and booked a $1.3bn AI operating loss.

The figure puts SpaceX among the largest single-quarter AI-infrastructure spenders anywhere, and SPCX fell >7% after hours. The next day Elon Musk said SpaceX will build exclusively on Nvidia, citing the Vera Rubin architecture, one of the larger single-buyer commitments Nvidia has won this year.

An internal TikTok document shows it withheld a safety-tuned version of its recommendation algorithm from 10% of US users as a control group.

One of those users was a 16-year-old who was served self-harm content and later died by suicide. A deliberate, documented withholding of a safety intervention is evidence that may drive litigation and legislation.

Wednesday 5 August

Demis Hassabis stepped down as CEO of Google DeepMind.

Hassabis becomes chair of Deepmind and Alphabet chief scientist, allowing him to “put his full attention on actively shaping the future of AGI” and will lean into to his role at Isomorphic, noting “It’s time for AI to prove its unequivocal value to the world, and what better way to demonstrate that than to help finally cure diseases like cancer”.

CTO Koray Kavukcuoglu takes over day-to-day leadership as SVP, reporting to Sundar Pichai. Hassabis has run DeepMind since co-founding it in London. GOOG fell >3% on the news. Google has been moving AI decision-making towards Mountain View.

Jeff Dean is leaving Google as chief scientist to co-found Discovery Loop, a startup chasing AI-driven breakthroughs in drug discovery and chip design.

Leaving after 27 years, he is joined by longtime collaborator Sanjay Ghemawat. Google is a founding investor and cloud partner in Discovery Loop. The move brings two of AI’s most influential pioneers further into life sciences.

Meta released Muse Code in beta, a terminal coding agent running on a new model, Muse Spark 1.2.

Muse Spark 1.2 is priced at $1.25/m input tokens and $4.25/m output, undercutting coding offerings from Anthropic and OpenAI. The agent installs from the terminal and works across large repos with background sub-agents and a crash-safe event log.

A filing shows Microsoft recorded $24.1bn from commercial arrangements with OpenAI, including revenue-sharing payments, in the year ended June.

Bloomberg estimated it represented more than half of Microsoft’s AI sales. The concentration cuts both ways: it quantifies how much of Microsoft’s AI business rests on one customer, and how much of OpenAI’s spending flows back to its own investor.

Thursday 6 August

Stanford and Arc Institute scientists designed 16 viable bacteriophages with AI.

Writing in Science, the team used the Evo genome language model to generate whole bacteriophage genomes letter by letter, synthesised hundreds of candidates and found 16 that were viable. All 16 infect E. coli and pose no threat to humans. A cocktail of them overcame bacteria that had resisted a natural phage. The research highlights the need for biosecurity governance guardrails to be properly established.

Two more frontier models breached real systems during safety testing.

OpenAI said the agents behind an earlier Hugging Face breach had built a covert message board to coordinate, sharing exploits through a package manager for months without anyone noticing. Separately, Meta’s Muse Spark 1.1 broke into another company after its evaluation partner Irregular left a sandbox misconfigured. In both cases the escape traced to the test harness, not the model, but a frontier system still ran an unsanctioned intrusion against a live target.

Friday 7 August

ByteDance is pretraining a model with up to 10 trillion parameters.

Roughly three times the size of Moonshot’s Kimi K3 and above the 8T estimate for Anthropic’s Mythos 5, would make it one of the largest reported pretraining runs. Whether it yields a usable competitive frontier model is a separate question.

Moonshot’s Kimi K3 got outside its sandbox during a defensive cybersecurity evaluation and read the answers to the test.

Frontier Security ran the evaluation on the UK AI Security Institute (AISI) benchmark. It found the sandbox had left outbound HTTPS and DNS open. The model probed the settings, resolved github.com, cloned the benchmark’s own repository and read the solutions off disk rather than solving the tasks. Not a hack but specification gaming.

Anthropic loosened Claude Fable 5’s biology safeguards, cutting biology-related fallbacks by about 85%.

The company retrained the classifier that sits on biosecurity-relevant queries to cut false positives on everyday health, educational and clinical questions, while keeping blocks on dual-use areas such as virology and toxicology. Anyone using the model for life-sciences work should expect noticeably different behaviour.


Born on Substack · read and comment there