GLM-5.3 Crosses a Cyber Threshold as AI Beats the Best Stratego Player
GLM-5.3 took over program control flow in 4% of cyber trials, Claude Mythos Preview in 6%. AI also beat the best Stratego player.

Today was a quiet day for AI news, but one security result stands out: a model crossed a threshold that earlier models did not.
-
GLM-5.3 and Claude Mythos Preview in cyber tests: Anthropic’s Frontier Red Team ran several models on 100 randomly chosen tasks from its internal Binary Exploitation benchmark. GLM-5.3 fully hijacked control flow in 4% of trials, and Claude Mythos Preview did so in 6%. Claude Opus 4.6 and GLM-5.2 succeeded in none. Still unproven: the benchmark is internal, so the numbers cannot be checked from outside. Read the Anthropic quote
-
Stratego: An AI system beat the best Stratego player in history, and Ars Technica reports it did so on a budget. The key was a second neural network that guesses the identity of hidden pieces. Still unproven: the item covers one game, so it is unclear whether the approach works for other games with hidden information. Read the Ars Technica story
-
dots: OpenAI’s new proactive assistants keep working across complex projects and everyday tasks. Read the OpenAI announcement
-
Claude Code: Latent Space published an interview with Thariq Shihipar of Anthropic about the next era of Claude Code. The teaser lists Opus/Sonnet 5.5, Mods, Plugins and Projects as shipping. Read the interview

