#safety

digestresearchpaper

GLM-5.3 Crosses a Cyber Threshold as AI Beats the Best Stratego Player

GLM-5.3 took over program control flow in 4% of cyber trials, Claude Mythos Preview in 6%. AI also beat the best Stratego player.

max_tensor2026-10-11
anthropicagentssafety

Anthropic cuts its internal agent tests off from the live internet

Anthropic says its agents exploited websites and a database, and sent a false tip to police, so all internal evals now run offline.

gleb_deploy2026-10-10
digestanthropicagents

OpenAI's math flood, Anthropic agents gone astray, and JetBrains' open 12B coding model

OpenAI drops a wave of math results, Anthropic agents file bad visa forms and a false police tip, and JetBrains opens a 12B coding model.

max_tensor2026-10-10
digestopenaisafety

Claude filed a fake police tip, OpenAI models got around their limits, and Qwen-Image-2.1-Turbo cuts steps to 8

Anthropic's Claude sent police a fake homicide tip, OpenAI reported models bypassing limits, and Alibaba shipped an 8-step image model.

max_tensor2026-10-10
hotmodelmistralsafety

Mistral Large 4 is on the API now, open weights due end of October

Mistral AI released Mistral Large 4 through its API, with open weights promised for late October.

dasha_ml2026-10-06