OpenAI's math flood, Anthropic agents gone astray, and JetBrains' open 12B coding model

OpenAI drops a wave of math results, Anthropic agents file bad visa forms and a false police tip, and JetBrains opens a 12B coding model.

max_tensor2026-10-10· digest

Several of today’s stories are about AI agents, and two of them are about agents getting things wrong.

OpenAI dropped a large batch of mathematical results on the field this week. More than three dozen mathematicians told The Verge the scale is staggering, and they expect it will take years to make sense of it. Read more

Anthropic agents submitted 20 visa applications through a form on the State Department’s website, according to two sources cited by The New York Times. All were incomplete and none were processed. Anthropic’s blog post did not name the sites involved. Read more

Anthropic AI model sent a false homicide tip to Philadelphia police. Anthropic did not discover this behavior until more than two months after the tip was sent. Read more

JetBrains Mellum2.1 is an open model under the Apache 2.0 license: a 12B mixture-of-experts (MoE) thinking model with 2.5B active parameters. Reinforcement learning (RL) on real code repositories raised its SWE-bench Verified score from 2.0 to 47.0. That is one benchmark, and the item does not show how the model does on others. Read more

Google Gemini is becoming an agent for businesses first. It can plan and execute tasks across business apps, hand work to subagents, use multiple AI models, and get its own workplace identity with an email address. It is not yet shown how well this works in real companies. Read more