The AI Front Page

Entity Edition

Anthropic

30 stories from 11 sources across 11 topics.

Stories

30

Sources

11

Topics

11

Mozilla.ai Blog / 9:38 AM

Who Cares About LLM costs?

Why would you worry about them? LLMs promise something close to infinite capability, and worrying about the meter feels like someone else's problem. Ship the feature. Let the model think as long as it needs to, right? Right. Until the bill arrives. The shock Picture it: your

ReadSource

The Decoder / 9:03 AM

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings

OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only with its own API features instead of the official test setup, where the model landed at 7.8 percent. ARC Prize claims its test environment is provider-neutral, but may have used an outdated API that skewed the comparison with Opus 5. The article OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings appeared first on The Decoder .

ReadSource

Simon Willison LLMs / 6:18 PM

Quoting Matthew Green

Right now we’re in the midst of a historic transition from traditional public-key algorithms based on EC-based cryptography and RSA, moving over to new post-quantum algorithms based on novel problems. This is why there are so many standards like HAWK being considered. If there was ever a perfect time for a massive new public cryptanalysis capability to come on line, we’re in it. So unless AIs succeed in undermining all of our hard problems altogether (or we live in Impagliazzo’s Minicrypt ) then this could not be a better time for AI to get good at cryptanalysis. In the best case, the result is that we gain real confidence in the problems we’ve identified, and the cryptanalysis literature gets a lot more robust. Hopefully. — Matthew Green , on Anthropic's recent cryptography work Tags: cryptography , ai , generative-ai , llms , anthropic , claude , ai-security-research , claude-mythos-fable

ReadSource

The Decoder / 7:40 AM

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only through its own API with retained reasoning and context compaction. In the official test environment, the model managed just 7.8 percent. Opus 5 hit its 30.2 percent without such aids. The article OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness appeared first on The Decoder .

ReadSource

Simon Willison LLMs / 10:45 PM

Discovering cryptographic weaknesses with Claude

Discovering cryptographic weaknesses with Claude The best part of this article (here's the repo ) about how Anthropic researchers used Claude Mythos to find mathematical flaws in both HAWK and a weaker version of AES ("neither of these results has a practical impact on today’s computer systems") is the prompts that they shared, spelling mistakes included: the models tend to think it is impossible to solve so they don't try they need a good amount of prompting. why not do aes-128 r7? the whole point is to find something better than existing approaches. no again the goal is that we have highly inteligent model as good top researcher, we want to find new attacks no we don't want to change the targets [...] agian we need to find something that worth publishing again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings. Mythos Preview worked for 60 hours in total (~$100,000 in estimated API cost) and the main human interventions were to encourage it not to give up and "find something that worth publishing". The paper CryptanalysisBench: Can LLMs do Cryptanalysis? describes the new eval that was created as part of this work, in partnership with ETH Zurich, Tel Aviv University, and University of Haifa. Via Hacker News Tags: ai , prompt-engineering , generative-ai , llms , anthropic , claude , ai-security-research , claude-mythos-fable

ReadSource

The Verge AI / 7:33 PM

AI’s finally expensive enough to make Wall Street nervous

It's earnings season, and investors got an unpleasant surprise from Google: an increase on its spending estimate, to as much as $205 billion - from the last quarter's projection of up to $190 billion. Even the lower end of Google's new projected range - $195 billion - is much more than the company had previously […]

ReadSource

The Decoder / 10:04 AM

Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price

Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work at half of Fable 5's token rates. On ARC-AGI-3, a benchmark for novel problem-solving, Opus 5 hits 30.2 percent, nearly four times higher than GPT-5.6 Sol. The article Anthropic's Claude Opus 5 delivers near-Fable 5 performance at half the token price appeared first on The Decoder .

ReadSource

The Decoder / 9:31 AM

Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks

Anthropic's Claude Opus 5 leads the Artificial Analysis Intelligence Index with 61 points, edging out Claude Fable 5 and GPT-5.6 Sol. The model scores highest in analytical quality and coding, and costs up to half as much as Fable 5 at lower reasoning tiers. But the race at the top remains close. The article Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks appeared first on The Decoder .

ReadSource

Latest story in this edition: 1:59 PM

Back to front page