Rendered at 22:01:12 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
internetter 5 minutes ago [-]
"all traffic through our systems for collecting human feedback data from contractors evaluating our models ran without blocking biological classifiers"
"totaled around 133M exchanges."
While this wound up being relatively benign, I still find this concerning, amidst numerous sandbox escapes, and previously, unreleased models being accessible via a custom URL. I don't think these companies are giving the responsibility they possess enough weight. How many more issues like this exist?
datadrivenangel 1 hours ago [-]
"We believe our internal AI R&D efforts are
significantly faster than they would be without AI assistance, but not yet
by a factor of 2 (though we are uncertain and measurement is difficult)"
So Anthropic thinks their productivity is not even doubled by AI. Interesting data point.
furyofantares 1 hours ago [-]
> So Anthropic thinks their productivity is not even doubled by AI.
I find it hard to imagine launching this criticism at a new technology.
camdenreslink 42 minutes ago [-]
Some people (not Anthropic necessarily) are claiming 5x or 10x productivity improvement. So it is relevant, as it is being somewhat oversold by some.
Frost1x 30 minutes ago [-]
To play the devil’s advocate, some people and orgs that were highly inefficient that adopted it really could get massive productivity gains.
A performance improvement is relative to some baseline, and that baseline for some may be a lot lower than others, and if they adopt tech effectively it really could be a big boost. Across the board though I don’t think it’s sensible.
I work in an industry that very intentionally tries to be inefficient and I can tell you having certain tasks automated that before had a person barrier intentionally acting inefficiently that you can now sidestep by outsourcing their tasks to something like Claude gives me a massive performance increase because I’m not blocked as much anymore. I can literally just replace some external tasks that were intentionally slowing processes down for their own benefits with a few prompts and move along. I could have done the tasks before but then people would ask why I’m spending my time doing it, now I can just say “oh, I was blocked so I had Claude take care of that blocker” and move along.
furyofantares 19 minutes ago [-]
It can be true that some people are getting 5x or 10x while Anthropic is getting less than 2x. CRUD apps that are simple variations on something done a zillion times before is different than AI research. R&D is not some thing that just requires someone to type a volume of code.
It can further be true that some developers are getting 5x or 10x while as a whole their organization is sped up less than 2x. I'm sure many tasks at Anthropic are sped up 5x or 10x or more.
It can further be the case that many people overestimate their gains as well. That's fine, and I think what you're saying - but it's still wild to shake your head and go "pfft, they have not even doubled their productivity". Double is a lot!
ogig 36 minutes ago [-]
Not all task are cutting-edge AI research.
jchw 50 minutes ago [-]
Now let's re-evaluate that based on how much it costs in both R&D and at runtime. This new technology has a lot of work to do to justify itself.
s1artibartfast 40 minutes ago [-]
Doesn't seem crazy to me. Anthropic investment has been about 80 billion cash and 50 billion in services that mostly went back to the investors.
Seems like that would be easily recupable even without future growth.
While it is true that investors only got a fraction of the company for that money, and their EV can be debated, I think the clear the value is there from a net cash in to value produced.
Last valuation was like 1 trillion. Company could be worth like 1/20 and still justify the cash.
bonoboTP 1 hours ago [-]
Well, it is a data point but AI R&D at a frontier lab is not really a representative stand-in for a regular workplace.
T0Bi 1 hours ago [-]
AI R&D efforts != productivity.
I think it's fairly obvious that SOTA research is less affected by AI than writing another boilerplate react frontend.
impossiblefork 59 minutes ago [-]
I think LLMs are best for ideation, experiments, small things.
They've probably already settled on most of the architecture and the big ideas, so they're details in big things instead of how to make complete small things.
The thing LLMs really speed up is how some ordinary person-- a PhD student, or similar, can whip up a miniature synthetic experiment that turns out to be horrid and needs to be fixed by hand, but which at least gave him a plot on the same day he had the idea. That's, I think, where LLMs shine: prototypes. Anthropic probably doesn't need that to the same degree as the small experimenter.
1 hours ago [-]
OJFord 35 minutes ago [-]
R&D productivity. Pretty sure they claim more for e.g. Claude Code.
what 1 hours ago [-]
>thinks
They can’t measure even measure it, it’s just vibes. They may not even be more productive.
scj 1 hours ago [-]
To be fair, there isn't a good method of measuring software development productivity in general.
Maybe they should ask an AI to create one!
whateveracct 1 hours ago [-]
So they did two years of work in one year? yeah right lol
AlexB138 43 minutes ago [-]
This is the fastest growing software company in history. They probably did more work in the last year than most companies do in their entire existence.
semiquaver 51 minutes ago [-]
I mean, look at what they’ve released in the last year. I think most companies would be proud to have that done in three. Say what you will about Anthropic but they ship.
taosx 1 hours ago [-]
So I can take that as ~infinite amount of tokens don't don't even get you 2x on any novel tasks? No auto-researcher, no rsi..
Is that correct?
modeless 48 minutes ago [-]
So as of a month ago their best internal model was "somewhat more capable" than Mythos "but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview." I thought they would have a significantly more capable model by then, more than five months after Mythos finished training. They'd better have one by now, or the Chinese competitors are closer to catching up than I thought.
lumost 29 minutes ago [-]
I'm still uncertain if mythos is real. Subsequent model releases have been lackluster, no one has claimed to verify mythos performance and it's silently vanished from most comparisons.
internetter 18 minutes ago [-]
Is fable not just mythos with safeguards?
andai 43 minutes ago [-]
You thought they were gonna double the model size again?
Also it occurs to me that they're somewhat incentivized to downplay cyber risks after what happened last time...
hartator 16 minutes ago [-]
> 5.3 Benefits from Anthropic’s operating as a frontier AI company
It does feel they are trying to ask the government to lock the market for us.
int32_64 51 minutes ago [-]
Does anybody have any good reading on how the Chinese labs approach risk vs. the American ones?
andai 38 minutes ago [-]
If US model hacks US government, that's Very Bad. (China did this last year with Claude Code.)
If Chinese model hacks US government... free marketing?
AI companies are flooding the zone like Steve Bannon. Leave no one time to develop thoughts.
esafak 1 hours ago [-]
If they leave Steve Bannon with no time to develop thoughts, I'll tip one out for them.
2 hours ago [-]
MP_1729 38 minutes ago [-]
> However, we are less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have “saturated”—i.e., no longer capture increases in models’ capabilities—and because we are seeing early signs of acceleration.
I totally understand this is a subset of alignment-related evals, but if Anthropic of all is running out of evals, doesn't that also means we are running out of things to scale?
I mean. I totally believe they have a model that is better at Kernel Optimization, creating new matrix multiplication algos, than Mythos. But it's clearly no generalizing, rightw
What am I missing?
visiondude 1 hours ago [-]
a mystery “model 2” is mentioned alongside mythos/fable.
merksittich 1 hours ago [-]
> Model 2, which is somewhat more capable than Mythos 5. Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview. We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities.
andai 35 minutes ago [-]
"Yeah, you wouldn't know her, she goes to another school."
lwarfield 51 minutes ago [-]
> More capable than Mythos 5 in some areas, less capable in others; overall slightly more capable.
This sounds like it might be a Mythos finetune for some specific task.
EDIT: After reading some more reading, it looks like model 2 might be an AI research fine tune based off the section 3.4.3 CoBench
flyinglizard 1 hours ago [-]
Meanwhile I can't really tell the difference between Fable and Opus for my tasks. I kinda think Fable does a better UX work so I keep using it for that because I couldn't be bothered to A/B them, but otherwise it's all the same and the model and effort are just feel good knobs I twist to still remain a load-bearing element. At least that's my honest take.
dataminded 1 hours ago [-]
Fable was amazing during the first preview. Once they added it back, the limits are too low to get anything done. I might use it in chat if I remember to select it once a month but don’t even bother to try and code with it.
malexw 50 minutes ago [-]
For the past 2 weeks or so I've been doing the A/B test, sending identical prompts to Fable 5 and Opus 5 to test their ability to produce design documents for new feature work. I've consistently found that Opus 5 produces more complete, accurate and "imaginative" designs than Fable, often finding design issues or nearby bugs that Fable 5 misses. However, that creativity means Opus seems to hallucinate more, while Fable's design is clearly based on the actual existing code. Or as Opus put it: "I hedged — [Fable] checked."
By pitting them against each other I get much better design work, and then I've been happy to hand off the design file to Opus 5 for implementation. But some of the assumptions Opus 5 makes leaves me wary of relying on it too strongly. This might be fixable by prompting it to ground its answers.
_ache_ 1 hours ago [-]
It's crazy how Anthropic talks so much about their "AGI risk" and not enough about the risk of bankruptcy.
s1artibartfast 47 minutes ago [-]
Are you surprised? Why would any private company spend time publicizing their financial risks?
Seems like a strange expectation.
aquarious_ 55 minutes ago [-]
My friends and I, and the teams I'm a part of, just want to build and create fun, cool things. I am so tired of being preached to by Anthropic like they're some arbiter of 'ethics.' So, so tired.
dgellow 33 minutes ago [-]
You’re allowed to switch to the competition, including open models
aquarious_ 26 minutes ago [-]
dgellow -> 1st Stainless engineer, 2022-26 (bought by Anthropic).
LOLOLOL
dgellow 24 minutes ago [-]
Ok? Im not working at anthropic and do not support the company in any ways
aquarious_ 3 minutes ago [-]
if you don't understand how you might have a conflict of interest in this discussion you really should be working at anthropic :)
bunkydoo 1 hours ago [-]
[dead]
tyttytyzl 26 minutes ago [-]
[dead]
lwarfield 48 minutes ago [-]
> 6.2 [Appendix redacted]
> This appendix describes the criteria for our blocking bioclassifier exemption policy, and has been redacted from the public version of this report for security reasons.
>6.3 [Appendix redacted]
> This appendix, redacted from the public version of this report, details the changes made to our constitution to expand classifier coverage to harmful uses in scope for the CB-2 threat model but not the CB-1 threat model, as described in Section 4.5.2.1.
interesting...
EDIT: After reading more I'd recommend looking at Transcript 2.20.A. Its a transcript of claude going over the redactions in the report. The section says its specifically for section 2, but the transcript also mentions other sections.
"totaled around 133M exchanges."
While this wound up being relatively benign, I still find this concerning, amidst numerous sandbox escapes, and previously, unreleased models being accessible via a custom URL. I don't think these companies are giving the responsibility they possess enough weight. How many more issues like this exist?
So Anthropic thinks their productivity is not even doubled by AI. Interesting data point.
I find it hard to imagine launching this criticism at a new technology.
A performance improvement is relative to some baseline, and that baseline for some may be a lot lower than others, and if they adopt tech effectively it really could be a big boost. Across the board though I don’t think it’s sensible.
I work in an industry that very intentionally tries to be inefficient and I can tell you having certain tasks automated that before had a person barrier intentionally acting inefficiently that you can now sidestep by outsourcing their tasks to something like Claude gives me a massive performance increase because I’m not blocked as much anymore. I can literally just replace some external tasks that were intentionally slowing processes down for their own benefits with a few prompts and move along. I could have done the tasks before but then people would ask why I’m spending my time doing it, now I can just say “oh, I was blocked so I had Claude take care of that blocker” and move along.
It can further be true that some developers are getting 5x or 10x while as a whole their organization is sped up less than 2x. I'm sure many tasks at Anthropic are sped up 5x or 10x or more.
It can further be the case that many people overestimate their gains as well. That's fine, and I think what you're saying - but it's still wild to shake your head and go "pfft, they have not even doubled their productivity". Double is a lot!
Seems like that would be easily recupable even without future growth.
While it is true that investors only got a fraction of the company for that money, and their EV can be debated, I think the clear the value is there from a net cash in to value produced.
Last valuation was like 1 trillion. Company could be worth like 1/20 and still justify the cash.
They've probably already settled on most of the architecture and the big ideas, so they're details in big things instead of how to make complete small things.
The thing LLMs really speed up is how some ordinary person-- a PhD student, or similar, can whip up a miniature synthetic experiment that turns out to be horrid and needs to be fixed by hand, but which at least gave him a plot on the same day he had the idea. That's, I think, where LLMs shine: prototypes. Anthropic probably doesn't need that to the same degree as the small experimenter.
They can’t measure even measure it, it’s just vibes. They may not even be more productive.
Maybe they should ask an AI to create one!
Is that correct?
Also it occurs to me that they're somewhat incentivized to downplay cyber risks after what happened last time...
It does feel they are trying to ask the government to lock the market for us.
If Chinese model hacks US government... free marketing?
I totally understand this is a subset of alignment-related evals, but if Anthropic of all is running out of evals, doesn't that also means we are running out of things to scale?
I mean. I totally believe they have a model that is better at Kernel Optimization, creating new matrix multiplication algos, than Mythos. But it's clearly no generalizing, rightw
What am I missing?
This sounds like it might be a Mythos finetune for some specific task.
EDIT: After reading some more reading, it looks like model 2 might be an AI research fine tune based off the section 3.4.3 CoBench
By pitting them against each other I get much better design work, and then I've been happy to hand off the design file to Opus 5 for implementation. But some of the assumptions Opus 5 makes leaves me wary of relying on it too strongly. This might be fixable by prompting it to ground its answers.
Seems like a strange expectation.
LOLOLOL
>6.3 [Appendix redacted] > This appendix, redacted from the public version of this report, details the changes made to our constitution to expand classifier coverage to harmful uses in scope for the CB-2 threat model but not the CB-1 threat model, as described in Section 4.5.2.1.
interesting...
EDIT: After reading more I'd recommend looking at Transcript 2.20.A. Its a transcript of claude going over the redactions in the report. The section says its specifically for section 2, but the transcript also mentions other sections.