Hacker Newsnew | past | comments | ask | show | jobs | submit | drbscl's commentslogin

So did I. Unfortunately it's even more verbose according to https://artificialanalysis.ai/models/claude-opus-5-5#token-u...

The problem with 5 wasn't just the verbosity, but its insane way of communicating. It had this bizarre circuitous sentence structure that always buried the lede, and always tried to be faux profound. I'm okay with verbosity if it's actually readable.

Unfortunately, they're full of it https://artificialanalysis.ai/models/claude-opus-5-5#token-u...

It does work out to be a similar cost per task though


You should probably look at the cost/score graph by effort level instead:

https://artificialanalysis.ai/models/claude-opus-5-5#intelli...

It is most of the pareto frontier.


Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort

so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)

You’re missing my point. I’m saying anthropic are exaggerating their results.

how are they exaggerating the results? Comparing the cost from that chart for 5 and 5.5 for medium-max effort paints a pretty clear picture:

         mean  median
 model
 5      4.135   4.245
 5.5    3.150   2.640

Again seeing how max is a clear outlier, the median cost saving is ~38%, not that far off from the proclaimed 40%.

Is verboseness the only measure of token efficiency towards overall task completion?

5.5 is higher for max effort, slightly higher for xhigh and lower for high, medium and low effort.

The biggest proportional difference seems to be at max (5.5 is 38% more) and at high (5.5 is 21% less).

I think most people run at high and xhigh. At xhigh it is close enough to be task dependent and I don't think most people will notice. At high effort I think it looks like it will be an improvement for most people.

5.5 Max should probably be compared to Fable - it performs a lot better than 5 Max.

https://artificialanalysis.ai/models/claude-opus-5-5?models=...


Disagree. Our internal company tests showed a cost per task drop from 0.35usd to 0.16usd . Opus 5low vs opus 5.5 low

Very fast you were.

Even created an account to tell us just that.

I don't think so, I typically use Opus 5 on High, and 5.5 scores lower on token use:

https://artificialanalysis.ai/models/claude-opus-5-5?models=...


My understanding of tech salaries in China is that they are pretty decent, but not as high as in SF; closer to typical European salaries.

Mostly due to lower cost of living; Shenzhen is way cheaper than SV


I seriously doubt salaries are included. It must be just the electricity and GPU costs.

In these metrics, yes. In the reported training budgets of anthropic/openai, who knows?

Omni usually means multimodality (in terms of input and/or output type, text, images, audio, etc)

The organisation has to still pay their suppliers, and have to comply with anti money laundering laws

I can pay cash at store, cinema, museum and not give my name. Why not at conference too? And it is not only this conference, all of them require it now.

Think about it :)


Yeah honestly fair point

Though I imagine it’s more about getting upfront payment for tickets to cover suppliers


I mean, DEF CON still allows cash at the door with no ID and manages to make it work.

Do they sell out of tickets within, literally literally, three seconds?

Hopefully this move to the bigger venue improves the situation.


> extremely right wing misogynism of DefCon.

Can you elaborate? I thought a lot of LGBT people attend regularly?


I was also wondering what they meant, tbh, but I have never been to DefCon, so I don't know.

One ex-colleague from my previous employer was there, though, and he absolutely loved it.


those can both be true at once though

Yes, but not commonly. Hence my request for elaboration

Breaking TOS isn't illegal per se. It just allows for denial of services, and may define terms by which the provider can reclaim costs.

>As we develop GLM, the model sometimes exhibits capabilities that surprise us, and even unsettle us.

Come on now

Also, why would they introduce themselves on their own blog?


No, just a push for regulatory capture, nothing more

I think technically it can, but the cameras are black and white, so it would be very subpar for AR out of the box.

There is some chatter about using the expansion port to add a 3rd party colour camera to get over this, but I wouldn't hold my breath for it.

Edit: I stand corrected, there is such a 3rd party camera linked on the page!


> There is some chatter about using the expansion port to add a 3rd party colour camera to get over this, but I wouldn't hold my breath for it.

The official product page above has a link to one you can already preorder: https://arcturus.vision/


Now I see that the page reveals this:

> Four high-resolution, monochrome cameras provide controller and headset tracking

(Emphasis mine)

So it's not really suitable for real AR out of the box. Thanks for the heads up!


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: