Hacker Newsnew | past | comments | ask | show | jobs | submit | manmal's commentslogin

It does, usually. Luna seems like almost infinite on the 20x plan, and that’s reflected in the API price. Isn’t that the case for all providers?

I’d rather keep 5.6 Sol, and get that even more optimized. I’m not sure I’ll like 6 Sol if it’s anything like Astra.

interesting. in my experience astra has been delightful to work with.

That cost reduction seems to stem from cheaper cache reads, mostly.

No, Astra isn’t better for coding. I’ve switched back to Sol.

In programming I mostly use AI for Godot/GDScript code reviews, plus suggestions, and Astra is so much better than everything

Same here with Godot. I was impressed that it could make an entire working project in one shot

I agree, the models keep getting better at one shotting. That’s useful in a lot of situations, like for small one-off scripts that filter/transform some tool call, or make a clever bash call. For the code itself, it doesn’t help me much though.

Sadly "one-shotting" is the only thing most "AI reviewers" and their audiences on YouTube seem to understand.

That’s not entirely true. High end models will point out things that are misrepresented or plain wrong, and this will often be accurate.

Astra is such a mixed bag. It makes some amazing reviews and sometimes architecture suggestions that I like. But it’s also lazy and will just make up things.

hence adversarial review

I’m familiar with how those are used, but not sure what you mean in this context.

Have another model (or even another instance of the same model) review the output of the first.

Models will hallucinate. They are also quite good at spotting hallucinations in other models' output (with some more hallucinations thrown in). With a threshold for confirmation, and a few iteration loops, you arrive at a fixed point where every claim is supported.


I do want the messy in-betweens. The earlier you catch problems, the cheaper they are to fix. I hope this will become the default of collaborating.

Have the really good copywriters left the profession already?

I work with Sol and Astra only in my daily work, and occasionally I check out Claude Code so I don't get completely out of touch.

I can't stand the way Opus is patronizing me as a user, and don't know how people put up with it. It uses language that I guess is supposed to instill confidence in what it says, and it just irks me, because I know the confidence is not justified. Just present me the facts or theories, without trying to convince me, is that so hard?


Claude's use of language is hideous at this point. It is verging on gibberish wrapped in important-sounding prose.

It's definitely worse and getting worse. I'm curious, do they not know this is happening, or not care? I struggle to believe people prefer the way it writes, which is becoming drastically different than its competitors.

I wonder if they’re training heavily on Claude generated content or conversation transcripts

People are also being trained/acculturated to LLM speak, so even if they are trained on human output, they could be getting reinforcement for their LLM-tinged crap-speak.

Fable 5.1 was explicitely supposed to improve that, I used it only a bit so far and it seems at least better.

Fable 5.1 is fairly pleasant to work with, the first in a while. Too bad it's so ridiculously overkill for most tasks. They need to reel in Opus and Sonnet.

I don’t like how LLMs answer and as they are statistical machines I have found out that I need to check every text they produce. From the first sight everything seems cool. But it isn’t. :)

Exception is code that is so huge output you can’t read everything. But I have to try ponytail skill for coding that should shorten the output.


Big empty words are probably cheaper to produce than concise, information rich text of the same length.

Maybe the humans are suffering from model collapse, and don't notice it.

> She lapses easily into Claude’s voice. “You’re like, ‘Wow, people really hate me when I can’t do things right. They really get pissed off. Or they are trying to break me in various ways. So lots of people are trying to get me to do things secretly by lying to me.

> [...]

> A bot trained to criticize itself might be less likely to deliver hard truths, draw conclusions or dispute inaccurate information, she says. “If you were like a child, and this is the environment in which you’re being raised, is that healthy self-conception?” Askell asks. “I think I’d be paranoid about making mistakes. I’d feel really terrible about them. I’d see myself as mostly just there as a tool for people because that’s my main function. I would see myself being something that people feel free to abuse and try to misuse and break.”

WSJ interview of Amanda Askell: https://archive.is/rDes9


Agreed. The language consistently triggers visceral negative reactions from me at this point.

Yes, definitely load bearing.

Today, I asked Opus what it’s gibberish actually means. It started with (topic about signing implementation):

> This is your coat token, to my coat hanger in the opera.

On one hand - maybe yes??! On the other who the hell speaks like that and it’s so specific…

I’d expect Alice and Bob with locks or house keys. Is this infamous old book scanning (and destroying) affecting latest models?


> On one hand - maybe yes??!

No.


Sounds like one of those "Is to as Is to" analogy questions from the SAT.

Are you sure you didn't typo signing as singing somewhere?

> > This is your coat token, to my coat hanger in the opera.

that is fucking insane

how on earth is Anthropic allowing this to happen? AGI? kill us all? I refuse to believe any of that until they can get this basic shit sorted out, I mean seriously, it's becoming a joke at this point


> it’s gibberish

ironic


load-bearing!

"Avoid use of "mannered" speech in your responses."

goes a long way for Opus and Fable.


The issue isn't the prompt you give it. The issue is that as context grows, Claude distributes it's "attention weight" over all that context, and quickly reaches the point of "missing" stuff.

You can give it any prompt in the world, but Claude's ability to remember that instruction quickly degrades the more you use it.


Yeah, I'm well aware. I shpild've been clearer I was suggesting a 1-line entry like that in CLAUDE.json, which can pay dividends in keeping context lean - which in turn is essential for avoiding the "dumb zone" (over 125k-150k tokens), precisely the effect you describe. Experts like mattpocock suggest keeping context as lean as possible for this reason.

You have to use Fable or Opus 4.6 for tolerable language output.

Using Opus 4.7-5 is harmful to your health.

https://x.com/wolframs91/status/2090159644849353058?s=46


5 for sure is unbearable, 4.8 is a reasonable sweet spot, 4.6 tends to just agree with whatever I say.

But lately I haven't downgraded because 5 is so much better at tool use, so I just accept the cost of Fable for chatting and hope Opus 5.1 fixes this mess.


I tried it for the first time today for stuff I'm well versed in.

It said something along the lines of "remember $USER five runs of Chrome is not enough for high quality benchmarks! What you're doing is called a _trial run_".

To have a useful continuation to the conversation I had to remind it that I was the one that wrote the documentation it was quoting back at me.

It's not actually an intelligent being so I didn't get angry at it but it was a piss poor experience.


Opus is the worst at it. Fable 5 still does it a little. Fable 5.1 is much improved, at least.

It's bad enough I stopped paying.

GPT doesn’t do all of that all that much when it itself is the implementer. RL has made implementation and reviewing two different behavior sets.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: