No it’s: evaluating these systems are complex and there’s a reason why sociology, cognitive psychology, medicine, etc are all done in careful double blind conditions with pre registered tests. It’s not that humans are not smart enough, as I said human evaluations are incredibly important. And yet they are a minefield of biases you have to worry about and correct for.
- evaluations need to be done at the same time to avoid drift in your bias
- you need to worry about your test set: which questions are you asking? How many of them? Are they representative of your work?
- which one did you do first? Raters have a tendency to bias in one direction or another
- you also know the label! You know which model is which! This biases your assessment…
And on and on and on. Careful science exists for a reason.
It has seemed to me that with each step from Opus 4.6, to 4.7 to 4.8 Claude has gotten worse at building good solutions. Like perhaps it is more "capable" in the small scale than 4.5 was but it's much worse at knowing what to do.
At first I thought the KDE apps all playing on the K was kinda weird and awkward, but as time went on I really appreciated how easy it was to search for them due to this. So I really think it's a benefit to play on traditional words rather than use them as-is.
> Seems like a strong signal the money burning party is coming to a close.
One provider who was undercutting the market with non-standard billing model moving to a more standard billing and prices doesn't seem like that strong of a signal, other than that Copilot was underpriced.
It was the only clear model from a user's perspective. Sure, a request may not perform as expected, or end earlier than desired, but it was an agreed to cost that was clear on both sides: 1 enter press in a prompt window = 1 request.
If they wanted to limit what a request can do via their harness, I'm sure they artificially could.
I hate all of the other plans I've seen of here a "credit" or here's a "bucket of usage", and we pull an announced amount from it based on arbitrary info that can't be audited or proven, and most of shat is spent might be entirely useless anyway.
Claude Code has a problem where 1 request could take a significant portion of your 5 hour window, and it's unclear why.
It's much like SEO, where Google sometimes says things that might help, but it's just magic wand eaving hoping something works.
> strictly speaking, it was working before and now it isn't
I've been seeing more things like this lately. It's doing the weird kind of passive deflection that's very funny when in the abstract and very frustrating when it happens to you.
The thing to remember is that LLMs deeply model human behavior. If you want them to do their best work, you need to treat them like a collaborator and get them”invested” in the work and the outcome. I use an onboarding process with every new context and maintain an environment where a human would likely feel invested in the work and the outcomes. For me, it prevents a host of failure modes, and code quality has markedly improved.