Claude Executive: Damn, we hardly know who our users are, can't we just force them to say their full name or ban them?
Claude Product Manager: No, that'll piss people off too much, and sadly we can't just ask for ID either...
Claude Executive: There must be some way we can force people to link their government IDs with our platform so our analytics get better and more accurate?
Claude Product Manager: We could limit the platform to 18+ and use "Age Verification" as the reason for people to hand over IDs, seems other platforms had success with this approach
Claude Executive: And we hardly have any users younger than 18 anyway, go for it!
If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).
Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Key for me is to "group" cables. For example, I have a bag of USB-C cables, a bag of USB-A cables, etc.
Grouping them is key to deduplicating.
It's easy to look at a single legacy USB A-to-B "printer" cable in isolation and think "I might need this someday!" Because you really might need it someday. However, if you group them you might see that you have ten of them. And then you can get rid of... maybe 8 of them.
I also (mostly) put individual cables into baggies. You can get clear 2mil generic ziploc style baggies for super cheap on Amazon or elsewhere. $15 for 200 or something. Prvents tangles and way less effort than wrapping or tying them.
It's so refreshing to see DeepSeek's tech report[1] full of juicy details; meanwhile, something like Fable's system card[2] is like 70% "safety", 10% "model welfare" to make sure little Claude isn't distressed, and 20% benchmark numbers.
> One person speculated that the timing may have something to do with Mullenweg’s annual trip to the Burning Man festival, after which he tends to return “with ideas.”
This line is so perfectly on the nose it could be in a McSweeney's article and yet here it is in reality. We are truly on the weirdest timeline.
> “And the biggest news: I won't be able to make the ELT tomorrow,” Matt wrote in an Announcement Slack. He continued, writing that Mark Davies, current Automattic Chief Financial Officer, has “conspired” with Automattic Board of Directors members Ann Dunwoody, Toni Schneider, and Sue Decker “behind my back and they voted to put me on a paid leave of absence. I voted against that.”
If you’re doing this type of accusation in the company wide announcements channel, odds are you’re never coming back. A board coup rarely happens without a good reason, but a “leave of absence” isn’t a firing. It’s still pretty bad, but it’s not a firing. However, crashing out and waging war on the board will lead to that outcome.
Software development untethered from the practical realities of the customer / user is what drives people insane.
When developers are required to interact with the customer on a regular basis, the freewheeling effects described in this article are damped massively.
The potential for insanity goes off the charts when the development team is siloed away in solitary confinement and the only interactions with the client occur via some prison guard known as "project manager" sliding notes under the door.
Working with the customer sometimes sucks. Just like exercise and eating vegetables sometimes suck. It's a temporary unhappiness that keeps us grounded in reality.
I'm imagining a future where a bunch of bizarre laws interact oddly (as they do), and now we've got websites with unnecessary nudity pasted in the corner.
"Oh, those? Those are just compliance tits. Ignore those. It's just a thing that came a few years after we finally got rid of the cookie banners. The companies wanted certain protections awarded only to 18+ sites. But you can't just declare yourself an 18+ site, so some sites post the most minimal amount of imagery that constitutes erotic nudity. That's why Google's graphic for the past few months has just been that one with the two dots in the middle of the o's."
Numbers from GPT Astra
- Shopify has 3000 engineers as of 2026
- Google Chrome when released in 2008 conservatively had ~ 60 engineers.
- GTA 5 in its credits had 150 software engineers. Surprising even to me who has had many an experience of being in a bloated FAANG team, this 150 includes GTA Online!
In a sane society, Shopify’s opinion on anything engineering related would be thrown into rubbish because they seem to have managed to complicate a simple app into requiring thousands of engineers and now maybe millions in cloud spending to Frontier labs. This is unfortunately not an isolated case, Spotify for one has the same issue, idk what “engineering” Spotify is doing, it’s the worst app I’ve used in my life.
There’s a standard for paper sizes called the ISO 216 international standard, aka A-series paper sizes. You may be familiar with it outside of the American 8.5”x11” letter size. The simplicity, for example, where an A3 sheet is exactly twice the area of A4. One A3 sheet folded in half becomes two A4 sheets. This is intentional: when you fold or divide an A-series sheet exactly in half, the resulting sheets have the same proportions as the prior. The aspect ratio is preserved across all A-series paper sizes.
The iPhone Duo is designed with this in mind. The aspect ratio any way you view it, folded or open, is preserved. Apple design knows a thing or two.
As I also said on Twitter - it really amazes me how fearless Deepseek are. Every single model release is packed with new and crazy clever ideas and somehow, they always commit to training them at near frontier scale.
I know everybody wants the tell all story of the clever ideas that were developed over the last ~3 years at Anthropic and OpenAI, but what I really want to thumb through is DeepSeek's notebook of "brilliant but didn't quite make the cut" ideas.
They must be trying some truely bonkers stuff to be able to land this much architecture novelty in their full releases.
If you put every company that needs/has an app on a spectrum, there is a line somewhere that roughly divides them into two groups: where Electron/React Native/etc. makes sense or not. It's just a normal engineering decision: solving problems given limited resources. Companies have different problems and different resources.
I think people in the tech community have probably also noticed that it's rather popular to have an absolute opinion on the goodness or badness of these tools. There's some magical thinking borne from ignorance that everyone just ought to go native or that React Native is the best thing ever to be used everywhere or that AI makes this line disappear entirely.
I think these takes serve little value and distract from what’s interesting, and what the subtitle to this article says: that this line is moving due to AI. And I think that’s probably right.
Better driver education and higher test standards would save lives too. So would banning alcohol. None of these solutions, including automated cars, will be mandated any time soon. It's not enough to have data, you also have to have societal buy-in.
Interesting that Waymo chooses to compare accident rates with the average driver vs the rideshare drivers their cars replace. The numbers wouldn't look as impressive because rideshare drivers are involved in fewer serious accidents than the average person.
Waymo says it has 170 million miles and I found 2 fatal accidents it was involved in (fault doesn't matter as fault is not included in the other stat either). Looks like the average is about 1.3 fatalities per 100 million miles, and about 1 per 100 million miles for CDLs. This doesn't seem that impressive.
If we want adoption, we would need it be placed in individual vehicles and it should look at removing driver liability. People will use it when they are drunk if it means not going to jail and it's readily available.
As a biker, I've had human drivers intentionally try to drive me off the road or hit me. Neither driver education nor higher test standards will solve that problem.
I say this as someone with tens of thousands in audio equipment, as someone who can hear the compression warble in the hihats and the cymbals - audio quality with earbuds doesn't actually matter. It's like the argument for audio quality in cars - you're listening in a hostile environment, and convenience wins over fidelity. I remember walking to work one time with over-the-ear Sennheisers attached to my iPod and it felt ridiculous, and unsafe.
I finally capitulated to earbuds when they started giving them away with the Pixel phones. The "find my" feature basically justifies the pro-prices, to me, because my biggest complaint about bluetooth earbuds in my own use case was how easily they're lost. I've lost and recovered my Pixel Buds Pro 2 a couple of times now. When traveling a lot by train and plane, the convenience of pocket-sized noise-cancelling earbuds has become worth it for me.
But when I'm on meetings, I'm wired in. During the pandemic, I lost track of how many times my colleagues' Airpods would sputter out and die in the middle of a client meeting. Then they'd switch to wired headphones and their Macbook microphone. They sounded better, they could hear better, and there was no battery power sword of Damocles anymore.
I work at an intersection of tech, applied research, and science.
Something I’ve noticed in collaboration that does occur is an increased confidence in people outside their domains to say things with conviction. I have people who have limited experience with software pushing out layers and layers of abstracted code that’s fairly sophisticated but often misguided in intent who will say what they’re doing is correct, with conviction.
I also hear a lot more questioning people in their domains and challenging opinions, then hearing what I can only imagine are fragmented pieces of conversations they had with an LLM thinking through some argument. Then there’s silence when you discuss shortcomings, then they come back later with their memorized fragments of what you said, combined with memorized fragments of the LLM response to the argument.
It’s occurring, a lot more. People are treating their LLMs in collaboration as a source of truth and using then to focus on their specific path or goals they think or have bias towards going down, vs just opening discussing things, considering tradeoffs from experts multiple disciplines weigh in on and then taking an approach that everyone finds most agreeable.
It’s making me want to be a lot less collaborative with such individuals. I don’t want to sit around and refute Claude text outputs all day.
I think it's a useful analogy to compare OpenAI to a human collaborator. These researchers willingly collaborated with an OpenAI model, giving it ideas, and OpenAI provided useful replies. Then, OpenAI goes ahead and publishes work along the lines of this collaboration, without attributing the researchers. If OpenAI was in fact a human researcher, this would be highly unethical.
Now, OpenAI is claiming that the model it used to generate the result was not trained on these collaborative communications with the researcher. This is a technical argument that is impossible to verify as an OpenAI outsider, and probably difficult to verify even for internal OpenAI employees. Provenance is hard to track - you would hope OpenAI has very good tools for this, but a full data trail of all inputs is difficult to trace through.
Another interesting thing to consider is if instead of OpenAI doing this, it was another research mathematician A using an OpenAI model just like the internal group at OpenAI did to publish these results. What if the model A used was trained with unpublished communications with other researchers B who were working on the same problem? Should researcher A technically include B as coauthors? How could they do this when they do not know the communications B had with OpenAI? In this scenario OpenAI, as a middle man, has laundered information from B to A, stripping out attribution. A scooped B without even knowing it!
> The motion says the PlayStation Terms of Service put a binding arbitration agreement and a class action waiver in Section 14, and quotes the opt-out clause: ...
> The clause requires a user who does not wish to be bound to notify Sony in writing within 30 days of accepting the agreement.
Binding arbitration on individuals should be illegal, full stop. The only use case is taking away people's rights as consumers and workers. Or dodging responsibility for deadly mistakes like the Disney+ incident.
This "opt out" mechanism is made to let Sony lawyers argue that accepting it was your choice so it can't be struck down as forced, even if 99% of users have no idea it exists, by design. Evil all the way down.
Well, a search for "youtube acoustic fire extinguisher" indicates it has already been invented a few times by people all over the world. The most interesting video is https://www.youtube.com/watch?v=ZvnCQg4w4o8
Unfortunately, this is for the best. WordPress powers a large percentage of the internet, and I'm not sure what's going on with Matt (who I've always liked) but something turned the past few years.
He's been behind unforced error after unforced error as CEO, and I hope a break helps him gain some perspective (although I don't see him being brought back).
WordPress has been his life for 23 years. Doing the same thing for 23 years likely makes it difficult to have a proper perspective.
We did the same thing - had 90% of it overnight. Then spent a few days in the background tweaking for polish.
Our app is smaller, and has about 15-20 screens. I started at about 12:30am by giving codex a goal and it inventoried every screen based on the react native code, then created android and iOS directories, used maestro (I had already set up this tooling for a previous personal app build a few weeks prior), and had the whole thing working in android and iOS in the morning. Took it about 6 hours while I slept.
The app is way smaller, launches instantly, and the android app is (supposedly) native looking. I say supposedly because I don't use android phones. But it's using Jetpack Compose and Kotlin.
And I don't know Swift or Kotlin. I honestly don't see the point of React Native anymore. I know Expo is doing very cool agentic stuff, but I'm just not sure why I'd need any of it when I can write a native app.
To show you how strong Stockfish 19 is compared to 18, I used to lose 100% against 18, and I now lose 100% against 19, probably faster. Time to fire up En Croissant and see :)
This is a man who, to be fair to him, has been beneficially bullish and iconoclastic in the WP community's favour for decades, as well as writing a big chunk of what was "early modern" WP. Before he came unglued.
The fact that he looked, sounded, talked the way he did, was involved, put his own money where his mouth is, and was press-available to the extent he was is a lot of why WordPress was ever taken seriously. Given the level of control he has, it was used pretty judiciously for the longest time.
I am of the opinion that his broad charge against WP Engine was valid; I think he handled it insanely.
But as I say, he has come unglued. I'd be surprised if he returns to the job. I wish him well as I think anyone who has made money because of WP should. Things have gone wrong but there was a long run of things going remarkably well.
I do think this was overdue, for everyone including Matt, even though he evidently cannot see it. I hope, but am not hopeful as it were, that he sees that this is a message from the world to change track. I would instead expect to see a bit of revenge.
These threads are often amusing with people obliviously asking: why doesn't Apple make exactly what I want? And then stating with certainty how well their custom requirements would sell.
I'll probably wait and see how this phone does over the next few years. If it is still highly regarded by version 3 then I could imagine making the switch to a folding phone. For now I will remain conservative since my phone is such a critical piece of tech.
When the code is shitty it becomes harder and harder for the models to make changes and this grinds progress down to a halt - this has been my experience with “factories” trying them and doing refining steps every few months.
I sincerely don’t understand what the people who say they no longer read any code are doing, because it must be somewhat trivial to not run headlong into these issues that stack up time after time - then people say to just prompt better and it doesn’t have that problem for them, but I look at those same people’s code and it’s horrific, and then I find they haven’t made it far past a proof of concept phase. I watch entire teams slow down to a crawl and not be able to handle changes, or production incidents. This seems common among many people I talk to.
I personally think that the boosters need to put up or shut up - the promises are way over the skis. Every single person I’ve seen being a strong proponent of these techniques both has nearly unlimited tokens to spend and also seems to be in the business of selling a solution. I can’t find many not-currently-marketing-something engineers succeeding using these techniques in production systems unless they’re quite simple, or doing a very specific task from a more mature codebase.
I've been wondering whether AI really is improving rapidly at open problems or we're being fooled.
- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay
- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]
- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.
- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?
This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.
The title (likely intentionally) is misleading, it should say "travelling faster than light in a medium". Nothing here travels faster than light in vacuum.
BTW there are special types of telescopes used to observe gamma rays - they cannot see gamma ray directly but observe a flash of Cherenkov light of a cascade of charged particles created when gamma ray hits atoms in the atmosphere. Those telescopes are Imaging Atmospheric Cherenkov Telescopes [1].
Claude Product Manager: No, that'll piss people off too much, and sadly we can't just ask for ID either...
Claude Executive: There must be some way we can force people to link their government IDs with our platform so our analytics get better and more accurate?
Claude Product Manager: We could limit the platform to 18+ and use "Age Verification" as the reason for people to hand over IDs, seems other platforms had success with this approach
Claude Executive: And we hardly have any users younger than 18 anyway, go for it!