> Claude Opus 5.5 is our first release since we called for pacing the frontier.
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
> It found the U.S. “failed in its obligation to do everything feasible to verify” that the school was a military objective and that the failure “went beyond mere negligence.” The report said the United States “directed the strikes at the building of the school while being aware of a substantial risk of striking a civilian object and acting recklessly as regards the possibility that this would happen.”
Reading the details, "AI" doesn't really seem like the culprit -it's a scapegoat.
The intelligence that it was no longer a military target never entered the target database, the team that was responsible for vetting the target list was gutted, and said team was never even consulted.
The White House wanted 1000 targets and pulled from their database without any due diligence. Whether it was an AI call or an SQL query - this was from pure human maliciousness and incompetence.
Sorry folks, this is a rollout artifact, we needed a way to turn this off remotely via feature flags if it broke something, and with telemetry off you don't get those. It's already been fixed as part of v2.1.281 releasing today.
Not trying to start flamewars, but more and more I wonder how much are people willing to put up with such practices. On linux for 15+ years and everytime I try to use mac or win, it is an ordeal. And ads. F*cking ads in a system someone has purchased. And spying. Really user hostile environment. No, thank you.
At this point, no one seems capable of keeping a large database safe. I assume all medical and biographical information that exists is in the hands of the major state actors.
China hacked 22.1 million records of US government employees:
The author observes that a call to GPT-5.6 Luna is only 4-5 orders of magnitude more expensive than grep, and then predicts that at current rates of progress, calling an LLM will soon be cheaper than a grep. I think this is a good time to invoke Stein's Law: "If something cannot go on forever, it will stop." These efficiency improvements won't continue forever. It's more likely that the per-call cost of high-quality, compiled software like grep will be a lower-bound that LLMs asymptotically approach, rather than a line that they blow past with perpetual exponential progress. (Barring a true breakthrough in something like quantum computing or room-temperature superconductors.)
This statistic constantly confuses people who haven't done hiring.
If we decide we're going to grow headcount by 10 senior software engineers, we don't copy and paste the same job listing 10 times. We leave it open and collect resumes through that. At a large or fast growing company, we might never stop interviewing and hiring senior software engineers. The listing stays open for a year or more.
There are also specialty roles where hiring takes months. I've had niche roles open where we didn't get any applicants with any experience on the topic for over 90 days on multiple occasions.
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.
If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
I doubt John Ternus will change direction anytime soon, since it will look like he’s reversing the plethora of eyesore ads that Tim Cook added over his tenure.
But the Apple with ads is not the Apple that had some taste and discernment in the past. For a long time I’ve visited the App Store’s app update page directly (tap and hold on App Store icon to see the context menu option). Anytime I inadvertently go to the App Store home page or the few times I search, it’s an ad filled disaster!
From this article
> repeatedly attempting to prod customers towards even more of the company’s products might seem cheap, even distasteful.
From a recent post by John Gruber:
> Steve Jobs in 2011: 'We Build Products That We Want for Ourselves, Too, and We Just Don't Want Ads' [1]
Looks like Tim Cook, John Ternus and Eddy Cue really enjoy being swamped with ads in their products. Will there soon be a time when Apple executives start carrying some other brand’s devices with them to avoid having a rotten experience?
I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.
The meaning isn't clear at all. So open for interpretation that it is meaningless. That's the whole fucking point. For all I know they are "pacing the frontier", or not. The fact that there's no meaning to it let's you know that it was a pointless waste of tokens and attention.
> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.
We ended up with a Samsung fridge in our new home. It has no external display or any other indication that it is a "smart" fridge. I only found out when a friend visited who has a Samsung phone and the phone offered to connect to the fridge. We spent the next 30 minutes removing panels from the fridge and found the wifi/bluetooth antenna on a small board under the top-right hinge cover. The board was disconnected and the fridge continues to operate normally.
Also the OS mostly “just works”, just last week my Linux laptop disabled NVIDIA GPU (and almost bricked itself?? Not sure, had to fix apt) during automated updates
The problem imo is the slow deterioration of institutional knowledge that offloading the mental task of wisdom gathering to AI is causing.
One interesting comparison is to the history of manufacturing. West/America decided one day that manufacturing would be cheaper to outsource and better (short term) profit was to be made by outsourcing it all to China. The institutional expertise started to deteriorate, to the point that America simply didn't even have the capacity, or expertise anymore to produce stuff (such as grill brush [1])
I feel like you could take all the handwavy comment that are made today to dismiss this caution, and find equal dismissal back then when companies were actively outsourcing the manufacturing.
"I'm coding 10x faster"
"look at the output velocity per employee"
"we are producing much more (in China)"
"look at profit / number of (manufacturing) employers"
Seems ok if you're American / Chinese but I'm struggling to understand how the rest can be OK with allowing institutional knowledge to deteriorate while having an active dependency to the former two. We already see this with the tech dependency towards USA and manufacturing competition from China.
I'm not sure if it was any kind of official rule, but I worked at times in a JTAC capacity for troops in contact. I mention the last bit because it's very high stakes where seconds and minutes count a great deal more than when you do coordinated tactical strikes against strategic objectives like weapons caches. When I did that kind of work we had to have three eyes-on forms of contact. Often that was the calling troops, the observer (like an Air officer), and an air asset like a drone. They all had to independently describe the target, its orientation, and surrounding activity.
However you sourced your target is irrelevant to the activity at hand. When I first heard of this news I knew immediately someone skipped target verification or that it was simply no longer a policy of the DOD.
I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
It seems like it would be more effective to simply put all the solar panels in a field, and construct a cheap shade over the entire canal no? The supports in the picture are massive and don't look cheap. Also I would imagine the solar array uses more copper than a similar capacity array just built in a field. You can't daisy chain multiple miles of solar panels together, so you need an extra power line to run alongside the whole thing.
Put the solar panels in a field: The solar array uses less copper.
The shade supports don't have to hold up solar panels: Shade supports cost less.
The best reasoning they give is that California has insane permitting requirements, and it takes 1/6 the time to build on developed land compared to undeveloped land.
No, OpenAI did not solve the "wrong" Navier-Stokes problem. OpenAI did not solve the hardest version of the problem (unforced blow-up), but did give a solution to the Clay Millennium Prize Problem as written and understood, choosing the explicitly allowed forced option.
SciAm writes "in a sense, the LLM found and exploited a loophole in the framing of the question". This is pure sensationalism. Choosing option (C) (out of an explicit list of four options) is neither a "loophole" nor something "found by the LLM"; everyone involved knew this was the option they were pursuing.
With the grumbling out the way, there is some actual scientific content to the article: there's a strong argument that OpenAI's method will not extend to the unforced case, leaving our understanding of NS incomplete. This negative result is itself new and interesting (and predicated entirely on the solution found by OpenAI)!
Meanwhile, the 47 Muni bus that connects the Van Ness transit corridor to Caltrain has been “suspended” since 2020, and the extension of Caltrain to the transit center is still unfunded.
I appreciate solutions that meet us where we are, but it’s depressing that we don’t seem to actually have the will to make a sustainable, integrated mass transit plan.
It's not just a scene; it's the whole premise of the setting. It's why the Galactica survived and the newer ships did not. It's why the new Vipers got wiped out and they had to pull the old ones out of mothballs.
In the pilot, the Galactica was literally being turned into a museum, and that's why they lived.
It is unthinkable to me that anyone believes there is such a thing as computer security after so many years of nonstop hacks and leaks. If you have a computer and it is connected to a network with access to the Internet, assume that computer is semi-public. Meaning, if someone was interested enough in accessing your computer, they could do it. Do not hook any computer with access to anything that would be devastating if it was made public to the Internet. Do not put anything that would be devastating if it was made public onto someone else's Internet-connected computers.
For example, do not hook your goddamn water or traffic or electricity infrastructure up to the goddamn Internet, and then, do fire the guy who suggested it.
The correct analogy for computer security is not locks and keys and doors and gates. It is a house in a floodplain. Your house will not survive the flood of it hits you. Do not store anything critical or irreplaceable in that house.
> The reason there's not more public transit is because the economics are terrible.
The reason we don't have roads is because the economics are terrible. The government gets $0 per trip [1], but it costs billions of dollars a year in operating expenses.
Wait, that's not how it works.
[1] Federal gas taxes don't count, because those go to the Highway Trust Fund which funds road construction, which is not an operating expense. State gas taxes may vary. Although toll roads also properly charge people user fees, although for some reason, lots of drivers complain about how expensive tolls are...
This happened after I stopped working for the DOD, but I am 99% sure the program in question during this incident was something I worked on. It was effectively an anomaly detection system for identifying any weird behaviors detectable in WAMI (wide area motion imagery, which is high resolution, low framerate drone video data that covers a large geographic area). As an example from the unclassified sample project they asked us to do as part of the bidding process, our code flagged a car doing donuts in a parking lot (the sample data was just from a random US city on a random day because this was an unclassified open bidding process). The system was not and never was intended to be a "terrorist" detection system. The idea was always to flag potentially interesting events for a human analyst to follow up on. The disconnect it took for someone to blindly trust the program as a target acquisition system boggles my mind. Someone doing donuts in a parking lot should not be automatically targeted with hellfire missiles, yet in practice that is how they were using the tool. I still lose sleep over it.
Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.