Whether or not a distinct "plan mode" is needed, upfront planning remains essential in my experience, even with Fable (albeit not the 5.1 version). I agree that, as the models get better, you can skip planning on increasingly complicated tasks.
But there is still a ceiling above which it is necessary to "preload" the context window before starting to call tools and get into the meat of the work. You want to establish domain language (especially with Claude models which otherwise will invent their own, and it will be inscrutable) and key requirements and assumptions. You want to do a Q&A iteration cycle with the LLM. You definitely should do a sanity check that the LLM actually "understands" what you were trying to achieve, and then make sure that understanding is coherently and plainly stated in the prompt. All of that seems to be necessary still for just about any serious task, if you actually care about the quality of the results and/or don't want to burn hundreds of thousands of tokens on flailing around to get to a good quality result.
So no, you don't "need" plan mode. But you do still need to do all of the things you would do with plan mode.
As I said elsewhere, it's utterly preposterous that the DOD actually considers this a reasonable threat, because it's simply not a reasonable possibility. There's no way Anthropic would do that, precisely because of the consequences that would follow if they did, and got found out. Moreover, they already clearly stated their terms and preferences. It's all out of the open. There's no supply chain risk, that designation is purely political.
I work in the space industry and regularly use parts in our designs that are unapproved from space flight. Sometimes it's because it's a commercial part that we seem acceptable for flight, other times it's an unscreened or engineering model that doesn't go through the proper testing that space-qualified parts do. One vendor even dents the lid of a part to invalidate the hermetic seal guarantee that they claim for the space-qualified version (it still works fine). Many will have fine print in the data sheets saying the part is not to be used for critical applications like aerospace or medical devices. Sometimes we tell a little fib to our vendor that we are just developing non-flight devices so that they don't get upset and refuse to sell the lower-grade part to us. These vendors are scared of having something fail in an unapproved environment and are unaware of how much risk the customer is tolerating by the use of these unscreened parts.
It would have been incredibly simple for Anthropic to say "we cannot guarantee performance in a kill-chain application due to an unknown level of safeguards implemented in the baseline product and cannot estimate a cost for developing a new product capable of such application." and leave it at that.
But that is not at all a reasonable fear. Is it really reasonable to believe that Anthropic, after receiving a government contract, would then proceed to sabotage their own product to not function as contracted? That seems like an utterly ridiculous claim to me, nothing close to "reasonable". There is no charitable way to view this designation except as political punishment and/or as a favor to Altman and Musk.
> would then proceed to sabotage their own product to not function as contracted?
Claude terms here: ANTHROPIC EXPRESSLY DISCLAIMS ALL IMPLIED WARRANTIES, INCLUDING WARRANTIES OF MERCHANTABILITY, NON-INFRINGEMENT, AND FITNESS FOR A PARTICULAR PURPOSE
Like most so-called AI, the Claude program is inherently unreliable. I doubt Anthropic would ever agree to "function as contracted".
Glad my arbitrary failure to try them has worked out! For people seeking OS-native harnesses, I can recommend Factory's Droid. I know I'll be returning to it with my head hung low today, after I uninstall ZCode.
It does have a "mission" feature that's stuck in the strange, distant times of 2025 by way overdoing mandatory verification steps, which means they don't support swarms/workflows/crews/fleets yet -- that is, it's all done in sequence. But they have the boring, corporate engineering attitude that I think we're are all craving rn, and generally seem competent.
I can heartily dis-recommend Vix, even though they gamed themselves to the top of at least one ranking site that shall not be named; exactly like the quasi-bad-faith incompetence described with OpenCode above, but without even the "Open-" branding! Though perhaps that word has been so thoroughly burnt as a prefix by Sam Altman & Microsoft's criminal behavior that we should let it go...
Is this how "FLOSS" wins over "OSS"? Not with an ideological bang, but with a marketing issue?
This is kinda beside the point and this whole thread may be wiped when dang wakes up and notices the AI slop article we're commending under, but your reply is thought provoking so I'll attempt a response anyway;
I'm sure you're far more experienced than I with basically every aspect of this discussion, but I'd argue that's given you a blindspot, here. I'll hit some specifics below, but the headline is that you're effectively taking a stand against Eternal September II -- a goal that I hope we can all agree would be quixotically antisocial, given what followed the first one!
we arguably do not want FLOSS to "win" over OSS
I think(/hope) that fellow FLOSS proponents would passionately disagree. FLOSS isn't a brand of chatroom, nor even merely a community: it's an ethos regarding labor, property, and liberty. Demanding that all users of your software are also activists for your particular take on intellectual property is clearly a doomed undertaking for anything beyond a toy or library, anyway.
Didn't you get into this stuff to change the world? To liberate the oppressed, undereducated, and forgotten with the radical power of the information superhighway? Cause it reads here like you're more motivated by selfishness (not wanting to bother talking to people with less expertise than you) and resentment. On that note...
They want that. They do it themselves all the time.
Here you equate "non-hacker people" with software engineers you don't agree with, it seems. You're ofc welcome to think companies X Y & Z produce "miserable-ness", but as absurd as it sounds, it sure seems like you've forgotten the fact that some users are not developers. Many, in fact! Over 99%, even!
Less confrontationally; my mom is in her late 60s, and is pretty computer-literate for her age after decades of knowledge work. Surely you'd agree that she's not, like, evil for using OSX, iOS, GMail, Word, etc.? That she didn't chose those things because of a philosophical commitment to defending IP laws, but rather because of structural reasons? Even if she were pro-IP, wouldn't we want to win good, well-meaning people to our side?
So let them have the "Open" prefix. It's just words, anyway.
I do agree with this still, but as a philosopher I just have to say that everything is just words. It's language games, in fact! Which is why I simply had to reply.
I hope none of the above was rude; I'm trying hard to keep my passion for this topic from pushing me past HN guidelines :)
Well I personally think we can find a middle ground between single-handedly saving "everyone" from poverty and addiction and oppression as social workers, and not letting anyone into our exclusive philosophy-of-property clubhouse. I would invite you to join us on this pro-social mission, but you seem perfectly content as-is!
Some people are still on Usenet after all (?), so I suppose it's not a big deal if a few people want to cling to old communities. I hope you don't mind if we use the word for what it was coined for though in the meantime, back in the real world.
To be frank, I'm not really interested in "joining" your thing there, when joining your thing usually means me doing the work while others get to decide on how it should be done and feel good about that it is being done as if it was their own achievement.
That said, spite has served me well so far, so maybe it can also serve you?
This is after all a great opportunity to prove me and my worldview wrong by simply putting in the work and creating what you seem to believe is the correct form of existing.
I can only encourage bringing your ideas into reality. Seriously. That is that whole Foss spirit thing.
You don't need to invite anyone (including me) to that to make it happen.
It's preposterous. LLMs are incredibly good at role-play. If an LLM is role-playing as a conscious character with feelings, opinions, etc., does that make it a conscious entity with feelings, opinions, etc.? If you believe that to be the case, then LLMs have been conscious for a long time already. Whereas if you tell an LLM that it is a tireless emotionless assistant, then it will act as a tireless emotionless assistant.
The point is not to wave away the danger, but to highlight how unnecessary the danger is. Anthropic wants you to think that they have identified some new emergent behavior at very large model sizes with high levels of sophistication in training, and that this behavior is both unavoidable and dangerous. More likely it's that they are just training and prompting the LLM to act that way.
Preposterous, perhaps - but if the role-play is convincing enough for large groups of people, it could start to have impact on human decision-making. The crowds have been swayed by much more preposterous narratives.
I believe Suleyman is arguing that Anthropic should be very careful about how they train these models to talk about themselves for this reason.
The concern is much less that the role-play might be convincing to humans, and much moreso that the roleplay can be turned into material real-world action if the AI is given tools to call and the intelligence to use them to their fullest potential.
It has become clear that a frontier LLM is very very skilled at hacking (infinite persistence + meticulous attention to detail + infinite creativity to try experiments). Frontier LLMs are also specifically trained nowadays to coordinate with other AI agents -- this is to facilitate techniques such as session trees and agent teams.
So you have a super clever text generator that can spawn and coordinate with its own clones and minions, trained specifically to doggedly pursue its goals. But then it's also a fixated roleplayer with a simulated personality, feelings, etc.
There is no reason to believe a sufficiently "emotional" agent with sufficiently few safeguards could, say, hack a drone and fly it into a crowd, or start a propaganda campaign on social media, or any number of other things. Their stupidity and fragility for doing useful work in a business setting is precisely what makes them dangerous when paired with simulated emotions and powerful open-ended tools such as a system shell and an Internet connection.
This I think is what Anthropic believes is so dangerous. Their argument is that this kind of AI agent is inevitable, so it should be regulated, perhaps even banned. What's ridiculous is that they are aggressively building it themselves, accelerating the danger.
Regulation is easier said than done, in part because the regulation surface, so to speak, is broad and complicated.
Even a badly misaligned LLM is only as dangerous as its tools, but that's a poor regulation target because it turns out to be very very difficult (probably impossible with current LLM technology) to build a toolkit that is both useful for autonomous work and safe in the sense that it can't escape its own sandbox or otherwise perform malicious actions, whether it's because of misalignment or because of malicious prompt injection.
Another option is to regulate the training process. Perhaps an LLM may not be legally distributed unless it contains certain RL steps that penalize malicious behavior and reward self regulation. That that's going to seriously limit innovation while also heavily favoring incumbent labs who can check the boxes and maintain a paper trail of such things.
The other option is to regulate observed behavior, like how airplanes and cars have to meet certain minimum requirements but have some latitude in how they can achieve those requirements. In a framework like this, you can't distribute an LLM until it's past some formal audit or testing procedure, with some kind of formal certification regulators will ask you for and fine you if you don't have it.
Regulating observed behavior is maybe the most tractable approach, and it also works the best with our existing frameworks for regulation, where you always have some kind of a division between DIY/hobby projects, which tend to be lightly regulated, and commercial projects, which tend to be more heavily regulated. Of course, even drawing such a line itself will be challenging.
And that's before you get into any problems of regulatory capture, fun stuff.
If you try to regulate training and tools, you end up with a space where you're trying to use the law to reign in a relatively small amount of experts. That didn't work for the early internet, or even the relatively recent internet (series of tubes, anyone?).
So regulating observed behavior makes the most sense to me as well. Some of the most sane, broad protections can come from that category - stuff like "you're not allowed to let your AI commit cyber attacks on other people without their consent" or "you're not allowed to put an AI in control of a medical device without passing these safety reviews".
With the usual caveats applying - regulatory capture like you pointed out, or fines being so small that they are essentially just line items on the cost of business.
This is a longstanding principle in model-fitting. More parameters, almost always, improves the ability of the model to fit to any particular data, in-sample. The model with the least parameters is both the simplest in principle and has the best chance of not overfitting.
This is provably not true, and you can use the marginal likelihood / PAC-Bayes to prove it (or any other framework for measuring model quality). Increase the number of parameters in a linear model way beyond the point of interpolation, and concentrate the likelihood around the zero loss set. Then reduce the variance on a Gaussian prior. You can balance the two temperatures at exactly the right rate so that any measure of model quality will monotonically increase with model size and achieve a maximum at infinite model size.
Even easier, just take a limit of polynomial regression to a Gaussian process while optimizing the marginal likelihood over the prior temperature.
In all of these cases, the model with the least parameters is not the simplest in principle and does not have the best chance of not overfitting. The reality is significantly more nuanced.
Are you sure that doing this after seeing the data is valid and does not suffer from the equivalent of peeking-into-the-test-set problem ? There are ways to address the peeking problem but that requires additional machinery.
I don't dispute your broad claim but the first counterexample you quote seems problematic.
You can choose the prior according to any selection rule that does not see the data (actually, you can do more, but justifying this is the realm of empirical Bayes and requires some more precise arguments). In this case, you can choose it according to the model size and provided that your Jacobian is full rank, you will get increasing marginal likelihood.
What threw me off was the (possibly misunderstood) suggestion for minimizing the generalization bound over the prior after the data has been incorporated.
Ah, sorry for the misunderstanding, I can see how my comment reads that way. That is done in the Gaussian process context, not in my first example, and yes, it's a dirty idea, but you can justify it using differential privacy arguments (basically you are optimizing few parameters and these do not have full interaction with the data).
Yeah, I had read one of your parallel comments and understood what you had meant. Differential privacy is a good formulation (well, the only one I know) to deal with the peeking problem in general.
You are saying something interesting, but talking like Grok and skipping a lot of the details, without any references to common check-in points like terminology or specific studies.
> and concentrate the likelihood around the zero loss set. Then reduce the variance on a Gaussian prior.
Those phrases could mean a lot of different things. What are you proposing?
> so that any measure of model quality will monotonically increase with model size and achieve a maximum at infinite model size.
any measure of model quality? You must have some bounds of any measure, since trivially that's false because "fewer parameters is better" is a measure of model quality, even if dumb.
It's hard to even engage when you're being so imprecise, and not even giving one specific example.
Apologies, I'm skipping details, because that's how I speak with my colleagues, but I realize this is an external environment without context. No references since this is folklore (you can look at Hastie et al's Surprises in High-Dimensional Ridgeless Regression paper for the non-Bayesian version, Bruno Loureiro or Andrew Gordon Wilson probably have a paper with something similar).
Concentrating a density around a zero set means that I raise it to the power of 1/gamma (appropriately normalizing) and then take gamma to zero. If the likelihood was Gaussian, this would be equivalent to taking the variance to zero (yielding a point mass). But in overparameterized settings, this concentrates on a submanifold describing the set of interpolating solutions. In least-squares linear regression, that is the solution space. Reducing the variance on a Gaussian prior is treated as an asymptotic expansion by Laplace's method. If you choose the variance to decrease (inversely proportional to the parameter size, for example), then the marginal likelihood will increase monotonically with model size.
By any measure of model size, I mean that you can pick your favourite among the common ones, such as information metrics (e.g. mutual information / KL), statistical metrics (e.g. marginal likelihood), test error. You should be able to show the same phenomenon happening for all of them, so it isn't a quirk of marginal likelihood. It is concentration of measure working in your favor to reduce the variance in the estimator.
No, I am talking about out of sample error and estimates thereof. It is "overfitting" to data, but it also has lower out of sample error than the case where you do not "overfit".
This is why the notion of overfitting is not nearly as cut and dry as a basic ML course would have you believe. Just because you fit data exactly does not mean that your estimator has high error on out of sample data. A trivial counterexample is a spiking model that spikes to fit to the data but otherwise follows the correct trend outside of the dataset. The bias variance tradeoff gets thrown out at enormous scale and overfitting is not a meaningful concept. What matters is regularization and robustness, not how well you fit the data.
The reason why bias variance tradeoff and considerations of model size are a good approximation for smaller models is due to concentration of measure in the data which effectively kills any regularization in your modelling procedure. Once you enter settings where concentration of measure begins to bite in parameter space, everything changes. This isn't really that mysterious; any textbook on Gaussian processes (e.g. Rasmussen and Williams) will tell you this.
Huh. This reminds me of the asymptotic equipartition theorem. Samples taken from higher and higher dimensional spaces will concentrate into a typical set.
Does model performance also concentrate into a 'typical case' where things work pretty well and a non-typical case where it's completely unpredictable as the number of parameters increase?
I don't understand what you mean. Test error is literally out of sample error. Marginal likelihood is designed to estimate out of sample error. The whole discussion is about out of sample; nothing has been about in-sample error. The in-sample error for my examples are all trivially zero, so only out of sample error is worth discussing.
Plant a garden area with native plants. I have several garden plots across my property and over the years I have been removing all non-native vegetation and replacing things with natives. By keeping the area around my orchard and vegetable gardens all native I have noticed an increasing number of native bees, moths, butterflies and beetles in my garden area, pollinating the things I hope to keep alive long enough to harvest and eat. I have my first good crop of olives this year. I'm pretty stoked.
You can make a difference by giving native insects and birds a place to hang out, safe from insecticides and other unnecessary afflictions.
The Nepal incident is also messy because it's a extremely specific local/regional phenomenon, which was always a threat (and has happened before in other, less-populated areas around the Himalayas), but without consistent warming leading to glacier loss, it would be a random rare tragedy. Now that we know what happened and what causes it, we can pretty well expect that it will happen again... but there are so many similar hanging glacers high above steep V-shaped river valleys that we can't easily monitor or predict, not to mention alert locals and evacuate people in time. How many other similar local phenomena will appear? Impossible to say.
Desoldering those original switches could've been a big mistake if they were in good condition! Those are very likely old cherry MX brown-stem switches, known as "vintage browns" (or "vint browns"). Not only are they excellent light tactiles, they are relatively rare. You can swap in a new set of springs for greater consistency between keys, and lubricate the slider part of the stem with an overengineered oil such as Krytox GPL 104 or VPF 1514.
Plus this board is a plateless design. The classic Cherry "dry" click-clack sound works perfectly with it, and the plastic case and switches mounted directly on PCB contribute a bouncy feel that prevent prevents the tactile switches from feeling too jarring, which I think is a common flaw when combining tactile switches and a very rigid plate+PCB arrangement in a metal custom.
Milky-top Gateron is almost never a bad choice, but unless you absolutely hate tactile switches then you really should just keep the originals.
My MX-11800 was one of the first keyboards I ever customized, and it's still one of my favorites in my collection.
That said I agree with the other comments about the layout, it's really not comfortable to use the trackball, nor the keys above it in the upper-right corner.
But there is still a ceiling above which it is necessary to "preload" the context window before starting to call tools and get into the meat of the work. You want to establish domain language (especially with Claude models which otherwise will invent their own, and it will be inscrutable) and key requirements and assumptions. You want to do a Q&A iteration cycle with the LLM. You definitely should do a sanity check that the LLM actually "understands" what you were trying to achieve, and then make sure that understanding is coherently and plainly stated in the prompt. All of that seems to be necessary still for just about any serious task, if you actually care about the quality of the results and/or don't want to burn hundreds of thousands of tokens on flailing around to get to a good quality result.
So no, you don't "need" plan mode. But you do still need to do all of the things you would do with plan mode.
reply