> I think it is fundamentally a crisis of trust. I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.
[...]
> I don’t think that a glitzy marketing campaign with a positive spin (which some have advocated that Anthropic do) is the way to win back that trust — at this point, saying that AI will cure cancer is more a cliche than it is inspiring, and most people think it is deceptive. The thing that will work is actually curing cancer. I think by far the most accurate criticism of AI companies including Anthropic is that we haven’t yet delivered on our big promises to benefit the world. That is totally on us, and I think it’s the criticism you should be making, instead of all this stuff about messaging and marketing.
> We are however doing our best to fix this: Anthropic is ramping up its efforts very quickly in biology and medicine, and we hope to have incredible results in the coming years and some early glimmers in the coming months. When we’ve actually accomplished something real, the whole world will hear about it, as loudly as possible, you have my word on that.
Turns out "winning back trust" doesn't have anything to do with any actual concerns people may have re: employment, electricity prices, stock market bubble, intellectual property, scams, cybersecurity, environmental issues etc. Rather we'll just do all of that even harder and the miracles ("curing cancer", lol) we've so far failed to deliver are bound to arrive in short order!
> The thing that will work is actually curing cancer.
This is honestly hilarious. Dario must think people are stupid.
First, cancer research charities are some of the most well funded on this planet. They all obtain donations on the basis of "together, we will cure cancer" messaging. They all fund the best science they can.
Second, the human body is a complicated thing. You can feed your fancy LLM as many textbooks and academic papers as you like, but the reality on the hospital ward will always be different. Why do you think student doctors have to spend so many years "doing the rounds" Dario ? They are all academically smart, they are all capable of memorizing text books ... but there is no substitute for seeing and doing the reality.
I barely trust Claude to write code, let alone find a cure to cancer.
> The thing that will work is actually curing cancer.
I think typical bubble behavior the leaders have set up the whole promise to fail.
Everyone is expecting some faux super intelligence to come and find a cancer solution everyone else missed. However, it is just as likely that vanilla current LLM's will create enough of a productivity boost for back office automations in research heavy hospitals to create the space for regular humans to create these breakthroughs, but LLM's won't be able to claim that for themselves and inevitably “fail”.
Savings from back office automation go entirely into administrative bloat. I find it even less likely such efficiency would lead to breakthroughs than I do LLMs becoming super intelligent and doing it themselves, which I don’t find likely at all.
> I mean they did solve the protein folding problem did they?
They made huge progress, but I would say that the vast majority of work on this problem was designing the harness for the model. That's a lot of work for each and every domain.
The effective altruist and longtermist crowd are into it as a religion. They would happily sacrifice anything human to satisfy their dreams of AI. They care way more about what the AI needs than their what their fellow humans need. We will end up with an AGI having better living and working conditions than humans and the whole AI industry will be clapping
It's an ideological thing. (We just had a conversation about ideology actually, remember?)
Capitalism has always been a tension between capitalists, who want maximum return on capital and workers, who want pesky things like a living wage, sick leave or safe work conditions.
For the longest time, the only power labor has had to get those things was the power of collective bargaining. No agreement with your workers meant no production happened.
Now, there's finally a chance at salvation. The capitalist Messiah is AGI and it will finally deliver them from those annoying laborers.
This ideological bent is why they're putting everything they have into AI. It's why VCs and their fellow capitalists are going so crazy.
They see it as a way to finally solve the contradictions of their ideology, but in reality it'd only create a new one: If everyone's out of a job, who will buy their products?
I think it's more reasonable than it sounds on the surface. People's jobs, the bubble, etc are societal level problems and well beyond what Anthropic could even hope to influence on their own.
Curing cancer sounds insane, but it's also a research problem, not a societal level coordination problem. And one AI has already proved to help with breakthroughs (alphafold). IMO it makes sense for them to shoot for something like that as proof of AI's beneficial sides.
>People's jobs, the bubble, etc are societal level problems and well beyond what Anthropic could even hope to influence on their own.
This is not true. There are two companies at the center of AI direction: Anthropic and OpenAI. If there were anyone on the plant who has the ability to influence our direction then it would be Dario Amodei.
Acknowledging the grievances is a good step, but it's not enough. There needs to be a clear explanation of actions to address them, a plan to enact those actions, and commitments with consequences in failure of those actions. Tell people how you're going to make them more employable and effective and needed. Tell people how your datacenters will be carbon neutral. Tell people how financial actions resulting in a frothy market will be coming to an end. He and Sam Altman alone have this power and their inaction says everything we need to know about their intent.
> Curing cancer sounds insane, but it's also a research problem,
I just don't buy the AI labs approach to this stuff. Like, unless we can basically simulate the entirety of human biology, I don't really see how LLMs can make progress here. Maths is different as it doesn't require a real-world interface, and programming already (by definition) can be simulated on a computer.
Without that, I can't see much (if any) progress being made on domains like biology.
I'm a noob on this topic, but I think drug discovery is more amenable to this structurally than other problems. Simulating biology is what we were doing with protein folding before Alphafold, and the search space was far too large to find stuff in reasonable timeframes. Alphafold showed that you could take a physical process and make a neural net clever enough to learn just enough structure that it starts finding things we might care about, and still physically accurate, much faster.
Drug discovery is similar AFAIK. The space of possibilities is even larger than protein folding, but it's structurally similar enough that I think AI will help to make progress on the discovery side. Actually getting the drug tested and approved is another matter though for sure.
For one, the question for Anthropic is whether LLMs, specifically, not AI techniques more generally, can help significantly with cancer research. And here, all experience so far is that LLMs only really work when they can easily automatically verify their own outputs and self correct - such as in math (using automatic proof verifiers) or programming (using compilers and unit tests).
The second problem is that biological research speed is highly dependent on slow biological processes, such as cultures and long term studies. In programming, if an LLM could provide excellent insights and research suggestions 100x faster than a human, it would speed up the work roughly 100x. But in biology, it would only speed up the total work by a small amount - as any insight, even if absolutely brilliant and spot on, would still require months and years of actual experimentation.
The real issue is that biology needs actual experiments done in the physical world, which isn't nice and orderly and well behaved and easily loadable onto a 19" rectangular box.
On the other hand, AI means actual experiments done in the physical world but coordinated by an entity that never sleeps and never gets depressed and can multiply itself manifold and always comes up with new ideas.
> People's jobs, the bubble, etc are societal level problems and well beyond what Anthropic could even hope to influence on their own.
Anthropic is literally creating the bubble. It is not beyond their scope of influence, it is literally what they are consciously achieving.
As for peoples jobs, same actually applies. Anthropic is selling itself on dream of replacing jobs, even or especially where they are well aware AI does not perform that well. They are actively trying to replace people quickly before management notices it does not work well.
And also, they can influence how much their data centers contribute to global warming.
This is actually crazy, because biochemistry is where LLMs are weakest. It's one of science's most fuzzy and unpredictable domains, in general, and a lot of published works which an LLM might take at face value might be unreplicable or subtly flawed. (See e.g. the entire history of Alzheimer's.) It is also where real-world lab and clinical trial work is most important.
So Anthropic are going to spin up a medicinal chemistry lab and start mouse experiments?
I mean, it would be nice, but I don't think that the guys at Anthropic know what they're talking about here, or what they might be getting themselves into. (If indeed this is more than just PR.)
AI bros consistently handwave away any notions of the material world imposing limits, because (spoiler!) it's a religion where you start from the miracle and work backwards. "Oh, we'll just set up fully automated robotic research labs everywhere!" Where will you get the raw materials? Oh we'll mine the asteroids! How will you find the energy to get there? Oh, the AI will invent new physics! (again without needing to run experiments). Et.c.
Didn't and don't mean to disparage anyone's medical struggles, but "a cure for cancer" is a well-worn strawman. One that's been achieved for the low hanging fruit, the higher ones are seeing steady progress (already before LLM chatbots, even!), the bottleneck isn't "intelligence" and to the extent there are socioeconomic (access to screening, treatment) or environmental/lifestyle factors involved the AI boom is likely just making things worse!
I'm sure Dario knows this, and it's anyway too pedestrian compared to the usual list of fruits of ASI. The text probably originally read "nanobots eating you alive and uploading to the cloud" or something, but they figured that wouldn't go over with the intended audience. "What do the peasants care about? Oh I know! Curing cancer!"
> AI bros consistently handwave away any notions of the material world imposing limits, because (spoiler!) it's a religion where you start from the miracle and work backwards.
How else are you going to work towards a dream? It's how Elon Musk managed to get reusable rockets when everyone said it's unfeasible.
Granted, most ideas don't work, this is why you need testing, but I am a bit surprised to see this attitude on hacker news.
If they said that their dream is to solve physics, I'd say "that's cool and I can't wait to see how it turns out." LLMs have gotten so good at mathematics, and are so adroit at analyzing vast quantities of data, that they might have a really good shot.
But they said that they want to cure cancer. That's far outside the core competencies of any LLM, and it requires a lot of real-world wetwork with liquids, chemicals, cell line experiments, animal experiments, etc. They can't merely analyze existing data -- they'd need to generate vast amounts of new data, which isn't really the case in physics. And then regulatory approvals and so forth.
"We're working on curing cancer" sounds more like a poor PR attempt than an actual effort, though I'd love to be wrong.
Not really. The problem is that the trust is violated here. What people fear is that government and large companies are going to use AI against them, and make the reverse impossible. Sadly, this is exactly what is happening.
And that's exactly what EU's AI regulation does. Of course, AI is being blamed for this happening, despite that last time I checked every EU parlementarian, every EU commission member is flesh and blood, the fact that it is 100% intentional government policy, humans, that are doing this to you using very un-AI methods.
You see, in this regulation, governments are allowed to use AI, and to approve uses of AI. So there is absolutely nothing in EU's AI directive at all that prevents government and large companies, with approval, from answering the phone, and all "support requests" with AI, and not with people. Also "if you agree" (just like you've agreed to selling your location data in your cell contract)
The watermarking algorithm itself has a VERY specific property that should have set off everyone's alarm bells: it is NOT the case that you can "check text for AI watermark". What it technically mandates is that if you provide access to a model, those people should be able to check if text is generated by that model. Not by any other model. It is NOT a general "is this AI?" check. So let's analyze how this works in 2 specific cases:
You get a mail from the government about taxes. You want to check if that text is generated by an AI and whether you even should ask to talk to a human. So you've got a piece of text, or an audio recording. Can you check if it's AI or not? NO YOU CAN'T. SynthID requires access to the model keys, and the legislation only mandates checks if you've got access to the model.
The government or some large company, however, puts in it's regulations that you're not allowed to use AI to communicate with them (we all know this is coming), and they get a piece of ChatGPT text (or from any public model) from you. They would like to enforce that they won't talk to AI can they do this? YES, since they have access to the model too, SynthID allows for this usecase AND the AI regulation makes this mandatory.
Is this in any way a problem with AI? No. Can anything be done about this by "fighting AI"? Well let's see ... will government institutions be forced to stop worsening services further if ChatGPT etc become inaccessible or obviously excluded by government and large company services? No.
What is being outlawed in the EU, in other words, is exactly ONE thing: that you use AI to help you in your dealings with governments and large companies. To assert your rights, to resolve issues faster, to make them wait rather than you, to ... THAT is being outlawed, nothing else. Only that YOU get help from AI.
This legislation will make it a big advantage to have access to "your own" model, because you'll have the SynthID keys, and the law behind you if you do that. And I don't mean local model, I mean your own model with your own SynthID keys, ie. not Google, not ChatGPT, not Claude, not Grok, and if the EU can make it happen, not Chinese models either. Finetuned models, however, will not have SynthID and so you'll never be able to prove text generated by them is AI, and you will not effectively have the power to refuse contracts or laws that specify they can use AI anyway.
It's the EU government making it clear they are using AI to avoid even having to talk to you at all while disallowing the use of AI by mere citizens to assert their rights, against them, or anything they have interests in, like phone or electricity companies.
>The watermarking algorithm itself has a VERY specific property that should have set off everyone's alarm bells: it is NOT the case that you can "check text for AI watermark". What it technically mandates is that if you provide access to a model, those people should be able to check if text is generated by that model. Not by any other model. It is NOT a general "is this AI?"
Where do I find that in the Act (or code of practice etc.)? IANAL, but Article 50 reads different to me but if there is a comment or guide how to read it - also fair enough.
There is nothing in the AI act that requires offline tooling, nor is there anything about requiring the passwords and/or API to work for every model, only for the people you are providing the model to.
“The detection solution may be made available in the Union as one or more of the following: (i) a public, ideally standardised, specification allowing any third party to implement a detection mechanism; (ii) a piece of software (e.g., a standalone executable or library); (iii) a cloud-based service accessible to users in the Union through an API.”
Option 3: "a cloud-based service".
And accessible "to users". NOT to the public.
Note: this is not the actual law, it is the code of practice, ie. guidance for model providers. The start of that document clearly states that you can comply with the law in other ways if you want, you'll just have to justify yourself. So you have a choice to not even do this.
As to who decides the law clearly states who decides if someone is legal, like in most EU legislation. It's not the courts, it's "National market surveillance authorities designated by each EU Member State", and there is an EU office as well (and it is explicitly stated that they are not allowed to override each other). So every EU country has the right to provide exceptions to the law, just like they do for the GPDR. You do not have any rights under this legislation as an individual. Only these "National market surveillance authorities" get rights under this legislation.
Secondary: article 7 additionally gives the EU commission the power to declare any code of conduct they want that declares what compliance with the AI act actually means.
(and, of course, this is yet another attempt at declaring math illegal. The only way to actually enforce this legislation is for all models to comply with this, all over the internet. Obviously the EU does not remotely have the power to make that happen)
While there is a large element of paranoia here, we're already seeing serious problems in e.g. planning consultations that people can just spam them with AI. AI makes a lot of traditional social processes stop working properly.
> puts in it's regulations that you're not allowed to use AI to communicate with them (we all know this is coming),
.. do we?
> You get a mail from the government about taxes. You want to check if that text is generated by an AI
No, actually, I want to know whether it's correct and whether it's legally binding. Using AI makes it less likely to be correct, sure, but ultimately whether it's human, spreadsheet, or AI the important thing is the legal right to correct process.
I genuinely think Dario is a well intentioned, intelligent dude, but I think him and Anthropic have a huge PR problem and are really out of touch with how they're perceived.
Anthropic in particular has developed this almost Orwellian like veil of condescending rhetoric that on the surface suggests they're looking out for you while underneath they're taking actions that suggest they do not trust you, Mr/Mrs Ordinary Person. All the safety rhetoric, never supporting open weight models, the lockdowns on harnesses outside claude code, etc.
Anthropic if you care about public good, do something to empower people. Release an OSS model. Open source Claude Code. Open source some inference tooling or something. Just give people anything except your words.
Sorry but this sounds like such a populist claim. They probably also write their war plans on Microsoft Word, and drive Fords to the office! Not to mention the simple fact that there are no wars without mistakes. I'm not justifying this specific war, but as a service provider you do not get to approve or reject specific operations.
> Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers).
This is actually a good point. Open models are getting better and better, some of them might even be useful in consumer hardware now. But if AI performance is still correlated with compute power, then no doubt power will remain with the people owning the chips.
Swap that point about AI with "electricity". Everything runs on electricity it's "a technology that tends to concentrate power" (no pun intended). The electricity providers must be too powerful... But somehow electricity providers aren't that powerful. Unless there is no competition in sight...
The point "AI is structurally a technology that tends to concentrate power" is not that correct. They need this statement to be true, otherwise no way to justify the trillion evaluations.
Swap that to "oil" and everyone goes "well, yes, obviously, oil is so powerful that people start wars over it".
Iran has hit the Amazon data center in Bahrain. The Ukranian deep drone bombing campaign has hit refineries, but also Wildberries, the "Russian Amazon" warehouses. The US campaign against Iran now, Iraq and Serbia previously, targeted power infrastructure. It will obviously be a target in the next war, and AI goes on that list too.
Every time the US completes an AI data center, someone in the Chinese nuclear command updates their target priority list. And vice versa.
> The point "AI is structurally a technology that tends to concentrate power" is not that correct. They need this statement to be true, otherwise no way to justify the trillion evaluations.
How is it not true?
The biggest companies in the world are, quite obviously (just look at the numbers) going to be AI companies. The most powerful governments in the world are quite clearly going to be those that are close to (or in control of) AI companies. People have a very hard time seeing the second order effects of the control of intelligence - we need to fix that.
but its difficult to use more electricity directly to get more utility. There are some rare ways - things like PtX systems - but even those are linear at best. AI is concentrating power in the sense that nobody wants to use the twentieth smartest AI - so if you can use the most compute, you get an outsized share of the rewards (in terms of paying customers) - over and above your share of the compute. There are other business sectors where similar dynamics exist - where the biggest capital tends to win - a sort of natural monopoly type situation, its nothing about hyperscaling or whatever.
Competition is only one of the ways to keep a corporation aligned. In the case of France, having a state monopoly on electricity even worked very well until we broke it for the sake of competition.
Current AI is at least a few orders of magnitude less efficient than it could be. At some point the labs put too much work and money into transformers and nearly abandoned fundamental research (in both ML and hardware). There are tons of low hanging fruits in efficiency but you'll have to redo everything from scratch so nobody bothers. Which is also pretty convenient and lets people like Dario Amodei speak about "natural concentrations of power".
The Chinese labs are picking up on the low hanging fruits on efficiency, and no, you do not need to abandon transformers, you just need to push them closer to the more computationally efficient architectures of the past. OpenAI and Google seem to be trying a few things too.
Anthropic clearly are not though, and to call their operations wasteful is an understatement.
Arguably the question is whether it's economically feasible to self-host something similar.
Are you self-hosting Google or Bing? No, but we have quite a huge ecosystem of full-text search tools with PageRank, with options to scale to almost Google scale (if you have the money). After all LLM training starts with the same crawl mechanism.
As long as barriers to entry is not too high (ie. it makes sense to take the risk to start a business that provides something similar - usually for a niche) market forces work.
We have the classic empirical chart reproducing microeconomics.
And setting up a pharma plant is also very capital intensive.
Here the obvious barrier to entry is completely artificial. (Which provides an incentive to spend a lot of money on R&D -- though it naturally raises the question of Pareto efficiency.)
You're missing the point, what will happen is this:
1) in things like tax law, registering with city hall, dealings with the DMV, your phone subscription, insurance contract, ... you will find that one of the new fine prints in the contract will be that you're not allowed to use AI to communicate with them.
2) because of how SynthID works (you need the SynthID keys to verify, which are secret. So the only way to find if text is ChatGPT/Google/Anthropic watermarked is to ask ChatGPT/Google/Anthropic), government and large companies can enforce this against you. That is what the watermark is for. To end any insurance claim written by AI with "you're not allowed to submit AI written insurance claims" and refuse it outright there and then.
"Sorry your request was AI watermarked and pursuant to law 234 of 2025/03/11 chapter 3258 paragraph 33 decile 1299 we hereby close it without response"
3) when they reply, however, they use a custom model that also has custom SynthID keys. You will not even be able to tell their responses are AI written, or at least, you won't be able to prove it. You won't be able to enforce any AI-related rights (ie. the right to talk to a human) you have under the law against large companies.
In other words: this is to make sure that all the advantages AI provides are available to deny your unemployment claim, and to Verizon to charge you more, but completely inaccessible TO YOU when you want to change to a cheaper subscription. They can inundate YOU with AI-written requests BUT YOU CAN'T.
Self-hosting helps because it prevents them from verifying if your responses are AI written, because you can generate non-watermarked AI text and so there is a level playing field.
For government, because it's in law or regulations (ministerial decisions in Europe). For large companies "You agreed to it" (you know, like you agreed to allow Verizon to sell your location data to Palantir)
The other 2 questions I don't understand. My point is that the EU AI directive makes this possible. Makes it possible in ONE direction, while prohibiting the other. AI can be used by government and large companies to spam you and deal with you, and can't be used by you without being 100% up front about that to them (ie. enabling refusal)
No. Here is the list of organizations that have the power to make laws in the EU (and JUST the across-the-EU part of that list, within countries, within states, within provinces, within towns there's another list). This is referred to in legal tradition as the "Hierarchy of norms", because there is also a clear order defined.
this seems exactly the usual anti-consumer bullshit that is regulated state-by-state (or sometimes by (lack of) FCC/FTC effort, or by the CFPB that is now a zombie)
however, AI doesn't really influence this. already there's a lot of problem with things like Ticketmaster, Apple's walled garden, abuses of IP law (patent trolls, DMCA trolls), etc.
the insurance industry is a prime example of this. the suffering caused by power imbalance is incomprehensible, and yet there's not enough political will to address this.
sure, it's easily possible that some important aspects of our everyday lives will be worsened by bad AI regulation. but IMHO this is wholly an upstream problem, it's a symptom of bad politics. (a byproduct of the Zip2 to Tesla to "democracy with roman salute characteristics" pipeline.)
that said, obviously the foundation to have any chance of a nonpatological market to exist is that self-hosting has to be legal.
I think his point is hand wavy at best. It presupposes infinite scaling and ignores all the algorithmic efficiency wins that are being discovered. Ironically, many of which are being discovered with autoresearch style workflows, using the very LLMs that his company builds.
The #1 post on HN right now[1] is full of people jubilating about how they can run Qwen 3.8 27B on their > 5 year old GPUs. If that isn't democratization of AI, I don't know what is.
I'm sure he's smart enough to instantaneously realize this too, but as the famous Upton Sinclair quote goes, he won't mention it even if he does.
Which percentage of people have GPUs capable of running Qwen3.8 27B? I am one of those, and for my job I am still resorting to hyperscalers because tasks are completed faster and more accurately that way. Even if we assume that models will no longer improve and we reach a point where everyone can run Fable in their laptop, surely running 1000x Fable agents would give you an advantage.
I think access to compute will matter just as much, if not more, as access to models.
NVIDIA's 3090 was released in September 2020. Apple's M1 was released ~2 months later. Anyone with a 5 year old M1 Mac with 32GB or more RAM can run a 4-bit quantized version of Qwen3.8 27B on their machine. AFAICT, there are ~110m Apple Silicon Macs in the world. I'd wager that at least ~20% of those have enough memory to run this model. And if you account for gamers with NVIDIA and AMD cards, I'd wager that the segment is at least an order of magnitude bigger.
> Even if we assume that models will no longer improve and we reach a point where everyone can run Fable in their laptop, surely running 1000x Fable agents would give you an advantage.
IMO, asymptotic advantages are marginal. At least for coding, we got a glimpse into how much of an advantage it gives (or doesn't) when the Claude Code codebase leaked[1] ~4 months ago. :)
As someone typing this on an M1 Mac with 32 GB of RAM who tried using 3.8 27B (Q4_K_M) yesterday in both LM Studio and llama.cpp, I wouldn't call it particularly usable in terms of token speed. (and that was with `--spec-type draft-mtp` for llama.cpp).
If you want to leave it running with the fans going crazy for 40 mins or overnight or something, fair enough, but otherwise it doesn't seem worth it to me. It's certainly not "interactive", even taking into account the over-thinking it does by default.
The 3.6 (maybe they'll release a 3.8?) MoE model is much more usable (but obviously not as good) on this machine spec.
That's actually an interesting question, and I don't think we have the data to answer it. But we do have the Steam data, and about 7% of Steam users have a GPU that can run it well at a 4-Bit quant (≥24 GB VRAM). About 30% can run a 3-bit quant - I'd say that's just barely usable (≥16 GB VRAM).
I'm not sure whether that's low or high, or how it compares to a general audience.
The very definition of (applied) technology is power amplification. Use a lever, move more weight than you could before, 1 person with the tool now wields the power of 3 without.
Making "tech" a career and a societal goal onto itself, without the adjoining understanding of and deep commitment to ethics and the responsible use of power, is why we're sliding into authoritarian rule by a small circle of techno-oligarchs.
We need less "move fast and break things" and more "plant trees you will not live to see bear fruit."
One thing Dario said is important reflecting on (but from another angle): where's the big deliverable from AI? If AI makes us 10x more productive, the 3 years since its popularization were enough for a product that would have taken 30 years to build without AI, for instance.
I'm an AI advocate but that question makes me feel that AI is simply "very useful" rather than being a historical game changer for humanity.
Anyone saying 10x as a serious claim is clearly using a round number and vibes; however, even if it were so, AI getting popular 3 years ago does not mean what you say.
The earliest used models in this category would easily 10x (what did I just say) the creation of one-off short scripts, but you're not writing a 30-year app out of just a bunch of short scripts.
The METR time horizons graph suggests we can now get 10x (ahem) speedup on solving most coding problems that take us a few hours and about half of problems that take 2 days. Amdahl's law bites: even infinity speedup on half your problems is only 2x overall.
If you let an LLM loose, with a huge budget, what's the biggest artefact it can make before it drowns under the weight of bad decisions? The C compiler and web browser headlines a while back? The maths papers we see now that solve problems which stumped the maths world for decades?
But this is the other side of the same coin: teamwork. One person getting a thing made in twelve months vs a team of a hundred, you can scale up fast with money when you have a proof of concept, and an LLM can make a lot of proofs of a lot of concepts even in free accounts.
LLM-assisted software engineering seems to be very efficient if you have a deterministic target. (bun rewrite from zig to Rust, 100% Node.js compatibility, pnpm compatibility -- https://github.com/oven-sh/bun/pull/38333)
Can only speak for my own project, but can give an example:
After an aquisition earlier this year I got the task of doing an SAP-Integration for the new company, last time I did this 5 years ago it was a 6 month task, but with the experience and skills ive gained since I estimated it would be a 3 month project (with or without AI, most work is just logistics, AI cant help much there).
In those 3 months I was able to not only integrate SAP but also deliver a completely modernised user-facing software for that integration. While I could have written that software myself in a vacuum it would have never been worth it financially, since it would have delayed the launch of the integration by 6+ months. Building the software post-launch of the integration would have easily taken 2.5 years at minimum.
But this is also basically a "spherical cow in a vacuum" scenario, where I was essentially acting as a solo dev, in full operational control of the project, with deep domain knowledge of the topic and an allready fully set up codebase that I knew perfectly while working down ideas I've had in my backlog for 5+ years.
> Building the software post-launch of the integration would have easily taken 2.5 years at minimum.
What you have said is correct, it lets you build software much faster. The question however is: is that software making money for the company? (Not talking about what built but in general)
I think, with AI, companies are saying yes to a lot of things they would have said No to ik say 2020. And as a result realizing “just building it” is not the answer.
Previously your GTM team or Product team would say “If we ship some big project X, we unlock $Y in revenue” but now people are realizing that those projections were really more of a hope. So companies are spending so much more tokens and shipping so many more PRs based on hope but a lot of it just doesn’t turn into meaningful revenue, especially not in short term
First, it's not three years since. The real improvements in programming ability arrived in the last 6-8 months.
Second, the Internet didn't show up much in GDP and similar measures either!
But your point stands. Where are the amazing digital products/stuff? I get that it might take time to arrive as we scale up compute and learn new paradigms. But so much infra already exists (deployment pipeliens, everyone reachable on a smartphone) that we should be seeing something.
> But your point stands. Where are the amazing digital products/stuff? I get that it might take time to arrive as we scale up compute and learn new paradigms. But so much infra already exists (deployment pipeliens, everyone reachable on a smartphone) that we should be seeing something.
I'm definitely seeing indie-sized games that appear to have had significant input from AI, though I'm not sure the balance between AI for coding and AI for assets. My experience attempting this directly suggests that the current level they work at can make very simple games as one-shots, but anything more than trivial will produce outputs only as good as the developer's combined willingness to put in effort tweaking things and taking it all one step at a time, and their taste about what "good" even is.
I'm using spare credits to build and improve an isochrone map renderer, which I otherwise wouldn't have had time for (apart from anything else, I'd have had to become skilled in JS+wasm, somewhat of a pivot from iOS). This also requires taking it all one step at a time, having UX and UI taste.
Having lived through GeoCities since before it was bought by Yahoo!, taste is… well. Most people make things that nobody else actually wants.
I've done amounts of refactoring and fixes and written tooling that just wouldn't have happened before.
I'm not sure what amazing new stuff y'all expect but the amount of technical debt in my projects is actually going down, cause I can finally get good enough test coverage, including E2E/load tests that actually prove whether the software works and scales or doesn't - just last week I diagnosed issues with SeaweedFS failing under concurrent writes when backing Sentry and could swap it out for Garage in a day, caught by a monitoring tool I slopped together that integrates with the Sentry API, no issues since.
The environment around me has gone from drowning in tech/ops debt to sort of swimming and at least holding above water for now (cause nobody will pay for 5x more tokens).
It's also insanely good for prototyping and being able to actually explore various ideas and shoot the bad ones down quickly instead of handwaving and looking at a loaded calendar, alongside being able to address well bounded tasks in parallel, better than human developers can - like I can give 5 GitHub issues to the slop machine and have it fix all of the annoying bugs. Issue with how some data shows up? Just feed it the DB dump and let it find out what's up.
Some projects have gone from around 500 code tests to around 4000, and before anyone says they're meaningless, at least 5% of those have caught real issues and helped a bunch, alongside linters and other tooling (including some tools I wrote myself). I've also written both native utilities and some web platforms for myself, side projects that I never would have gotten around to.
I'm measurably more productive than I've ever been (since I did measure that, looking at my commits over the last 2 years) but also burnt out. Still, it's the kind of burnout that's the consequence of context switching and lots of work, rather than the kind that I had years ago, where I had to manually untangle deeply nested Spring Boot service logic all over the place at like 2 AM cause the made up deadlines were kicking my butt.
In contrast to others, I don't need to move the goalposts - the productivity for me is here and now. Any future models will just make it better, unless we experience model collapse.
Disclaimer: you do need a LOT of code tests and validations, otherwise it all goes to shit. Maybe I'm just extending how much time it will be until it goes to shit for me as well, but go figure. You also have to babysit the models more than anyone would like or should, most of my work usually has 20-60 minutes of planning before dispatching the agent.
What's moving the goalposts? I am very much amazed at what Fable can do. I push its code straight to prod.
But I am just as amazed with how little real life consequence it seems to have! Even software houses were hit more by interest rates than by this magical revolution.
If I couldn't directly observe Fable in action, I wouldn't believe in AI.
I think it's more that it shifts the thought process from "If this is going to take 30 years then we won't bother, because the investment can be spent on things that pay off sooner" to "If we can do this in 3 years then we'll make that investment because that's a good bet."
The way it changes the game is by lowering the cost of making radical bets so we end up trying more moonshots.
Right now we are in the golden age where we do the same and take time off. Employers have not yet fully caught up with the workforce. I can't think of people that are not putting less hours this year for the same salaries.
The questions is what happens when they catch up. They'll cut like 50%+ of the workforce? What happens then to the demand that makes their companies work?
Or an example of MS - their main cost like most software companies are people, especially software devs, which are to be replaced by AI so on the surface they would greatly benefit from it. But their products are centered around helping out people do stuff on the computer. Why would you need that when the AI will do it better and faster directly operating on the data or using e.g. Python?
Who knows. What I tell is that the AI productivity boom is here. Just for a change right now it is not shown on businesses balance sheets, because it is captured by the workforce in non monetary ways.
> If AI makes us 10x more productive, the 3 years since its popularization were enough for a product that would have taken 30 years to build without AI, for instance.
A product still requires a lot of handholding and human thinking, at least if one does not want everyone even throwing a glance at it to immediately be repulsed by the usual AI slop tells.
> I'm an AI advocate but that question makes me feel that AI is simply "very useful" rather than being a historical game changer for humanity.
It absolutely already is a historical game changer on par with the Industrial Revolution when it comes to the amount of jobs destroyed and economies screwed up - and the impact will be even worse in 10+ years as existing seniors retire but no new seniors rise as AI has destroyed entry level career paths.
> It absolutely already is a historical game changer on par with the Industrial Revolution when it comes to the amount of jobs destroyed and economies screwed up
Which economies are already screwed up?
IMO it can’t ever be on par with the Industrial Revolution because AI can only really affect the information economy. Things people do with their hands/bodies have either already been automated or can’t be with current tech. If you’d asked people decades ago they might say no one will ever work in factories by 2026 because they’ll all be automated. It didn’t work out that way. I think AI will go the same way: absolutely game changing to some industries (of which software engineering will be one) but a great many will still survive with less dramatic changes.
If anything it might result in more focus on the human aspects. How many people out there earn their stripes putting together slide decks? In a world where an AI can put together the snazziest presentation you’ve ever seen in a heartbeat it’s going to matter more how you stand at the front of the room and present those slides than it does today.
Not disagreeing with you, but adding to my argument. Despite the human handholding, I feel that 3 years would have been enough for 18 months of thinking about the product plus 18 months where a team could get 3-5 years worth of coding/development done. Yet, we haven't seen anything big yet, like a new Youtube/Instragram, an amazing videogame, a major cure etc. Maybe they are coming, but every day that passes is an indication that the net productivity positive of the technology isn't as big as advertised. In my own personal use of AI, I experienced a 30-50% increase in productivity, but no more than that.
It's only been at most the past year where AI has been unambiguously helpful and not a hindrance. 3 years ago it gave the appearance of being helpful but it tended to be more of a hindrance.
Maybe so, but least for me in my personal coding projects, I'm going through my own task list noticeably faster. I had created this list before getting a ChatGPT Pro subscription and the number of bugs found in AI reviews and rate of closing tasks has certainly increased. I'm not saying my anecdote scales to teams or even other people, but I have no doubt about my personal productivity change. I wouldn't be paying for it otherwise.
IMO the constraint is that AI is still rather anaemic at sustainable green-field projects: It's good at one-off oneshots, and also at refactoring or fixing bugs or adding features to existing projects, where test suites and significant architectural scaffolding already exists, but the more you move away from that, the more wobbly the results get, and the more the human once again becomes the bottleneck, for all the hard work of coming up with all the conceptual scaffolding in the first place. Typing speed rarely was the bottleneck there anyway.
Any argument about regulation in the US which doesn't mention that China and other countries won't follow that regulation should be heavily questioned. Exactly who are you stopping from doing what you don't like?!?
Law abiding US citizens won't be able to run them, but you didn't solve anything other than making sure US citizens pay Sam or Dario.
Why the comment submitted by jacquesm, who posted the link and it's not exactly an anonymous poster, that said "Apologies for linking to X but this is worth reading." has been flagged to death??
I'm more interested in why somebody would suddenly flag that post. It's not spam, bot, offensive or anything else. Maybe you could downvote it if you don't like the take on X - it still would be childish but at least would be proportionate.
It will be interesting to see what kind of progress in biology and medicine they will preview during the fall. He does make a good point that most of the progress so far have not been material in the sense of providing real positive outcomes for ordinary people.
Or even for the HN crowd, when will we see e.g. "Mythos aided research discovers 10 new viable battery technologies"
So the solution to convince the public that these companies aren’t “looking for new ways to screw them over” is to try to go into biomedical research.
I guess it’ll be great for Anthropic to have the cure for cancer, but what’s that gonna mean for people with cancer? Funny how he doesn’t talk about that part.
Called it! He was bound to squeal after the release of QWEN 3.8 in one way or another and here it is. I wasn't even a teenager but I'm getting so much Jobs/Ballmer flashbacks with internet explorer and microsoft office.
Edit: Does anyone else notice the switching between dashes and em-dashes between paragraphs? Tells you a lot about the man, doesn't it.
> Edit: Does anyone else notice the switching between dashes and em-dashes between paragraphs? Tells you a lot about the man, doesn't it.
That he doesn’t know how to use basic word processing, or even agent SKILLS.md or whatever it’s called now. Or maybe it’s just the next generation of vagueposting.
The tech industry cooking up some new way to screw people over for 25 years
The tech execs waking up on a random monday:
> I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.
I can’t help but wonder if this essay was timed in an effort to compete with that bombshell story by WSJ a few days ago. It reports that Dario’s wife, Cami Clark, wields secret influence at Anthropic, had her existence scrubbed from the internet - and also tried to get Jeffrey Epstein to invest in her porn startup. Quite a read!
> Overall my view is that AI is structurally a technology that tends to concentrate power
> Open-weights do help some with this but are nowhere near a sufficient solution
Which is exactly why Anthropic contributes nothing to, and actively pushes for roadblocks and regulations for open weight models. Can't allow any hope to the masses.
Only people with $$$$$ are allowed to touch Fable. Which of course we have have aggressive guardrails for in case you even try to use it for something dangerous like AI model developement. Plus we are going to retain all data submitted to it just in case someone is trying to be sneaky.
> I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.
To paraphrase Sinclair "It's difficult to make a man realize the issues with regulatory capture, if they're to be the benefactor of regulatory capture".
> A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.
Not so sure. After all lack of the latter did help establish the current tech overlord rule.
> This is why Anthropic has always made its policy proposals very carefully. We try very hard to make proposals that disadvantage (slow down) frontier AI companies while advantaging smaller competitors.
> ...
> completely exempt any company below a certain amount of revenue or model training costs from being covered at all
One could argue that "Frontier AI" company know they have nothing to fear from company with less than XM$ revenue, and so their support for this type of regulation is still a way to force regulatory capture. In any case, whatever regulation you support, it's a regulation that you didn't have to handle when you were growing, but that incumbent will have to deal with.
Whatever it is Dario's intention or not does not matter. Capitalism push to consolidation and the eventual regulation that will need to be applied to mitigate the externality from a new industry, will mean that their will only be a handful of "very big" winner. Same story since the beginning of the industrial era. Even if that's not what Dario personally want, it's in the best interest of the shareholders, which will force Anthropic to do everything it can to be one of the big one.
> Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).
The same can be said about a lot of other industry (aviation, oil, chip manufacturing, ...). Every country / union big enough will finance their own champion to try to keep a foot in the industry even if they are not the best.
When people talk about Qwen 3.8 being on a par with Fable, they're really talking about Qwen 3.8 Max aka Qwen3.8-2.4T-A95B. That's a 2.4 trillion parameter Mixture of Experts model with 95B active parameters. You need about 400GB of RAM to run it. No one is running that locally.
The distillations of Qwen 3.8 down to a 27B model are good, but they're not on a par with frontier models.
When Dario talks about open weights not being a solution this is what he means - if you don't have 400GB of VRAM lying around the fact that there's an open model like Qwen3.8-2.4T-A95B doesn't really help much. If we're not regulating how models are available, or making sure access is open, then RAM prices will mean everything concentrates on a few very rich companies.
I don't really understand the argument you're making, but just to add a data point:
DeepSeek V4 Flash 0731 is 167 gigabytes from the developer and as a GGUF with no additional quantization. It limps along on my 192GB M2 Mac from several years ago [0]. This model tests better[1] than Claude Opus 4.6 released in February. That's six months ago - what will be available 6 months from now?
So yeah, enthusiasts aren't going to run frontier models on their gaming machines, but a small office could easily justify the $30k - $100k cost to run something like this at high speed. The small company I worked for routinely spent that kind of money on Dec Alphas twenty five years ago, and that's not accounting for inflation adjustment.
And this is completely discounting the advances smaller models are making. You're right that Qwen 3.8 comes in different sizes. However, Qwen 3.8 27B and Qwen 3.6 27B do run on gaming cards, and they're better than the frontier models from twelve months ago.
I have no idea what will happen in the future, but I wouldn't base my guesses solely on the largest open weight models.
[0] Yes, it's unpleasantly slow (5-8 tok/sec)
[1] Yes, benchmarks should be taken with a lot of salt.
Practically, no, the distill is great. It's fine to use it.
However, if you're having a discussion about access to frontier models, and using Qwen 3.8 as an example of how open weights is a solution, then you should be honest and accurate about what you're talking about. Making an argument like "People can run Qwen 3.8 at home. That shows open weights are great." is a bit disingenuous if you're not also making it clear that you're not talking about Qwen 3.8 Max (or that you have a beast of a PC at home :D ).
I agree 100%. IMO, it all started with ollama misrepresenting the Deepseek R1 distills as Deepseek R1, all for hype and marketing. I've had so many ostensibly technical people telling me: "I tried DeepSeek R1 and it was terrible", and every time when I probe further they'd tried the tiny 1.5B Qwen2.5 distill model that was further brain damaged by ollama's naive RTN quantization[1]. DeepSeek themselves were very forthright about it by naming it DeepSeek-R1-Distill-Qwen-1.5B[2].
I suppose the road to technical hell is paved with marketers and grifters. :)
Exactly. There are common coding tasks that these models can adequately do. They are absolute trash for anything that isn't coding. And even with coding, they are so, so far behind frontier models.
> "I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world."
> wall of text follows
> doesnt proceed to clearly tell us what the actual picture of the world is, then
am i correct in summarizing that the line of reasoning is
- frontier llm access means you are at an economic advantage
- a big risk of this is ongoing wealth concentration
- the "open weights" approach can't solve the problem of wealth and llm access being linked; you need compute too, and compute is expensive, thus "open weights" still favors the wealthy
- instead we need "objective and fair institutional processes", as this will allow small labs cook up their stuff while frontier labs get regulated
why would i care about what these smaller players do, if economic advantage = frontier model access? also, the reasoning around "why bother with open weights because compute is expensive too" seems like the kind of logic a motivated 12 year old could work their way around in 30 seconds. things are not either/or dario, you said so yourself.
seems like a 400 word corpo misdirection essay. par for the course.
> Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).
If only money thirsty capitals didn’t drain the HBM/DRAM/NAND capacity to rush to build out all the data centers so that they can make money off of inference, driving up memory prices like mad man leaving consumers with not only no chance of local inference setup, but also higher price of consumer electronics. Of course it has nothing to do with regulation, more to do with the greed.
> I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.
It is so tragicomic to see people whose lives are built around companies coming so close to realizing that everything they do is bad, and then at the last minute swerving aside to convince themselves that no, if they just do more of it and somehow do it "better", then it will all be okay. The reason people don't trust companies and believe they are cooking up some new way to screw them over is because that is what they are doing. If Dario or anyone else really wanted to dispel that perception there's an easy way: do a total 180 and start fighting against everything you've been pushing. But none of them will do that because they still fundamentally believe that what they are doing is good, and are unable to see that fundamentally it's bad.
I think that Dario is intelligent, incisive, has good judgement, and is well-intentioned. I hope he continues to wield influence. I think Sam is also most of those things, but I think power has a One Ring-like effect on him. My most contrarian view is that I think Elon’s problems are primarily appalling judgement when it comes to areas outside of tech, and has the emotional regulation of a child, although somewhat incredibly I believe he is fundamentally well intentioned. I don’t think any of those three really want tech feudalism, although I think the current US administration would press a “turn us into Russia” button the second they caught sight of it, which is to say, I think they have the worst of intentions.
> I think Sam is also most of those things, but I think power has a One Ring-like effect on him.
That's putting it mildly, very mildly. The guy all but wrecked the entire world's DRAM supply chain by using negotiation tactics that, if there were any semblance of regulatory authority left in the US, would lead to criminal charges for market manipulation.
Arrest Sam Altman, he is one of the most dangerous people on the planet. Not a single shred of respect for the 99% and the consequences his actions have on them. (And similar things apply to the other controversial figure you mentioned, Elon Musk actively enjoys hurting people by the looks of it.)
Dario's messaging has been extremely confusing and HE'S RESPONSIBLE for further increasing the negative sentiment of the public towards AI.
He's clearly focusing on being on the news to provoke emotions on people, preparing for the Anthropic big blockbuster IPO.
Dario, Sam Altman and others should instead be praying every night that this will take us to the Singularity VERY SOON, because if it doesn't, the entire country will blame them for absolutely destroying their retirement and the US' economy once this AI Datacenter bubble pops.
> I think it is fundamentally a crisis of trust. I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.
[...]
> I don’t think that a glitzy marketing campaign with a positive spin (which some have advocated that Anthropic do) is the way to win back that trust — at this point, saying that AI will cure cancer is more a cliche than it is inspiring, and most people think it is deceptive. The thing that will work is actually curing cancer. I think by far the most accurate criticism of AI companies including Anthropic is that we haven’t yet delivered on our big promises to benefit the world. That is totally on us, and I think it’s the criticism you should be making, instead of all this stuff about messaging and marketing.
> We are however doing our best to fix this: Anthropic is ramping up its efforts very quickly in biology and medicine, and we hope to have incredible results in the coming years and some early glimmers in the coming months. When we’ve actually accomplished something real, the whole world will hear about it, as loudly as possible, you have my word on that.
Turns out "winning back trust" doesn't have anything to do with any actual concerns people may have re: employment, electricity prices, stock market bubble, intellectual property, scams, cybersecurity, environmental issues etc. Rather we'll just do all of that even harder and the miracles ("curing cancer", lol) we've so far failed to deliver are bound to arrive in short order!
> The thing that will work is actually curing cancer.
This is honestly hilarious. Dario must think people are stupid.
First, cancer research charities are some of the most well funded on this planet. They all obtain donations on the basis of "together, we will cure cancer" messaging. They all fund the best science they can.
Second, the human body is a complicated thing. You can feed your fancy LLM as many textbooks and academic papers as you like, but the reality on the hospital ward will always be different. Why do you think student doctors have to spend so many years "doing the rounds" Dario ? They are all academically smart, they are all capable of memorizing text books ... but there is no substitute for seeing and doing the reality.
I barely trust Claude to write code, let alone find a cure to cancer.
> The thing that will work is actually curing cancer.
I think typical bubble behavior the leaders have set up the whole promise to fail. Everyone is expecting some faux super intelligence to come and find a cancer solution everyone else missed. However, it is just as likely that vanilla current LLM's will create enough of a productivity boost for back office automations in research heavy hospitals to create the space for regular humans to create these breakthroughs, but LLM's won't be able to claim that for themselves and inevitably “fail”.
Savings from back office automation go entirely into administrative bloat. I find it even less likely such efficiency would lead to breakthroughs than I do LLMs becoming super intelligent and doing it themselves, which I don’t find likely at all.
I mean they did solve the protein folding problem did they? NOt too far fetch to think it can automate cancer research or problem finding in some way.
The field of AI did very well in protein folding, but it was completely different from LLMs.
> I mean they did solve the protein folding problem did they?
They made huge progress, but I would say that the vast majority of work on this problem was designing the harness for the model. That's a lot of work for each and every domain.
There is already (partial) automation there, for example for finding new drug candidates.
Isn't that so far only static folding?
As usual with LLM era craze, finding candidates was never the bottleneck. And, so far, it failed to actually provide any tangible results. [0]
[0] https://www.science.org/content/blog-post/so-how-ai-drug-dis...
The effective altruist and longtermist crowd are into it as a religion. They would happily sacrifice anything human to satisfy their dreams of AI. They care way more about what the AI needs than their what their fellow humans need. We will end up with an AGI having better living and working conditions than humans and the whole AI industry will be clapping
It's an ideological thing. (We just had a conversation about ideology actually, remember?)
Capitalism has always been a tension between capitalists, who want maximum return on capital and workers, who want pesky things like a living wage, sick leave or safe work conditions.
For the longest time, the only power labor has had to get those things was the power of collective bargaining. No agreement with your workers meant no production happened.
Now, there's finally a chance at salvation. The capitalist Messiah is AGI and it will finally deliver them from those annoying laborers.
This ideological bent is why they're putting everything they have into AI. It's why VCs and their fellow capitalists are going so crazy.
They see it as a way to finally solve the contradictions of their ideology, but in reality it'd only create a new one: If everyone's out of a job, who will buy their products?
I wish I was smart enough to credibly promise capitalists they could replace all their workers.
It doesn't matter that I can't do it. They'll pay me lots of money if they think I can.
I think it's more reasonable than it sounds on the surface. People's jobs, the bubble, etc are societal level problems and well beyond what Anthropic could even hope to influence on their own.
Curing cancer sounds insane, but it's also a research problem, not a societal level coordination problem. And one AI has already proved to help with breakthroughs (alphafold). IMO it makes sense for them to shoot for something like that as proof of AI's beneficial sides.
>People's jobs, the bubble, etc are societal level problems and well beyond what Anthropic could even hope to influence on their own.
This is not true. There are two companies at the center of AI direction: Anthropic and OpenAI. If there were anyone on the plant who has the ability to influence our direction then it would be Dario Amodei.
Acknowledging the grievances is a good step, but it's not enough. There needs to be a clear explanation of actions to address them, a plan to enact those actions, and commitments with consequences in failure of those actions. Tell people how you're going to make them more employable and effective and needed. Tell people how your datacenters will be carbon neutral. Tell people how financial actions resulting in a frothy market will be coming to an end. He and Sam Altman alone have this power and their inaction says everything we need to know about their intent.
> Curing cancer sounds insane, but it's also a research problem,
I just don't buy the AI labs approach to this stuff. Like, unless we can basically simulate the entirety of human biology, I don't really see how LLMs can make progress here. Maths is different as it doesn't require a real-world interface, and programming already (by definition) can be simulated on a computer.
Without that, I can't see much (if any) progress being made on domains like biology.
I'm a noob on this topic, but I think drug discovery is more amenable to this structurally than other problems. Simulating biology is what we were doing with protein folding before Alphafold, and the search space was far too large to find stuff in reasonable timeframes. Alphafold showed that you could take a physical process and make a neural net clever enough to learn just enough structure that it starts finding things we might care about, and still physically accurate, much faster.
Drug discovery is similar AFAIK. The space of possibilities is even larger than protein folding, but it's structurally similar enough that I think AI will help to make progress on the discovery side. Actually getting the drug tested and approved is another matter though for sure.
There are two relevant problems here, I think.
For one, the question for Anthropic is whether LLMs, specifically, not AI techniques more generally, can help significantly with cancer research. And here, all experience so far is that LLMs only really work when they can easily automatically verify their own outputs and self correct - such as in math (using automatic proof verifiers) or programming (using compilers and unit tests).
The second problem is that biological research speed is highly dependent on slow biological processes, such as cultures and long term studies. In programming, if an LLM could provide excellent insights and research suggestions 100x faster than a human, it would speed up the work roughly 100x. But in biology, it would only speed up the total work by a small amount - as any insight, even if absolutely brilliant and spot on, would still require months and years of actual experimentation.
The real issue is that biology needs actual experiments done in the physical world, which isn't nice and orderly and well behaved and easily loadable onto a 19" rectangular box.
On the other hand, AI means actual experiments done in the physical world but coordinated by an entity that never sleeps and never gets depressed and can multiply itself manifold and always comes up with new ideas.
That fanatical tireless entity will need access to human tissue.
> People's jobs, the bubble, etc are societal level problems and well beyond what Anthropic could even hope to influence on their own.
Anthropic is literally creating the bubble. It is not beyond their scope of influence, it is literally what they are consciously achieving.
As for peoples jobs, same actually applies. Anthropic is selling itself on dream of replacing jobs, even or especially where they are well aware AI does not perform that well. They are actively trying to replace people quickly before management notices it does not work well.
And also, they can influence how much their data centers contribute to global warming.
This is actually crazy, because biochemistry is where LLMs are weakest. It's one of science's most fuzzy and unpredictable domains, in general, and a lot of published works which an LLM might take at face value might be unreplicable or subtly flawed. (See e.g. the entire history of Alzheimer's.) It is also where real-world lab and clinical trial work is most important.
So Anthropic are going to spin up a medicinal chemistry lab and start mouse experiments?
I mean, it would be nice, but I don't think that the guys at Anthropic know what they're talking about here, or what they might be getting themselves into. (If indeed this is more than just PR.)
See also this article by a genomics PhD about how intelligence is not the bottleneck in medical research.
https://www.writingruxandrabio.com/p/intelligence-is-not-the...
AI bros consistently handwave away any notions of the material world imposing limits, because (spoiler!) it's a religion where you start from the miracle and work backwards. "Oh, we'll just set up fully automated robotic research labs everywhere!" Where will you get the raw materials? Oh we'll mine the asteroids! How will you find the energy to get there? Oh, the AI will invent new physics! (again without needing to run experiments). Et.c.
Didn't and don't mean to disparage anyone's medical struggles, but "a cure for cancer" is a well-worn strawman. One that's been achieved for the low hanging fruit, the higher ones are seeing steady progress (already before LLM chatbots, even!), the bottleneck isn't "intelligence" and to the extent there are socioeconomic (access to screening, treatment) or environmental/lifestyle factors involved the AI boom is likely just making things worse!
I'm sure Dario knows this, and it's anyway too pedestrian compared to the usual list of fruits of ASI. The text probably originally read "nanobots eating you alive and uploading to the cloud" or something, but they figured that wouldn't go over with the intended audience. "What do the peasants care about? Oh I know! Curing cancer!"
> AI bros consistently handwave away any notions of the material world imposing limits, because (spoiler!) it's a religion where you start from the miracle and work backwards.
How else are you going to work towards a dream? It's how Elon Musk managed to get reusable rockets when everyone said it's unfeasible.
Granted, most ideas don't work, this is why you need testing, but I am a bit surprised to see this attitude on hacker news.
If they said that their dream is to solve physics, I'd say "that's cool and I can't wait to see how it turns out." LLMs have gotten so good at mathematics, and are so adroit at analyzing vast quantities of data, that they might have a really good shot.
But they said that they want to cure cancer. That's far outside the core competencies of any LLM, and it requires a lot of real-world wetwork with liquids, chemicals, cell line experiments, animal experiments, etc. They can't merely analyze existing data -- they'd need to generate vast amounts of new data, which isn't really the case in physics. And then regulatory approvals and so forth.
"We're working on curing cancer" sounds more like a poor PR attempt than an actual effort, though I'd love to be wrong.
Fair enough, but taking the history-ending outcome as given so externalities don't matter doesn't seem like a great approach.
rapture but for tech people
Not really. The problem is that the trust is violated here. What people fear is that government and large companies are going to use AI against them, and make the reverse impossible. Sadly, this is exactly what is happening.
And that's exactly what EU's AI regulation does. Of course, AI is being blamed for this happening, despite that last time I checked every EU parlementarian, every EU commission member is flesh and blood, the fact that it is 100% intentional government policy, humans, that are doing this to you using very un-AI methods.
You see, in this regulation, governments are allowed to use AI, and to approve uses of AI. So there is absolutely nothing in EU's AI directive at all that prevents government and large companies, with approval, from answering the phone, and all "support requests" with AI, and not with people. Also "if you agree" (just like you've agreed to selling your location data in your cell contract)
The watermarking algorithm itself has a VERY specific property that should have set off everyone's alarm bells: it is NOT the case that you can "check text for AI watermark". What it technically mandates is that if you provide access to a model, those people should be able to check if text is generated by that model. Not by any other model. It is NOT a general "is this AI?" check. So let's analyze how this works in 2 specific cases:
You get a mail from the government about taxes. You want to check if that text is generated by an AI and whether you even should ask to talk to a human. So you've got a piece of text, or an audio recording. Can you check if it's AI or not? NO YOU CAN'T. SynthID requires access to the model keys, and the legislation only mandates checks if you've got access to the model.
The government or some large company, however, puts in it's regulations that you're not allowed to use AI to communicate with them (we all know this is coming), and they get a piece of ChatGPT text (or from any public model) from you. They would like to enforce that they won't talk to AI can they do this? YES, since they have access to the model too, SynthID allows for this usecase AND the AI regulation makes this mandatory.
Is this in any way a problem with AI? No. Can anything be done about this by "fighting AI"? Well let's see ... will government institutions be forced to stop worsening services further if ChatGPT etc become inaccessible or obviously excluded by government and large company services? No.
What is being outlawed in the EU, in other words, is exactly ONE thing: that you use AI to help you in your dealings with governments and large companies. To assert your rights, to resolve issues faster, to make them wait rather than you, to ... THAT is being outlawed, nothing else. Only that YOU get help from AI.
This legislation will make it a big advantage to have access to "your own" model, because you'll have the SynthID keys, and the law behind you if you do that. And I don't mean local model, I mean your own model with your own SynthID keys, ie. not Google, not ChatGPT, not Claude, not Grok, and if the EU can make it happen, not Chinese models either. Finetuned models, however, will not have SynthID and so you'll never be able to prove text generated by them is AI, and you will not effectively have the power to refuse contracts or laws that specify they can use AI anyway.
It's the EU government making it clear they are using AI to avoid even having to talk to you at all while disallowing the use of AI by mere citizens to assert their rights, against them, or anything they have interests in, like phone or electricity companies.
>The watermarking algorithm itself has a VERY specific property that should have set off everyone's alarm bells: it is NOT the case that you can "check text for AI watermark". What it technically mandates is that if you provide access to a model, those people should be able to check if text is generated by that model. Not by any other model. It is NOT a general "is this AI?"
Where do I find that in the Act (or code of practice etc.)? IANAL, but Article 50 reads different to me but if there is a comment or guide how to read it - also fair enough.
It sounds like anti EU propaganda.
The only factual point you made is solved on February second. All Model providers have to provide offline tooling to check for invisible watermarking.
Yeah, it's particularly common when the EU does stuff.
There is nothing in the AI act that requires offline tooling, nor is there anything about requiring the passwords and/or API to work for every model, only for the people you are providing the model to.
“The detection solution may be made available in the Union as one or more of the following: (i) a public, ideally standardised, specification allowing any third party to implement a detection mechanism; (ii) a piece of software (e.g., a standalone executable or library); (iii) a cloud-based service accessible to users in the Union through an API.”
Option 3: "a cloud-based service".
And accessible "to users". NOT to the public.
Note: this is not the actual law, it is the code of practice, ie. guidance for model providers. The start of that document clearly states that you can comply with the law in other ways if you want, you'll just have to justify yourself. So you have a choice to not even do this.
The actual law is here:
https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-5...
As to who decides the law clearly states who decides if someone is legal, like in most EU legislation. It's not the courts, it's "National market surveillance authorities designated by each EU Member State", and there is an EU office as well (and it is explicitly stated that they are not allowed to override each other). So every EU country has the right to provide exceptions to the law, just like they do for the GPDR. You do not have any rights under this legislation as an individual. Only these "National market surveillance authorities" get rights under this legislation.
Secondary: article 7 additionally gives the EU commission the power to declare any code of conduct they want that declares what compliance with the AI act actually means.
(and, of course, this is yet another attempt at declaring math illegal. The only way to actually enforce this legislation is for all models to comply with this, all over the internet. Obviously the EU does not remotely have the power to make that happen)
While there is a large element of paranoia here, we're already seeing serious problems in e.g. planning consultations that people can just spam them with AI. AI makes a lot of traditional social processes stop working properly.
> puts in it's regulations that you're not allowed to use AI to communicate with them (we all know this is coming),
.. do we?
> You get a mail from the government about taxes. You want to check if that text is generated by an AI
No, actually, I want to know whether it's correct and whether it's legally binding. Using AI makes it less likely to be correct, sure, but ultimately whether it's human, spreadsheet, or AI the important thing is the legal right to correct process.
I genuinely think Dario is a well intentioned, intelligent dude, but I think him and Anthropic have a huge PR problem and are really out of touch with how they're perceived.
Anthropic in particular has developed this almost Orwellian like veil of condescending rhetoric that on the surface suggests they're looking out for you while underneath they're taking actions that suggest they do not trust you, Mr/Mrs Ordinary Person. All the safety rhetoric, never supporting open weight models, the lockdowns on harnesses outside claude code, etc.
Anthropic if you care about public good, do something to empower people. Release an OSS model. Open source Claude Code. Open source some inference tooling or something. Just give people anything except your words.
> I genuinely think Dario is a well intentioned, intelligent dude, but I think him and Anthropic have a huge PR problem
Is this what kids call rage bait these days?
No big tech founder CEO ever was or ever will be well intentioned.
And maybe don't allow the military to use your AI WHEN YOU CAN'T TELL WHETHER IT WAS INVOLVED IN THE KILLING OF HUNDREDS OF SCHOOL CHILDREN
Sorry but this sounds like such a populist claim. They probably also write their war plans on Microsoft Word, and drive Fords to the office! Not to mention the simple fact that there are no wars without mistakes. I'm not justifying this specific war, but as a service provider you do not get to approve or reject specific operations.
> Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers).
This is actually a good point. Open models are getting better and better, some of them might even be useful in consumer hardware now. But if AI performance is still correlated with compute power, then no doubt power will remain with the people owning the chips.
Swap that point about AI with "electricity". Everything runs on electricity it's "a technology that tends to concentrate power" (no pun intended). The electricity providers must be too powerful... But somehow electricity providers aren't that powerful. Unless there is no competition in sight...
The point "AI is structurally a technology that tends to concentrate power" is not that correct. They need this statement to be true, otherwise no way to justify the trillion evaluations.
Swap that to "oil" and everyone goes "well, yes, obviously, oil is so powerful that people start wars over it".
Iran has hit the Amazon data center in Bahrain. The Ukranian deep drone bombing campaign has hit refineries, but also Wildberries, the "Russian Amazon" warehouses. The US campaign against Iran now, Iraq and Serbia previously, targeted power infrastructure. It will obviously be a target in the next war, and AI goes on that list too.
Every time the US completes an AI data center, someone in the Chinese nuclear command updates their target priority list. And vice versa.
>Every time the US completes an AI data center, someone in the Chinese nuclear command updates their target priority list
You might have stumbled on the way to convince people that more data centers are a good thing
> The point "AI is structurally a technology that tends to concentrate power" is not that correct. They need this statement to be true, otherwise no way to justify the trillion evaluations.
How is it not true?
The biggest companies in the world are, quite obviously (just look at the numbers) going to be AI companies. The most powerful governments in the world are quite clearly going to be those that are close to (or in control of) AI companies. People have a very hard time seeing the second order effects of the control of intelligence - we need to fix that.
but its difficult to use more electricity directly to get more utility. There are some rare ways - things like PtX systems - but even those are linear at best. AI is concentrating power in the sense that nobody wants to use the twentieth smartest AI - so if you can use the most compute, you get an outsized share of the rewards (in terms of paying customers) - over and above your share of the compute. There are other business sectors where similar dynamics exist - where the biggest capital tends to win - a sort of natural monopoly type situation, its nothing about hyperscaling or whatever.
Competition is only one of the ways to keep a corporation aligned. In the case of France, having a state monopoly on electricity even worked very well until we broke it for the sake of competition.
Trump already banned solar and offshore wind, in order to concentrate power (literally) with the oil companies.
Except that's a bad analogy: electricity is the means by which to do work.
The outputs of models in AI datacenters are the work.
The electricity company does not get the entirety of my useful input to the work as a result of me using electricity, but an AI company does.
Current AI is at least a few orders of magnitude less efficient than it could be. At some point the labs put too much work and money into transformers and nearly abandoned fundamental research (in both ML and hardware). There are tons of low hanging fruits in efficiency but you'll have to redo everything from scratch so nobody bothers. Which is also pretty convenient and lets people like Dario Amodei speak about "natural concentrations of power".
The Chinese labs are picking up on the low hanging fruits on efficiency, and no, you do not need to abandon transformers, you just need to push them closer to the more computationally efficient architectures of the past. OpenAI and Google seem to be trying a few things too.
Anthropic clearly are not though, and to call their operations wasteful is an understatement.
Arguably the question is whether it's economically feasible to self-host something similar.
Are you self-hosting Google or Bing? No, but we have quite a huge ecosystem of full-text search tools with PageRank, with options to scale to almost Google scale (if you have the money). After all LLM training starts with the same crawl mechanism.
As long as barriers to entry is not too high (ie. it makes sense to take the risk to start a business that provides something similar - usually for a niche) market forces work.
We have the classic empirical chart reproducing microeconomics.
https://www.fda.gov/about-fda/center-drug-evaluation-and-res...
And setting up a pharma plant is also very capital intensive.
Here the obvious barrier to entry is completely artificial. (Which provides an incentive to spend a lot of money on R&D -- though it naturally raises the question of Pareto efficiency.)
You're missing the point, what will happen is this:
1) in things like tax law, registering with city hall, dealings with the DMV, your phone subscription, insurance contract, ... you will find that one of the new fine prints in the contract will be that you're not allowed to use AI to communicate with them.
2) because of how SynthID works (you need the SynthID keys to verify, which are secret. So the only way to find if text is ChatGPT/Google/Anthropic watermarked is to ask ChatGPT/Google/Anthropic), government and large companies can enforce this against you. That is what the watermark is for. To end any insurance claim written by AI with "you're not allowed to submit AI written insurance claims" and refuse it outright there and then.
"Sorry your request was AI watermarked and pursuant to law 234 of 2025/03/11 chapter 3258 paragraph 33 decile 1299 we hereby close it without response"
3) when they reply, however, they use a custom model that also has custom SynthID keys. You will not even be able to tell their responses are AI written, or at least, you won't be able to prove it. You won't be able to enforce any AI-related rights (ie. the right to talk to a human) you have under the law against large companies.
In other words: this is to make sure that all the advantages AI provides are available to deny your unemployment claim, and to Verizon to charge you more, but completely inaccessible TO YOU when you want to change to a cheaper subscription. They can inundate YOU with AI-written requests BUT YOU CAN'T.
Self-hosting helps because it prevents them from verifying if your responses are AI written, because you can generate non-watermarked AI text and so there is a level playing field.
1) Would that fine print be binding?
2) What actual law/regulation would that currently be that would be used for such an outright refusal?
3) Would that comply with current regulation?
Have you ever dealt with a government or government adjacent company?
It'll be spec'd out so that it's cheaper to bend over and take it than take it to court and prove them wrong.
> 1) Would that fine print be binding?
For government, because it's in law or regulations (ministerial decisions in Europe). For large companies "You agreed to it" (you know, like you agreed to allow Verizon to sell your location data to Palantir)
The other 2 questions I don't understand. My point is that the EU AI directive makes this possible. Makes it possible in ONE direction, while prohibiting the other. AI can be used by government and large companies to spam you and deal with you, and can't be used by you without being 100% up front about that to them (ie. enabling refusal)
1) Laws need parliament, no?
Just agreeing to it isn't necessarily enough at least in some countries.
2) What laws and sections specifically makes that possible? Are there examples of that happening?
3) Where can I find that interpretation of article 50(?)? Some other article? (To the extent that things would need to be labelled/watermarked etc.)
1) Laws need parliament, no?
No. Here is the list of organizations that have the power to make laws in the EU (and JUST the across-the-EU part of that list, within countries, within states, within provinces, within towns there's another list). This is referred to in legal tradition as the "Hierarchy of norms", because there is also a clear order defined.
https://eur-lex.europa.eu/EN/legal-content/glossary/eu-hiera...
I think you need to be more precise then on what you mean by "law" and differentiate there and also by what authorises that.
this seems exactly the usual anti-consumer bullshit that is regulated state-by-state (or sometimes by (lack of) FCC/FTC effort, or by the CFPB that is now a zombie)
however, AI doesn't really influence this. already there's a lot of problem with things like Ticketmaster, Apple's walled garden, abuses of IP law (patent trolls, DMCA trolls), etc.
the insurance industry is a prime example of this. the suffering caused by power imbalance is incomprehensible, and yet there's not enough political will to address this.
sure, it's easily possible that some important aspects of our everyday lives will be worsened by bad AI regulation. but IMHO this is wholly an upstream problem, it's a symptom of bad politics. (a byproduct of the Zip2 to Tesla to "democracy with roman salute characteristics" pipeline.)
that said, obviously the foundation to have any chance of a nonpatological market to exist is that self-hosting has to be legal.
I think his point is hand wavy at best. It presupposes infinite scaling and ignores all the algorithmic efficiency wins that are being discovered. Ironically, many of which are being discovered with autoresearch style workflows, using the very LLMs that his company builds.
The #1 post on HN right now[1] is full of people jubilating about how they can run Qwen 3.8 27B on their > 5 year old GPUs. If that isn't democratization of AI, I don't know what is.
I'm sure he's smart enough to instantaneously realize this too, but as the famous Upton Sinclair quote goes, he won't mention it even if he does.
[1]: https://news.ycombinator.com/item?id=49324985
Which percentage of people have GPUs capable of running Qwen3.8 27B? I am one of those, and for my job I am still resorting to hyperscalers because tasks are completed faster and more accurately that way. Even if we assume that models will no longer improve and we reach a point where everyone can run Fable in their laptop, surely running 1000x Fable agents would give you an advantage.
I think access to compute will matter just as much, if not more, as access to models.
NVIDIA's 3090 was released in September 2020. Apple's M1 was released ~2 months later. Anyone with a 5 year old M1 Mac with 32GB or more RAM can run a 4-bit quantized version of Qwen3.8 27B on their machine. AFAICT, there are ~110m Apple Silicon Macs in the world. I'd wager that at least ~20% of those have enough memory to run this model. And if you account for gamers with NVIDIA and AMD cards, I'd wager that the segment is at least an order of magnitude bigger.
> Even if we assume that models will no longer improve and we reach a point where everyone can run Fable in their laptop, surely running 1000x Fable agents would give you an advantage.
IMO, asymptotic advantages are marginal. At least for coding, we got a glimpse into how much of an advantage it gives (or doesn't) when the Claude Code codebase leaked[1] ~4 months ago. :)
[1]: https://news.ycombinator.com/item?id=47586778
As someone typing this on an M1 Mac with 32 GB of RAM who tried using 3.8 27B (Q4_K_M) yesterday in both LM Studio and llama.cpp, I wouldn't call it particularly usable in terms of token speed. (and that was with `--spec-type draft-mtp` for llama.cpp).
If you want to leave it running with the fans going crazy for 40 mins or overnight or something, fair enough, but otherwise it doesn't seem worth it to me. It's certainly not "interactive", even taking into account the over-thinking it does by default.
The 3.6 (maybe they'll release a 3.8?) MoE model is much more usable (but obviously not as good) on this machine spec.
That's actually an interesting question, and I don't think we have the data to answer it. But we do have the Steam data, and about 7% of Steam users have a GPU that can run it well at a 4-Bit quant (≥24 GB VRAM). About 30% can run a 3-bit quant - I'd say that's just barely usable (≥16 GB VRAM).
I'm not sure whether that's low or high, or how it compares to a general audience.
Physical constraints like time and compute make it so that AI does not concentrate power due to scaling laws.
The very definition of (applied) technology is power amplification. Use a lever, move more weight than you could before, 1 person with the tool now wields the power of 3 without.
Making "tech" a career and a societal goal onto itself, without the adjoining understanding of and deep commitment to ethics and the responsible use of power, is why we're sliding into authoritarian rule by a small circle of techno-oligarchs.
We need less "move fast and break things" and more "plant trees you will not live to see bear fruit."
One thing Dario said is important reflecting on (but from another angle): where's the big deliverable from AI? If AI makes us 10x more productive, the 3 years since its popularization were enough for a product that would have taken 30 years to build without AI, for instance.
I'm an AI advocate but that question makes me feel that AI is simply "very useful" rather than being a historical game changer for humanity.
Anyone saying 10x as a serious claim is clearly using a round number and vibes; however, even if it were so, AI getting popular 3 years ago does not mean what you say.
The earliest used models in this category would easily 10x (what did I just say) the creation of one-off short scripts, but you're not writing a 30-year app out of just a bunch of short scripts.
The METR time horizons graph suggests we can now get 10x (ahem) speedup on solving most coding problems that take us a few hours and about half of problems that take 2 days. Amdahl's law bites: even infinity speedup on half your problems is only 2x overall.
If you let an LLM loose, with a huge budget, what's the biggest artefact it can make before it drowns under the weight of bad decisions? The C compiler and web browser headlines a while back? The maths papers we see now that solve problems which stumped the maths world for decades?
But this is the other side of the same coin: teamwork. One person getting a thing made in twelve months vs a team of a hundred, you can scale up fast with money when you have a proof of concept, and an LLM can make a lot of proofs of a lot of concepts even in free accounts.
LLM-assisted software engineering seems to be very efficient if you have a deterministic target. (bun rewrite from zig to Rust, 100% Node.js compatibility, pnpm compatibility -- https://github.com/oven-sh/bun/pull/38333)
My pushback on that would be that all of the "engineering" concering a rewrite was already done the first time around. It's a glorified translation.
I do agree that the port would've taken a lot longer without LLMs though.
Can only speak for my own project, but can give an example:
After an aquisition earlier this year I got the task of doing an SAP-Integration for the new company, last time I did this 5 years ago it was a 6 month task, but with the experience and skills ive gained since I estimated it would be a 3 month project (with or without AI, most work is just logistics, AI cant help much there).
In those 3 months I was able to not only integrate SAP but also deliver a completely modernised user-facing software for that integration. While I could have written that software myself in a vacuum it would have never been worth it financially, since it would have delayed the launch of the integration by 6+ months. Building the software post-launch of the integration would have easily taken 2.5 years at minimum.
But this is also basically a "spherical cow in a vacuum" scenario, where I was essentially acting as a solo dev, in full operational control of the project, with deep domain knowledge of the topic and an allready fully set up codebase that I knew perfectly while working down ideas I've had in my backlog for 5+ years.
> Building the software post-launch of the integration would have easily taken 2.5 years at minimum.
What you have said is correct, it lets you build software much faster. The question however is: is that software making money for the company? (Not talking about what built but in general)
I think, with AI, companies are saying yes to a lot of things they would have said No to ik say 2020. And as a result realizing “just building it” is not the answer.
Previously your GTM team or Product team would say “If we ship some big project X, we unlock $Y in revenue” but now people are realizing that those projections were really more of a hope. So companies are spending so much more tokens and shipping so many more PRs based on hope but a lot of it just doesn’t turn into meaningful revenue, especially not in short term
First, it's not three years since. The real improvements in programming ability arrived in the last 6-8 months.
Second, the Internet didn't show up much in GDP and similar measures either!
But your point stands. Where are the amazing digital products/stuff? I get that it might take time to arrive as we scale up compute and learn new paradigms. But so much infra already exists (deployment pipeliens, everyone reachable on a smartphone) that we should be seeing something.
> But your point stands. Where are the amazing digital products/stuff? I get that it might take time to arrive as we scale up compute and learn new paradigms. But so much infra already exists (deployment pipeliens, everyone reachable on a smartphone) that we should be seeing something.
I'm definitely seeing indie-sized games that appear to have had significant input from AI, though I'm not sure the balance between AI for coding and AI for assets. My experience attempting this directly suggests that the current level they work at can make very simple games as one-shots, but anything more than trivial will produce outputs only as good as the developer's combined willingness to put in effort tweaking things and taking it all one step at a time, and their taste about what "good" even is.
I'm using spare credits to build and improve an isochrone map renderer, which I otherwise wouldn't have had time for (apart from anything else, I'd have had to become skilled in JS+wasm, somewhat of a pivot from iOS). This also requires taking it all one step at a time, having UX and UI taste.
Having lived through GeoCities since before it was bought by Yahoo!, taste is… well. Most people make things that nobody else actually wants.
> Where are the amazing digital products/stuff?
I've done amounts of refactoring and fixes and written tooling that just wouldn't have happened before.
I'm not sure what amazing new stuff y'all expect but the amount of technical debt in my projects is actually going down, cause I can finally get good enough test coverage, including E2E/load tests that actually prove whether the software works and scales or doesn't - just last week I diagnosed issues with SeaweedFS failing under concurrent writes when backing Sentry and could swap it out for Garage in a day, caught by a monitoring tool I slopped together that integrates with the Sentry API, no issues since.
The environment around me has gone from drowning in tech/ops debt to sort of swimming and at least holding above water for now (cause nobody will pay for 5x more tokens).
It's also insanely good for prototyping and being able to actually explore various ideas and shoot the bad ones down quickly instead of handwaving and looking at a loaded calendar, alongside being able to address well bounded tasks in parallel, better than human developers can - like I can give 5 GitHub issues to the slop machine and have it fix all of the annoying bugs. Issue with how some data shows up? Just feed it the DB dump and let it find out what's up.
Some projects have gone from around 500 code tests to around 4000, and before anyone says they're meaningless, at least 5% of those have caught real issues and helped a bunch, alongside linters and other tooling (including some tools I wrote myself). I've also written both native utilities and some web platforms for myself, side projects that I never would have gotten around to.
I'm measurably more productive than I've ever been (since I did measure that, looking at my commits over the last 2 years) but also burnt out. Still, it's the kind of burnout that's the consequence of context switching and lots of work, rather than the kind that I had years ago, where I had to manually untangle deeply nested Spring Boot service logic all over the place at like 2 AM cause the made up deadlines were kicking my butt.
In contrast to others, I don't need to move the goalposts - the productivity for me is here and now. Any future models will just make it better, unless we experience model collapse.
Disclaimer: you do need a LOT of code tests and validations, otherwise it all goes to shit. Maybe I'm just extending how much time it will be until it goes to shit for me as well, but go figure. You also have to babysit the models more than anyone would like or should, most of my work usually has 20-60 minutes of planning before dispatching the agent.
I remember people moving the goalposts like that since, let me check... when did AI Dungeon go viral? Wow, 2019.
What's moving the goalposts? I am very much amazed at what Fable can do. I push its code straight to prod.
But I am just as amazed with how little real life consequence it seems to have! Even software houses were hit more by interest rates than by this magical revolution.
If I couldn't directly observe Fable in action, I wouldn't believe in AI.
> What's moving the goalposts? I am very much amazed at what Fable can do. I push its code straight to prod.
>> What's moving the goalposts? I am very much amazed at what Opus 4.8 can do. I push its code straight to prod.
>>> What's moving the goalposts? I am very much amazed at what Opus 4.6 can do. I push its code straight to prod.
>>>> What's moving the goalposts? I am very much amazed at what GPT5 can do. I push its code straight to prod.
>>>>>> What's moving the goalposts? I am very much amazed at what Opus 3.5 can do. I push its code straight to prod.
I am definitely not impressed by Opus. Quite the opposite usually.
I think it's more that it shifts the thought process from "If this is going to take 30 years then we won't bother, because the investment can be spent on things that pay off sooner" to "If we can do this in 3 years then we'll make that investment because that's a good bet."
The way it changes the game is by lowering the cost of making radical bets so we end up trying more moonshots.
Do you have examples of such moonshots? Ideally not themselves AI related
Right now we are in the golden age where we do the same and take time off. Employers have not yet fully caught up with the workforce. I can't think of people that are not putting less hours this year for the same salaries.
The questions is what happens when they catch up. They'll cut like 50%+ of the workforce? What happens then to the demand that makes their companies work?
Or an example of MS - their main cost like most software companies are people, especially software devs, which are to be replaced by AI so on the surface they would greatly benefit from it. But their products are centered around helping out people do stuff on the computer. Why would you need that when the AI will do it better and faster directly operating on the data or using e.g. Python?
Who knows. What I tell is that the AI productivity boom is here. Just for a change right now it is not shown on businesses balance sheets, because it is captured by the workforce in non monetary ways.
> If AI makes us 10x more productive, the 3 years since its popularization were enough for a product that would have taken 30 years to build without AI, for instance.
A product still requires a lot of handholding and human thinking, at least if one does not want everyone even throwing a glance at it to immediately be repulsed by the usual AI slop tells.
> I'm an AI advocate but that question makes me feel that AI is simply "very useful" rather than being a historical game changer for humanity.
It absolutely already is a historical game changer on par with the Industrial Revolution when it comes to the amount of jobs destroyed and economies screwed up - and the impact will be even worse in 10+ years as existing seniors retire but no new seniors rise as AI has destroyed entry level career paths.
> It absolutely already is a historical game changer on par with the Industrial Revolution when it comes to the amount of jobs destroyed and economies screwed up
Which economies are already screwed up?
IMO it can’t ever be on par with the Industrial Revolution because AI can only really affect the information economy. Things people do with their hands/bodies have either already been automated or can’t be with current tech. If you’d asked people decades ago they might say no one will ever work in factories by 2026 because they’ll all be automated. It didn’t work out that way. I think AI will go the same way: absolutely game changing to some industries (of which software engineering will be one) but a great many will still survive with less dramatic changes.
If anything it might result in more focus on the human aspects. How many people out there earn their stripes putting together slide decks? In a world where an AI can put together the snazziest presentation you’ve ever seen in a heartbeat it’s going to matter more how you stand at the front of the room and present those slides than it does today.
Not disagreeing with you, but adding to my argument. Despite the human handholding, I feel that 3 years would have been enough for 18 months of thinking about the product plus 18 months where a team could get 3-5 years worth of coding/development done. Yet, we haven't seen anything big yet, like a new Youtube/Instragram, an amazing videogame, a major cure etc. Maybe they are coming, but every day that passes is an indication that the net productivity positive of the technology isn't as big as advertised. In my own personal use of AI, I experienced a 30-50% increase in productivity, but no more than that.
It's only been at most the past year where AI has been unambiguously helpful and not a hindrance. 3 years ago it gave the appearance of being helpful but it tended to be more of a hindrance.
perhaps being 30-50% ‘busier’ is something quite different from being 30-50% more ‘productive’ and ppl are just bad at detecting the difference
Maybe so, but least for me in my personal coding projects, I'm going through my own task list noticeably faster. I had created this list before getting a ChatGPT Pro subscription and the number of bugs found in AI reviews and rate of closing tasks has certainly increased. I'm not saying my anecdote scales to teams or even other people, but I have no doubt about my personal productivity change. I wouldn't be paying for it otherwise.
IMO the constraint is that AI is still rather anaemic at sustainable green-field projects: It's good at one-off oneshots, and also at refactoring or fixing bugs or adding features to existing projects, where test suites and significant architectural scaffolding already exists, but the more you move away from that, the more wobbly the results get, and the more the human once again becomes the bottleneck, for all the hard work of coming up with all the conceptual scaffolding in the first place. Typing speed rarely was the bottleneck there anyway.
Any argument about regulation in the US which doesn't mention that China and other countries won't follow that regulation should be heavily questioned. Exactly who are you stopping from doing what you don't like?!?
Law abiding US citizens won't be able to run them, but you didn't solve anything other than making sure US citizens pay Sam or Dario.
Any arguments that US laws or regulations should apply to people in other countries should be heavily questioned.
That's definitely not an argument I was making.
Why the comment submitted by jacquesm, who posted the link and it's not exactly an anonymous poster, that said "Apologies for linking to X but this is worth reading." has been flagged to death??
Maybe because "apologies for linking to X" is needless gaudy virtue signalling.
Make it "Apologies for linking to Amodei" and we're talking.
Just seems like virtue to me.
Vouched. (A few more might resurrect it)
You can 'vouch' for a comment, if you think it shouldn't be flagged.
I'm more interested in why somebody would suddenly flag that post. It's not spam, bot, offensive or anything else. Maybe you could downvote it if you don't like the take on X - it still would be childish but at least would be proportionate.
Why? I'm sure there are plenty of Anthropic employees on here...
i was about to post the same thing. i vouched for the comment, but it's too dead to recover from my single vouching action.
This site's audience has been less of the technical curious kind and more of the pathological misalignment one
It will be interesting to see what kind of progress in biology and medicine they will preview during the fall. He does make a good point that most of the progress so far have not been material in the sense of providing real positive outcomes for ordinary people.
Or even for the HN crowd, when will we see e.g. "Mythos aided research discovers 10 new viable battery technologies"
None?
Fair enough, help navigating complex and bureaucratic processes, personal validation and a more tailored search engine do provide some value
So the solution to convince the public that these companies aren’t “looking for new ways to screw them over” is to try to go into biomedical research.
I guess it’ll be great for Anthropic to have the cure for cancer, but what’s that gonna mean for people with cancer? Funny how he doesn’t talk about that part.
Called it! He was bound to squeal after the release of QWEN 3.8 in one way or another and here it is. I wasn't even a teenager but I'm getting so much Jobs/Ballmer flashbacks with internet explorer and microsoft office.
Edit: Does anyone else notice the switching between dashes and em-dashes between paragraphs? Tells you a lot about the man, doesn't it.
> Edit: Does anyone else notice the switching between dashes and em-dashes between paragraphs? Tells you a lot about the man, doesn't it.
That he doesn’t know how to use basic word processing, or even agent SKILLS.md or whatever it’s called now. Or maybe it’s just the next generation of vagueposting.
Tells you a lot about the man, doesn't it.
The CEO of Anthropic not using AI extensively would be news.
I don’t understand your edit; can you quote an example switch?
The tech industry cooking up some new way to screw people over for 25 years
The tech execs waking up on a random monday:
> I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.
No shit Dario, no shit...
Yeah and we’re not even done with dealing with the societal consequences of social media and dopamine algorithm driven everything
I can’t help but wonder if this essay was timed in an effort to compete with that bombshell story by WSJ a few days ago. It reports that Dario’s wife, Cami Clark, wields secret influence at Anthropic, had her existence scrubbed from the internet - and also tried to get Jeffrey Epstein to invest in her porn startup. Quite a read!
https://www.wsj.com/tech/ai/claude-dario-amodei-wife-anthrop...
> Overall my view is that AI is structurally a technology that tends to concentrate power
> Open-weights do help some with this but are nowhere near a sufficient solution
Which is exactly why Anthropic contributes nothing to, and actively pushes for roadblocks and regulations for open weight models. Can't allow any hope to the masses.
Only people with $$$$$ are allowed to touch Fable. Which of course we have have aggressive guardrails for in case you even try to use it for something dangerous like AI model developement. Plus we are going to retain all data submitted to it just in case someone is trying to be sneaky.
> I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world.
To paraphrase Sinclair "It's difficult to make a man realize the issues with regulatory capture, if they're to be the benefactor of regulatory capture".
> A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice.
Not so sure. After all lack of the latter did help establish the current tech overlord rule.
> This is why Anthropic has always made its policy proposals very carefully. We try very hard to make proposals that disadvantage (slow down) frontier AI companies while advantaging smaller competitors.
> ...
> completely exempt any company below a certain amount of revenue or model training costs from being covered at all
One could argue that "Frontier AI" company know they have nothing to fear from company with less than XM$ revenue, and so their support for this type of regulation is still a way to force regulatory capture. In any case, whatever regulation you support, it's a regulation that you didn't have to handle when you were growing, but that incumbent will have to deal with.
Whatever it is Dario's intention or not does not matter. Capitalism push to consolidation and the eventual regulation that will need to be applied to mitigate the externality from a new industry, will mean that their will only be a handful of "very big" winner. Same story since the beginning of the industrial era. Even if that's not what Dario personally want, it's in the best interest of the shareholders, which will force Anthropic to do everything it can to be one of the big one.
> Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).
The same can be said about a lot of other industry (aviation, oil, chip manufacturing, ...). Every country / union big enough will finance their own champion to try to keep a foot in the industry even if they are not the best.
Is this cure of cancer in the room with us?
When people talk about Qwen 3.8 being on a par with Fable, they're really talking about Qwen 3.8 Max aka Qwen3.8-2.4T-A95B. That's a 2.4 trillion parameter Mixture of Experts model with 95B active parameters. You need about 400GB of RAM to run it. No one is running that locally.
The distillations of Qwen 3.8 down to a 27B model are good, but they're not on a par with frontier models.
When Dario talks about open weights not being a solution this is what he means - if you don't have 400GB of VRAM lying around the fact that there's an open model like Qwen3.8-2.4T-A95B doesn't really help much. If we're not regulating how models are available, or making sure access is open, then RAM prices will mean everything concentrates on a few very rich companies.
I don't really understand the argument you're making, but just to add a data point:
DeepSeek V4 Flash 0731 is 167 gigabytes from the developer and as a GGUF with no additional quantization. It limps along on my 192GB M2 Mac from several years ago [0]. This model tests better[1] than Claude Opus 4.6 released in February. That's six months ago - what will be available 6 months from now?
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/tr...
https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF
So yeah, enthusiasts aren't going to run frontier models on their gaming machines, but a small office could easily justify the $30k - $100k cost to run something like this at high speed. The small company I worked for routinely spent that kind of money on Dec Alphas twenty five years ago, and that's not accounting for inflation adjustment.
And this is completely discounting the advances smaller models are making. You're right that Qwen 3.8 comes in different sizes. However, Qwen 3.8 27B and Qwen 3.6 27B do run on gaming cards, and they're better than the frontier models from twelve months ago.
I have no idea what will happen in the future, but I wouldn't base my guesses solely on the largest open weight models.
[0] Yes, it's unpleasantly slow (5-8 tok/sec)
[1] Yes, benchmarks should be taken with a lot of salt.
> The distillations of Qwen 3.8 down to a 27B model are good, but they're not on a par with frontier models.
Does it have to be? There are plenty of coding tasks, where it's good enough.
Practically, no, the distill is great. It's fine to use it.
However, if you're having a discussion about access to frontier models, and using Qwen 3.8 as an example of how open weights is a solution, then you should be honest and accurate about what you're talking about. Making an argument like "People can run Qwen 3.8 at home. That shows open weights are great." is a bit disingenuous if you're not also making it clear that you're not talking about Qwen 3.8 Max (or that you have a beast of a PC at home :D ).
I agree 100%. IMO, it all started with ollama misrepresenting the Deepseek R1 distills as Deepseek R1, all for hype and marketing. I've had so many ostensibly technical people telling me: "I tried DeepSeek R1 and it was terrible", and every time when I probe further they'd tried the tiny 1.5B Qwen2.5 distill model that was further brain damaged by ollama's naive RTN quantization[1]. DeepSeek themselves were very forthright about it by naming it DeepSeek-R1-Distill-Qwen-1.5B[2].
I suppose the road to technical hell is paved with marketers and grifters. :)
[1]: https://ollama.com/library/deepseek-r1:1.5b [2]: https://huggingface.co/deepseek-ai/DeepSeek-R1-Distill-Qwen-...
> coding tasks
Exactly. There are common coding tasks that these models can adequately do. They are absolute trash for anything that isn't coding. And even with coding, they are so, so far behind frontier models.
No one is running that locally because of the AI bubble consuming all the hardware in the industry. That won’t be the case long term though
- frontier llm access means you are at an economic advantage
- a big risk of this is ongoing wealth concentration
- the "open weights" approach can't solve the problem of wealth and llm access being linked; you need compute too, and compute is expensive, thus "open weights" still favors the wealthy
- instead we need "objective and fair institutional processes", as this will allow small labs cook up their stuff while frontier labs get regulated
why would i care about what these smaller players do, if economic advantage = frontier model access? also, the reasoning around "why bother with open weights because compute is expensive too" seems like the kind of logic a motivated 12 year old could work their way around in 30 seconds. things are not either/or dario, you said so yourself.
seems like a 400 word corpo misdirection essay. par for the course.
He seems to be grasping for a point but not providing argument for what it is exactly. And certainly not support that is logically coherent
So, what does a for-profit CEO, beholden to profit-seeking investors, have to say?
> Overall my view is that AI is structurally a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws).
If only money thirsty capitals didn’t drain the HBM/DRAM/NAND capacity to rush to build out all the data centers so that they can make money off of inference, driving up memory prices like mad man leaving consumers with not only no chance of local inference setup, but also higher price of consumer electronics. Of course it has nothing to do with regulation, more to do with the greed.
For $20 a month you can experience what it’s like to have a 130+ iq for a few minutes a day.
Why does this need a master?
Do we regulate bio geniuses for the danger they pose? No, we do not.
Do we regulate other tools like hammers and drills? No. Books? No.
From anthropic we now have Science as a PR strategy. God I can't wait until this bubble pops.
> I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over.
It is so tragicomic to see people whose lives are built around companies coming so close to realizing that everything they do is bad, and then at the last minute swerving aside to convince themselves that no, if they just do more of it and somehow do it "better", then it will all be okay. The reason people don't trust companies and believe they are cooking up some new way to screw them over is because that is what they are doing. If Dario or anyone else really wanted to dispel that perception there's an easy way: do a total 180 and start fighting against everything you've been pushing. But none of them will do that because they still fundamentally believe that what they are doing is good, and are unable to see that fundamentally it's bad.
I think that Dario is intelligent, incisive, has good judgement, and is well-intentioned. I hope he continues to wield influence. I think Sam is also most of those things, but I think power has a One Ring-like effect on him. My most contrarian view is that I think Elon’s problems are primarily appalling judgement when it comes to areas outside of tech, and has the emotional regulation of a child, although somewhat incredibly I believe he is fundamentally well intentioned. I don’t think any of those three really want tech feudalism, although I think the current US administration would press a “turn us into Russia” button the second they caught sight of it, which is to say, I think they have the worst of intentions.
> I think Sam is also most of those things, but I think power has a One Ring-like effect on him.
That's putting it mildly, very mildly. The guy all but wrecked the entire world's DRAM supply chain by using negotiation tactics that, if there were any semblance of regulatory authority left in the US, would lead to criminal charges for market manipulation.
Arrest Sam Altman, he is one of the most dangerous people on the planet. Not a single shred of respect for the 99% and the consequences his actions have on them. (And similar things apply to the other controversial figure you mentioned, Elon Musk actively enjoys hurting people by the looks of it.)
Dario's messaging has been extremely confusing and HE'S RESPONSIBLE for further increasing the negative sentiment of the public towards AI.
He's clearly focusing on being on the news to provoke emotions on people, preparing for the Anthropic big blockbuster IPO.
Dario, Sam Altman and others should instead be praying every night that this will take us to the Singularity VERY SOON, because if it doesn't, the entire country will blame them for absolutely destroying their retirement and the US' economy once this AI Datacenter bubble pops.
There are other countries.
* * *