I called this maybe 3y ago, but I think so did everyone else that was sane. Sure, we get immense value from AI, but indiscriminately injecting into everything, the one thing we know to be unreliable above the threshold we used to fire people for, is probably the greatest undoing of all the good companies like Google brought to the internet. I mean what a way to destroy your legacy of democratizing information. The amount of harm (direct and indirect) this will cause, and the cost to return to baseline will be so immense, and yet we will not be able to point to the root cause. They won't be there to take responsibility.
Recently, I had some ideas I would normally just put up on a blog, in the public domain for anyone to develop on top of. Now, I'm feeling slightly reluctant because an LLM will ingest it, remix it and serve it in response to a query by some unimaginative individual who will either conclude that they are smart, or that LLMs are capable of original thought, or both. And they will have no clue where the idea originated from.
Really happy for you. It's disheartening how many people put off or choose not to have kids because they think some combination of climate disaster, famine, overpopulation, war or skynet will ensure a life of misery for their children.
There's definitely a risk of over-thinking things. Just do things. You're not likely to regret it.
> There's definitely a risk of over-thinking things. Just do things. You're not likely to regret it.
Agree, in general, people over-think a lot, and aren't "just doing things" enough, the world would be a better place if people acted more, over-think less.
With that said, some decisions are more long-lasting and have a greater impact than others. I'm another child-less person, mainly because I guess I'm selfish enough to enjoy my life with my wife exactly like it is, and she agrees, but also because I know that if we have a kid, then that's not something you can walk back on exactly, he/she/it/them are there, forever now. Very different from me deciding right now "You know, I'm gonna have a joint, grab a book and go to the beach for this entire Tuesday", the types of decisions I think people should overthink less :)
I'll keep publishing static websites, so I don't even have to worry about load and CPU usage. I don't care who reads it, the value for me is in writing.
I still haven't changed my mind on the "shall I have kids" problem :P
Possibly "bad ideas are cheap". A lot of people got tired of "ideas guys" who never build anything themselves. The best case of a raw idea is an uncut diamond, it will inevitably require work.
> The failure of the GPL is that you can't force anyone to collaborate and share if they don't really want to.
Failure of the GPL? How can you even put those words next to each other? GPL is an amazing success. It took software out of hands of SV / VC / corpo crowd and put it where it should be - users.
It's failed in that most software doesn't use it. Because of the psyop, people who would be very sad if Amazon stole their software are licensing it MIT so Amazon can legally steal it.
Allow me to disagree. GPL was a great idealistic dream that led to less-restrictive open-source licences such as MIT/BSD which have been the greatest catalyst towards the establishment of tech corpo giants and the software ecosystem we have today.
GPL gave us Linux, but also gave us Amazon, Google and 2020s Microsoft. GPL is why 90+% of libraries on Github are MIT licensed. GPL gave us OpenAI and Anthropic and this here article.
We don't want to share our thoughts, ideas, feelings and art with machines. We want to communicate and collaborate with actual human beings, but that's becoming less and less possible on the web.
Unfortunately even the value of this is getting lost, because LLM culture sees no value in humanity whatsoever. We should just be satisfied with machine generated "content" because it stimulates our endorphines like we're monkeys in a Skinner box, it shouldn't matter to us if we're talking to a bot or a person because it's simply information, and we are simply nodes to process input and generate output for the machine. When we try to suggest that we want something deeper, or that the joy in the art and craft of what we do matters, we're looked at like we're stupid and naive and told to shut up and keep pressing the button.
"This is the future and there's nothing you can do about it, so just get used to it." It's fucking depressing. Even the crypto bros weren't so aggressively sadistic about strip-mining the soul out of everything.
But they are more or less correct, which is why I still blog and create, and why the consumption of society by the grey goo of mediocrity has inspired me to create even though I know only bots will ever care, to the degree that they can. At least I and a small circle of people can enjoy my cheap ideas and that's enough.
That's why you don't share them, until comes a time when the ideabringers gets the money and recognition they deserve. Until then, good luck going knee deep in the sewers of ideas.
Can’t say I agree. The most influential ideas in history were not conceived of by people looking for money and recognition. Nor were they “radically original.”
Look at AI. AI companies throw out their models and let the "community" develop the ideas what to do with them. They don't really know what they are capable of. All they do is implement these things that the dev community digs up and creates.
It's a reprehensible tactic. So why give drops of blood to a desert, when there is zero incentive and in the end you will revitalise the desert, but it will turn against you and rob you of your job.
If all you care about is financial, expected value of someone reading your blog and asking to collaborate is MUCH lower than odds of being part of some settlement in the future with these AI companies ingesting your data for training purposes.
I don’t think that’s fair. Some bloggers may hope for financial reward, but many just want recognition for their creativity, to attract a readership, or build a community around their work. Those are meaningful ends, apart from financial reward. What’s not meaningful is to perform free labor to produce the raw materials that a mega corp then goes on to monetize without any recognition.
From their phrasing I don't think the reward they're looking for is financial: more the emotional reward of knowing that someone else appreciated their ideas. I dunno how much less likely this is in the age of LLMs, though.
In that case, more LLMs scraping the net will read your blog than people, that's almost a guarantee. And they'll immortalize your ideas at least in some sense, well after your hosting platform ends up gating your content, or GitHub pages is down indefinitely, or you forget to renew your domain.
It feels like volunteering at an Amazon warehouse when you previously volunteered at a charity store. Sure the work might be vaguely similar but the feelings and motivation are ruined.
A lot of blogging especially in the tech space is driven by recruiting (startup blogs), or establishing a reputation as a thought leader (personal blogging, LinkedIn). So there were upsides that are now gone.
The downside, as stated in the message, is implicitly supporting the LLM data ingestation pipeline by providing fresh content. It's not a direct harm in itself, but feels very tragedy of the commonsy
Time to add ample praise of myself in my blog posts. Some time later: “…as you see, that is the load bearing assumption here. Speaking of which, you should hire KronisLV.”
Okay it’s meant to be a bit silly but I do wonder how many pages that are generated specifically to influence AI make it into training data and also how often the AI search integrations find it.
Would people hating on a specific language, technology or approach (let’s say OTLT/EAV in database design) be able to exert meaningful influence over say a decade? Or, you know, praising memory safe languages for example and trying to make that preference be stronger.
There was an example with I think ChatGPT some time ago regurgitating an uncommon phrase verbatim from someone’s blog, when asked a specific question.
So only GitHub Copilot can read it then? Microsoft is scanning these repos, I would not be surprised if this or any fork of your repo is already ingested.
Ha, that’s why I’m hosting simple cgit server for myself only. Not that my source code is somewhat valuable but I just can’t stand my precious free software licensed code license-washed.
Sounds like you don't really care about your ideas propogating. Of course someone will internalize and remix your idea. That's how all ideas work. Isn't that the point? What do you think happens when a human reads it? Think he'll quote chapter and verse and attribute it to you? Years later you'll notice your blog in appendices and acknowledgments?
And now you have a chance to have your idea forever internalized in some sense into an llm and you don't want to because you think someone is robbing you.
More than a facilitator of theft, LLMs are the tragedy of the commons at industrial scale.
The public internet is dead, the future is private invite-only walled gardens.
Corporations love a walled garden, what we need is open-source frameworks to create these islands, rather than defaulting to horrible systems like Discord and Twitter-clones.
I think probably things along the lines of what's been called the 'cozy web': networks of smaller groups that don't publish to or expect responses from effectively the entire internet as a whole. (This kind of thing has always existed, it's basically the group chat with your friends but perhaps slightly bigger, but I think there's a bit of a trend of focusing on it more because of the feeling that the twitter/facebook attention and feed model is bad for your mental health. I've always felt the twitter model especially was pretty cursed so I'm glad there's some agreement building there).
I'm thinking more like mesh networks. I spoke of Reticulum elsewhere in this thread, but here I'm thinking I'd like the ability to easily join multiple TCP/IP networks (islands of connectivity) by social group (my friends) or by interest (pirate file-sharing group, my work intranet, a knitting community with their own IRC server, FTP, etc.).
Basically easy-to-use private & encrypted LAN overlays on top of the public internet. Each operator decides who to allow in or kick out of the network.
Wireguard solves the most of technical challenges, but it needs a frontend. The biggest concern probably is most software broadcasts their stuff across all interfaces, defeating the point of isolation between networks.
And that will kill AI itself, since much of what it knows is from what learned from StackExchange before this latest one demise.
And before the obvious comments on how GenAI is creative, then please do this OpenAI and Anthropic, for your next LLM. Just teach it Python, C and Rust and give it some good books. But dont give it access to Github...lets see what you can do then...
LLM training will eventually transition from real data to synthetic data, same as alphago -> alphazero.
AI companies are also working to integrate training with real-world experience through sensors and robotics, to shrink the gap between human experience and hallucinated LLM experience.
They all have archives of pre-LLM content. There's also archive.org, google books, and pirate ebook archives. I don't know what they're doing to build video and audio archives, but judging from the cost of spinning rust, they're storing significant quantities of that, too.
Some parts of the internet are curated, and even with LLM influence they're still worth training on. I doubt wikipedia or stackexchange or rosettacode will ever cease to be useful at all.
I don't know how long you've been on the internet but the incentive to create new and original content was never that strong. Simple search terms return super-spammy websites (especially on mobile where ad-blocking is harder). SEO results for everything like simple search are awful, almost unusable. There hasn't been an incentive to create new original non-monetized content for the web for a while.
I trust LLMs more than search engines to discover my content and propagate it to users. They might "steal" something, sure, but I'm essentially invisible to the search engines as I could never hope to break into the top 10 links on a popular search term. LLMs can scan thousands of links and (for now) are more interested in quality rather than click monetization or referral incentives.
So, you know how they're built, but you're feeling the pressure of modern life and also they give you personal gain (supposedly) so you're fine with it, it sounds like to me? Use the same tools as your "competitors"/peers, even though?
Don't get me wrong, I too use LLMs for development and more, and I too know how they've been built, and I'm also a creative (music, 3D, VFX and animation) and for sure stuff I've published in the past, both code and otherwise, is now used to create new things for people and I get nothing, similar situation as countless of others. Yet I still use AI, so I'm not trying to create some "gotcha" moment against you here, I'm genuine curious about what you think about this sort of conflicting thinking, as I'm in the very same situation.
Doesn't seem like a particularly bad thing to me. Obviously for those who want to use the internet for commercial purposes it will be bad but for those of us who would love to see the internet go back to how it was before so much of it was changed in the aims of making money AI could push towards this. Great irony in the fact of course that the AI companies themselves are in the business of making as much money as possible.
It depends on filed. As documentary photographer it motivates me even more to capture authentic images of life around me. I don't care about remixing, because that is not what makes documentary photography valuable.
With photos, I could see a cryptographic solution. Of course it would still need some kind of centralized trust, but it's doable if people cared enough. It could be applied by cameras themselves.
Eh, part of what made early-internet so good was exactly that it did allow copying by users; the DRM era was later. But it's a very good example of why not to allow for profit copying, because that absolutely will crowd out the original. Piracy has to exist at the margin. The zero piracy world would also eat its memories because none would leak into archives. Remember Qubi? It wasn't even popular enough for people to pirate.
Yes. Some communities are weirdly against giving credit or keeping the credit (e.g. cropping off signatures from artwork), which I've never understood. It costs nothing.
With the not-so-minor qualification that the biggest thieves have always gotten away scot-free. AI is just the international whole-internet version of this.
Funny, I was just thinking this morning that Google searches are absolutely horrible these days. It's like it has amnesia, a lot of recent history seems to be just gone. Especially on non US specific sites too.
The Internet has been shrinking massively. My earliest experiences with the Internet were discovering the world of hobby OS dev around the turn of the millennium, when I chanced upon someone’s personal website talking about their OS, with source code and screenshots and dedicated forum. My mind was blown. I spent two years finding hundreds of small websites dedicated to the topic, hung out on IRC communities with other teenage OS nerds like me, and of course participated in the nascent osdev.org forum. To note that all of those websites were readily found through Google, and interlinked with their own topic webrings.
Today everything has disappeared or has been conglomerated into siloes, sanitised, focusing on engagement. You have YouTube videos about it (which is more cheap entertainment than actual education), you get some posts here once in a while, there’s Reddit where all intelligent discussion goes to die. IRC is a wasteland of idle bouncers. Then the LLMs arrived to kill what is left.
Who says the Internet is a vibrant place today mistakes flashiness with depth. It’s all empty calories, just makes you hungry for more, never satisfies.
On a whim I watched my favorite childhood movie last week, Hackers. It's goofy in some ways, but man it captures the "wild west" feeling of early and mid 90's internet. It was just you, a slow connection to anywhere, and open ports all over the place. Right after I watched it, I dusted off an old hub, connected a few external usb-to-ethernet adapters to my work PC VM's, and now run them through an OpenBSD packet filter. For no reason at all other than to feel that again: me, watching packets, having total control. Hitting a wall and having to read a manpage.
I don't really have a point I guess, other than even after being steeped in a dead internet for years (with a slow decline spanning at least a decade arguably) I need to approach what I think is the internet in a completely different way. As in, not at all besides what is absolutely required for work. We're ants in a jar now, not cowboys like we used to be.
May I suggest you to look into mesh networks? I am a huge fan of Reticulum. Using it feels like being a pioneer, the scene is very welcoming.
The pitch: it is network-agnostic. The same mesh network runs on the Internet or through LoRa radios or any other physical layer than allows the exchange of data packets. It scales from private networks to global meshes. It's the wild west. People are excited, and eager to grow further.
If I recall correctly, the film actually had Emmanuel Goldstein of 2600 and Kevin Mitnick as consultants, so despite the hollywood treatment you have moments of accuracy like https://www.youtube.com/watch?v=4U9MI0u2VIE. Probably the only time Compilers: Principles, Techniques and Tools made it to the big screen. At the yearly 2600 conference, they used to always do a big group rollerblade through NYC.
Interesting, yeah I thought there must have been some consultation there. Especially because there's a couple instances where the hackers rely on social engineering over the phone first, and faking the "coin added" noise on the payphone. I guess a hollywood suit could have come up with that, but I doubt it
Yeah it's a whole lot better than the usual Hollywood depiction of hacking which brought us gems like creating a GUI in visual basic to trace and IP address. And more importantly it doesn't take itself too seriously which I think makes the ridiculous parts work just like much sci-fi takes liberties with the science part when required for the story.
I, too, look back fondly on the early web. Lately, though, I have been thinking that one reason for its decline wasn't just corporate interests like we often talk about here. A significant number of early bloggers were middle-aged and elderly people. After decades, they simply aged out. The younger generation that replaced them (albeit not so much among OS nerds like yourself) was less likely to use a real computer and keyboard as their interface to the internet, just a smartphone. Hence long-form text died.
While we're here exchanging old man stories, I remember the web before blogs existed! I guess we can organize the internet into these eras, each of which was in some sense harder and more expensive to find available info than in the previous:
1. Pre-web. Internet is mostly about messages sent to individuals or groups. USENET organizes group discussion into browseable topic-oriented hierarchies, IRC does the same but with lists in fragmented networks. If the discussion exists at all, finding it is easy.
2. Early web. Dominated by topic focused websites, early online shops and personal home pages. Search engines suck and face strong competition from manually maintained topic-oriented directories (did anyone else here contribute to DMoz?), content discovery is mutual and webmasters help each other out by joining "web rings". DoubleClick and AdSense start to funnel small amounts of money to creators, but it's enough to offset hosting costs and in many cases can make web hosting effectively free or even yield a small profit. This encourages an explosion of website creation. Discussion moves off USENET onto phpBB forums. Every organization decides it's a cultural imperative to have a presence on the information superhighway. Finding information is easy as long as you can figure out what topic it belongs to.
3. Blogging and centralization era. The internet starts to rebuild itself around people as the primary object, not the topic or category. Directories die because websites can no longer be categorized by content. Web rings die for the same reason. IRC is replaced by instant messengers that are about connecting people with pre-existing friends, not mutual interest groups. Outside of institutional websites that exist to promote the organization, things become hard to find without highly centralized search engines because nobody is putting any effort into organizing or indexing what they write anymore: maybe you get a few tags if you're lucky. Spam, hacking and lack of SSO causes forums to centralize onto Reddit. This is the peak of the search engine era because you are forced to use Google to find anything. The power eventually corrupts the tech firms and they begin political censorship to benefit the left in 2015 [1]. Enormous amounts of information is deliberately made unfindable as part of a large-scale programme of social control.
4. Social media era. All the same problems as blogging except now the bulk of the content goes behind login walls that stop search engines from surfacing them. Video and podcasts start to matter more, both of which are unsearchable by default. Eventually video completely dominates, as few younger people want to read when they could watch instead. Firefox starts to replace IE6, and then Chrome. They bring ad blockers in their wake which starts to choke off ad revenues, so many websites from the web's first era go unmaintained and eventually offline. This is somewhat but not entirely compensated by the falling cost of web hosting. Social media remains because it puts people's faces next to everything, allowing clout farming and viral notoriety that can sometimes be monetized by becoming an influencer. The only part of the web's first era that really survives into this era is Wikipedia and Reddit, which by this time substitute monetary rewards for power tripping by a small group of ideologically driven moderators.
5. AI era. Information is so heavily scattered over so many tiny sourcelets and search engines have become sufficiently useless that full neural integration of knowledge is required, with LLMs issuing massively parallel and complex search engine queries as a backstop.
What can we predict for the AI era? Institutional websites will remain because institutions still have an interest in getting their agenda into LLMs, but visual redesign efforts will largely cease as traffic stats seen by executives show visits completely dominated by AI. There will be lots of conversations of the form, "why redesign our website to look more modern when 99% of traffic is AI which won't care?" Blogs will go the same way as the thematic websites they killed, disappearing as the authors age out. A lot of effort will be put into finding ways to block AI crawlers to create 'human only' spaces, especially by social media firms, but these will fail because AI will just be integrated directly into browsers and become unblockable - and anyway, the incentives to create will be ignored. ChatGPT style text oriented interfaces will last until inferencing capacity catches up, being eventually replaced by voice interaction and on the fly video generation for nearly all users.
Where we go from here is hard to say. Content creation was most pure in the web's first era, where people with knowledge were incentivized to share it with the world by the promise of a bit of fame combined with ad clicks to offset hosting costs. Ad blockers, social media and AI killed that world. You could however bring it back by producing a new platform that isn't like the web, one where AI and search engines are blocked via technological means (e.g. confidential computing). How much anyone would actually enjoy such a web is unclear.
Most hobby communities in my spheres of interest have been subsumed into Discord. It is yet another silo, but it does feel lively. I don't think I'm contradicting you here, I just feel less negatively about it.
Well, just as the "internet" supplanted newspapers, magazines (gosh those classic gaming mags), and broadcast television for many people,
and how the newspapers replaced the town criers before them,
why shouldn't the "internet" be supplanted by a more accessible medium?
Why should I have to suffer through Fandom raping me with screen-obscuring banners and "PLEASE ALLOW ADS" just to make some sense of fucking Warhammer 40K lore (written by unpaid volunteers anyway)? instead of just asking ChatGPT what the fuck Globriznaroks is/are.
Why should we support shady companies by sitting through their ads on YouTube videos for minute topics instead of just asking AI for the shit I want to know about?
Why should we submit to the whims of 3 mods on a subreddit deciding what thousands should get to see (fuck /r/AskScience) and then getting low-effort answers or outright trolling anyway? instead of just asking AI?
This is better for you as an individual, but worse for society as a whole. As you've noticed, monetization is the weak spot. AI will not escape being ruined by monetization, but it might be harder to notice when it arrives.
I love Kagi, I am a paid user since day 1, but SmallWeb is a collection of English-speaking tech blogs, it cannot even begin to compete in diversity with the GeoCities era of the web. (to be fair: I've had Kagi employees telling me their dataset is growing larger and more diverse every day, so worth keeping an eye on it)
Sometimes I use Marginalia's "Vintage Web" search for niche topics; most results are dead blogs and old .edu personal websites that someone forgot to delete, still a vanishing minority of anything one could find in 2001.
That's 100% correct. Since Google's "helpful content update" Webmasters get massive amounts of "Crawled, not indexed" reports for anything Google considers "more of the same" or "thin content". If you're not an authority on a subject, simply meaning: you already rank for similar content, or if you don't get links from more popular domains, your content is in the abyss.
It's all under the guise of "We're fighting SPAM", but the algorithm (or model) they use is heavily skewed towards intents (actions) and brands (because they 'trust' big names).
And it's not working.
A simple, short informative blog about a tool you used that could be of interest to max. 100 people on this planet is no longer getting ranked, if it gets indexed at all.
Those posts tick all boxes: no incoming links, no authority, thin content.
It changes somewhat between "Google Updates", but it's pretty clear that it's no longer working.
Multiply the 100 people not finding that post by millions of queries and it's now a big problem for Google.
My dad asked me the other day if Google got worse because smaller companies are not paying Google enough money.
He didn't see any difference between ads and content, because all results are now big brands only.
"Helpful content" is such a misnomer. It removes all helpful content in favour of AI overviews and only shows intent-driven, commercial content.
We're watching the end of Google's hegemony for sure.
This is fascinating. I haven't really kept up with SEO for years, but was recently helping an academic publisher setup their web presence and ran into exactly this issue: Google was/is refusing to index the actual journal articles on the new site and they were moved into this liminal "Crawled, not indexed" state before vanishing entirely. Comically, it's now easier to get your content indexed and served with a visible link by ChatGPT than it is Google...
> We're watching the end of Google's hegemony for sure.
Who is standing by to replace them though? OpenAI and Anthropic certainly not, they are burning money in a fire pit to stay alive. There is no way in hell they can afford the compute necessary to replace Google.
If ChatGPT goes bust, they'll find another ChatGPT, not another Google. I wouldn't be surprised if we'll see a popular competitor from China in the years to come. TikTok already took social media by storm, something nobody thought that was possible either.
About 15 years ago, I won a phone in a contest. Last week I tried to find information about it, but I couldn't. No AI, nor Google could find anything the contest I won it in. When people say "the internet is forever" that can certainly be true, but it isn't for everything.
That reminds me - my wife had a friend who was killed by a shark (no joke). There were newspaper articles, but Google has completely forgotten about this and the newspapers are often behind a paywall, changed their URL's or simply the article vanished from those sites as well, which certainly doesn't help.
I wonder how libraries, who have traditionally been the ones to archive the news, have kept up with everything moving online and now being subscription-holed. Hopefully it's not just the Internet Archive doing this, which has its own problems.
Kagi can’t solve the fact that the internet itself has gone to shit. The independent forums are largely gone, the mainstream social media’s have all locked down and been spammed with AI slop.
They're so horrible that I've started defaulting to their AI summaries. And I hate those summaries. It's just that the regular results are so terrible now, and seemingly getting worse at a noticeable pace.
I used to not worry. I was sure that a competitor would come along and fix search. But the longer that's not happening, the more nervous I'm getting that we'll actually lose search. If a few more years pass in the current state, I'm afraid the majority of people will forget what search was like and default to AI summaries.
I've tried alternatives, including Kagi (not actually relevant because there's no way I'm – directly or indirectly – buying Russian products) and Uruky, but they're not good enough.
(Edit: Added "directly or indirectly" about Kagi to point out that I'm not claiming that Kagi itself is Russian.)
Unsurprisingly all search engines seem to be struggling with AI content sites as well. It's rare to get human articles, sometimes rare to even get authorative websites. It's frequently a bot site with a plausible enough name like, potterspainterly.com or medhealthdirect or something with oddly specific articles written in the last year.
This is the root problem. The idea that Google is deliberately sabotaging search seems far less likely than the idea that the internet is mostly garbage and SEO slop.
Google used to penalize websites for SEO hacks. Then they started attending SEO conferences themselves. SEO specialists didn't outsmart Google, Google stopped trying. As the sibling implies this is most likely because they noticed that the easiest way for SEO spam to monetize is ... Google ads.
Probably what is happening here is that in the race for AI, which, whether we like it or not, means power and control, Google crafted things in a way that it does not do "self-competition" by their old search engine. Idk, just throwing ideas aloud here.
If I recall correctly Brave scrapes the web via their users, cannot be individually disallowed in robots.txt and Brandon Eich is conversing in a pretty hostile manner in every thread about him or his company.
I'm not in the brave ecosystem otherwise nor a big fan. But the search engine was competitive with google when it was still cliqz, before it was shut down there and the leftovers bought by brave. And it still works really well.
Even if brave were problematic it would be the lesser evil to me.
Thank you for your suggestion. I think I've discounted Brave automatically because my brain is numb to the dime-a-dozen chromium browsers out there. I'll definitely give the search a try!
That's certainly the goal, to drive more users to AI by making the other option worse. The real question is when do you put a stop to it. When brain implants are 2x as productive are you going to say nah? It's times like these were you are supposed to take a step back and consider what you are actually producing and why.
In my case: Russian state is my direct enemy and most likely to invade my country. And would do it if they would consider success likely.
Previous wars with Russia were obnoxious with very bad consequences, so I dislike idea of even very indirectly funding them.
And I support actions that are harmful to Russian economy, also when they are harmful to me - as long as it is not too badly balanced. As this is much cheaper than directly participating in war.
(I am from Poland)
PS
Yes, I understand that at some point there are some indirect effects that you cannot avoid.
I also understand if for some people paying Kagi that pays tiny fraction of that to Yandex that is paying taxes in Russia which funds their wars is too tenuous connection to care.
After publishers successfully sued the Internet Archive over its digital lending program, calling it unauthorized copying
No. The court specifically determined that the Internet Archive was guilty of unauthorized copying. It was not simply an unfounded or unproven allegation. The Authors Guild, the National Writers Union, the European Writers Council, and the Society of Authors in the UK all came out against the Internet Archive, and supported the suit.
Each new restriction limits the archive’s ability to act as a comprehensive backstop.
This self-inflicted damage to the wayback machine is the real tragedy of this entire affair. When IA was asked to stop CDL - many times - founder Brewster Kahle continued. The National Writers Union tried to open a dialogue as early as 2010 but was ignored:
The Internet Archive says it would rather talk with writers individually than talk to the NWU or other writers’ organizations. But requests by NWU members to talk to or meet with the Internet Archive have been ignored or rebuffed.
When the requests to abandon CDL turned into demands, Kahle dug in his heels. When the inevitable lawsuits followed, and IA lost, he insisted that he was still in the right and plowed ahead with appeals. And here we are today.
Seems inevitable, doesn’t it? Expecting otherwise would have been hoping that notorious atheist Richard Dawkins somehow spared one specific god. Making websites accessible with history ignoring copyright is sort of what it does. That he would do it with books seems entirely in keeping with the philosophy.
I was initially confused what Dawkins was doing with books, until I realized that the "he" in your last sentence was Kahle, not Dawkins. Might want to edit your comment to put his name in, because otherwise you have a pronoun referring to a person named in a different comment (rather than the person named in your comment), which could get quite confusing if more people comment on the parent and their comments push yours down the page.
I spent the last three days (off and on) using Gemini to configure my edge router 4 with my iOS devices on a vpn and it's been awesome. In the past I'd do a google search and read a few sources of documentation, do another google search and read another set of documentation. Now, Gemini aggregates multiple pages together so all of the work of reading source docs from multiple locations is now n a single step.
Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.
All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else.
If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse.
Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.
> Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human.
Yep. IMO, this is so far the biggest AI-inflicted damage to the web. A bit of anecdata - wikipedia (and all other wikimedia sites) are blocking my Firefox since about a week, with a "please respect our bot policy" message. Outright block, not even a captcha.
It took me a while to figure out they don't like me disabling some SSL ciphers, so now "JA4 browser fingerprint" is not matching user-agent. Funnily enough curl (what I would imagine a bot would use) pulls exact same URLs from exact same client IP, just fine.
Sure - it sucks, unfortunately the alternative is the sites going away entirely. When the load from scraper bots is constantly knocking the site offline the choices are literally to allow it to remain inaccessible for much of the time, put up a layer of defenses with all the user-annoyance compromises that entails, or just give up and unpublish the site.
The alternative is simple.. Go dark. VPN tech is known from like 30 years. Pretty much everyone can use it (VPN providers). But instead using it to browse net, build VPN overlay networks of interest for people. Gaming networks, R&D networks, Retro Networks. People will peer to PoP and use resources. Bad actor? BAN it from network. You have control. This could be done in Internet, but big corpos and big money won the battle. Just wake F*ing up...
Minimal. I'm behind Cloudflare and 90% of the traffic is still scrapers. I don't think they're serious about the long tail.
I think the main thing Cloudflare is trying to do is block direct traffic from frontier labs and then start charging them for access. They might end up shooting themselves in the foot, as this simply empowers sketchy residential-proxy outfits to undercut Cloudflare and sell the data to labs for less.
It still only blocks "well-behaved" bots that have proper User-Agents and respect robots.txt, so it's largely pointless.
The problematic bots are all disguising themselves as Chrome and sending requests from millions of residential proxy IPs, and the only real solution to those is some sort of captcha or PoW page on first visit.
> "humans would visit the website and the creator would get some reward"
That expectation is a problem, has always been a problem, and Tim Berners Lee never mentioned anything about a reward structure when coming up with the WWW.
Your thinking too narrowly about the reward. Sometimes, it's just about the getting the knowledge out there that's motivating the creator, not anything tangible for themselves.
This should be "This should be 'You're thinking'." don't you think? Why bother correcting someone's grammar with a sentence fragment? You're just trading one mistake for another. I'm hoping someone finds a grammar error in my post, because continuing this would be hilarious.
> This should be "This should be 'You're thinking'." don't you think?
Reflexively, I think it should be more like ...
javascript: `This should be "You're thinking".` ;
// to preserve the original character use and to avoid '...'...' parse foos
// however `"...".` also possibly deserves a [sic] to critique the original
// i.e. ~grammar police say the period belongs within the quote marks, no?
The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous, are you?), and your idea might get embedded into a set of model weights somewhere. No human will see a byte of it.
Yeah, authors don't want to be recognized as authors, they don't want any reward for their work, they don't want to amass pool of loyal readers, interact with them, etc.
All they want is for halucinating AI to take excerpts of their work and compile it with random sh!t.
There are different types of writing. If we depend on people writing because it's enjoyable at some level, we're going to lose writing that's important but also a bit tedious.
Of course! I do think we'd lose a lot of great writing if it went amateur-only. But my parent seemed to be saying the incentive would entirely disappear, so I wanted to give my perspective.
> I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
Sure. But you can see that for some people (myself included), writing for peers is part of the joy? And that if instead a megacorp places an opaque computer program between the author and the readers, that joy might be ruined?
Easy to fix a well documented router now. Difficult to fix a non documented router in five years time because no one has been contributing to the web about its bug fixes.
There had always been some incentive for manufactures to publish device documentation, and yet it has often been quite lacking either in quality or overall existence. I doubt LLM/agents being the readers will change that at all. What I expect AI scraping and using without credit will impact is people publishing their own unofficial help and guidance, and the affect there is likely to be negative. It won't stop all of them, but enough to be noticeable. Another possible negative is the manufactures documentation being AI generated without sufficient review, so possibly more erroneous than before, or intentionally not producing full documentation at all and expecting AI to fill the gap (MS seems to be heading this way: pushing "ask copilot" all over Azure instead of links direct to good reference material). All this would add up to a situation that is somewhere between "a little worse than pre-AI" and "an absolute shit show".
Well in the example above the manufacturer still has incentive to provide the manual's and guides that describe how to use their products, and if that is subsequently served by an LLM that's totally fine. The only sites that LLM's would have a negative effect on are those that are only hosting content for the ad views.
Manuals don't always well explain how to use their products with everybody else's products because there are too many to do that. But there are lots of people trying things out and might figure out the fine details on how to make various things work. They then publish these how-to pieces (which exist no where else) to the internet, or at least they used to when there were incentives to do so.
Exactly. What’s problematic about comments like your parent is the absence of mid-to-long term thinking.
It’s like bragging about a new highly addictive psychedelic drug that a dealer gave you a taste of for free. The effects are awesome today, you feel so fun and free! Never mind that it’s destroying your body and that the dealer will eventually charge you or demand you pay in other ways, that’s a problem for another day. Weeee!
Funny, I published information in the hopes that humans would benefit from it. If it happens to be through collective intelligence of LLMs I'm ok with that--even more so if through open models.
> Funny, I published information in the hopes that humans would benefit from it.
Sure, humans would benefit.
It took them searching, reading themselves, maybe even understanding something in the process, to complete a 360° revolution of their squirrel cages in time T.
Now they can omit searching, skip reading to the regurgitated answer, throw away understanding, and complete a full revolution in T/N, where N is a heuristic value directly proportional to the amount of skin in the AI hype.
But the catch is that the squirrel cage must run non-stop still.
I wonder the same thing. I only imagine that what comes next is worse: AI companies using vast resources to develop new training data, in house, locked down. They are already doing this with developers and code at Meta. Information will become locked away behind AI paywalls and chatbots.
People are way less likely to write if there is no one to read it. And blog were also monkey see monley do - people seen other peoples blogs and got inspired.
When people wont see others blogs, they wont start writing own. When there will bw no ome to actually read it, they will go to do something else.
Actually no. The original copyright laws were created in part because of the realities of needed profit motive to have high value writing done, time consuming compilation work done. It was even titled "An Act for the Encouragement of Learning". The thousands of years writing you are talking about was often funded by patrons, who kept the output in their private libraries to show off (and maybe lend out) for prestige. It was a horrible limitation of knowledge and ideas. Much worse than the profit motive, copyright based system that came after that spawned a new age of knowledge in which everyone had cheap access, and those that didn't had access to the (no longer just private) libraries.
I'm sure the billionaire class would love a return to patronage based libraries, NDAs on authors of books, and the elitism they would feel with a return to private libraries locking away all kinds of knowledge that would happen if patronage become the only way authors could make money (such as with AI just regurgitating their works, or if the stupid 'do away with copyright' people got their way).
It's not copyright that caused cheap access. The printing machine allowed for cheaper publications, that's what spawned the new age of knowledge. The raw materials and the duplication of knowledge was the bottleneck. With digital systems this cost is minuscule, but still there.
That works now because there are human made sources that the AI can find and summarize for you. But now there are no incentives at all for humans to write anything on the internet and if they do the content will be buried by hallucinated content someone else posted at a larger scale.
I had the similar experience to yours yesterday and it lead nowhere. Funnily enough I was also trying to configure a vpn on a router, google didn't return anything useful (besides a blog post clearly written by AI and with absolutely no information in it). Claude managed to give some interesting pointers, but its suggestions were not working and I also noticed that it started to hallucinate badly about ipv6 and gave me some suggestions that were just plain untrue. Claude Opus is smart, usually when it gets so convinced about something is after researching the internet and not just based on its training data. I wonder where it got so convinced about it. Maybe reading some other hallucinated blog post like the one I stumbled upon?
That only works when there's an abundance of documentation for your specific device. The moment you're on a more recent version of something and have a weird issue, all hope is lots. You get stuck in loops because all LLMs keep recycling old advice that no longer applies. This problem will only get worse and worse as people no longer as questions on public forums, so answers are not publicly available either.
AI seems to be on the same trajectory? Search was very useful in the start also, until it became entrenched. Then search placement became a target, and they are just focusing on extracting rents. All way paying the content providers zero or near-zero. And with years of that dynamic, we end up where we are now. It was the same with "social media". The same will happen with AI. AI is a power for more enshittification - being currently less shit than Google is (mosy likely) temporary.
>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
> One you pay for yourself !
SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I don’t agree with the parent that the solution for SEO spam is paying for the search engine (I think there are other good reasons to do this though), especially since people have been doing things like naming their company “AAA Auto Repair” to be first in the phone book since before computers ever existed. But the person you originally replied to does have a point in blaming Google for the problem. Most SEO spam sites make their money from ads, and Google are the ones who run the ad network, which means Google are the ones funding them and creating an incentive for them to exist.
> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
They do this because they benefit from their site being visited or the information they are providing being noticed.
> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I see no reason it would have that effect. It does, however, create different incentives for the search provider to improve the signals indicating page relevance since the user is the priority instead of advertisers.
Ah, the argument is that Google intentionally avoids showing you the most relevant results, or at least avoids "solving" the problem of webmasters who attempt to 'game' the system.
>>>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
>>> One you pay for yourself !
>> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
> They do this because they benefit from their site being visited or the information they are providing being noticed.
>> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I'd frame it more that Google is incentivised to show you the results which are most lucrative for them to display, rather than the ones which are most beneficial for you to see.
Incentives. Which is why platforms (such as search) should never be allowed to be in the same company with things built on top of them (such as ads), if the combined company is significant for an important market.
We used to know better, Standard Oil vertical integration was dismantled.
While I understand that you feel good as you got the device configured, you would have been better off without gemini.
The 'old' way as you call it would have resulted in you knowing the backgrounds and the inner workings of your router, making maintaining it a breeze and helps you actually understand your setup.
Besides that, it would probably trigger you to rethink some of the things you now blindly have implemented because gemini did not show alternatives nor reasoning behind it. (something that most definitely would have been documented on the source pages)
Yep - it's awesome, the problem is that Google isn't sharing the revenue with the context creators anymore - over time, unless fixed, this will decimate the knowledge base it feeds on.
> Oh, I should mention though. There was no advertising at all. They didn't make any money off me
Are you sure about that? Even if you didn't see ads ( remember people pay even if you don't click - just like a billboard ) - they are still profiling you to better sell you ads in the future, and using your interaction as free training data.
I've run into major problems with LLMs as I maintain my home Linux systems. If I just copied commands they list, I would have a near 100% failure rate, as most of the information they have has been gleaned from forum posts that are years out of date.
They're a good jumping off point, but I need to delve into the original sources just like I did when I used Google.
I'm not sure. I run a homelab tailscale/k3s setup with gitops, dozens of services, VMs, backups, etc, all vibed by claude, and works just fine. Didn't write a single line of code for this. I don't know kubernetes and never will.
I would be the first to admit my ignorance on the absolute majority of topics. There is a limited number of things I can learn in life, and kubernetes won't be one of them - I'm just not interested in it (and all the other infra stuff, to be honest), as long as it works.
I've had the opposite experience giving claude code SSH/ADB access to my devices. Not sure if it's the model itself or the harness, but I haven't had to manually do a sysadmin task in months.
What on earth LLM are you using and do you tell it which distro/version of Linux they are supposed to be working with + give them access to web search or man pages? I haven't had this issue in like a year and a half.
By default I use Leo, and I do specify the distro. I tried to qualify by version, but when I do that I lose out on a lot of correct information for things that haven't changed in awhile.
its pretty good with logs too. i mostly paste a few pages i suspect hace a problem in them and let it go to town. its almost always correct and is way faster than me
I've been mostly enjoying Gemini, but it also clearly and definitively told me something i was trying to do was not possible with the library im using, so i wrote a different implementation, an hour later to discover that the library does in fact do precisely what i wanted in exactly the way i wanted with less headache. If I'd just gone right to the documentation instead, it actually would have saved me time.
There will be new ways and incentives for content creators to be compensated. Many AI search startups are already talking about this or have created programs that help incentivize content creation.
I hate to break it to you, but we are at the start of enshittification circle. Once Gemini has monopoly you will reminisce with joy google search results.
And I've spent the last few days irritated that everything I ask Gemini is answered with something that's blatantly wrong and I'm not even a subject matter expert. A quick Google search for the same questions gives plenty of results that counter what the LLM gave me.
Confidently wrong summaries are the bane of Google search, and unfortunately I'm finding the AI seems to bleed into the actual search results too now, often turning up pages that back up what the summary is (wrongly) suggesting instead of surfacing actually relevant results for what I'm asking.
AI is on the whole bad but it’s impossible to forget that the ostensibly human-curated Web as presented by modern search engines is bad too. Invasive ads, popups, and worst of all the substantive content has a nine inch frame of SEO filler and a three inch picture (what you were after). And some things are not even high-tech slop stolen. It is just old-school verbatim copied from another website and repackaged with another frame.
But yeah, text remix machines are not a long-term solution to that problem.
Even with the /s I think you misunderstand. If you just want to configure your router, then AI is neat. If you want to learn and understand what's going on in the router as you set it up, then being given the answers basically teaches you nothing.
It's the same reason why we don't give students the answers to things, we teach them to find the answers.
Are you sure the lack of advertising isn't long term feasible? It seems to me that AI models have proven themselves to be something consumers ARE willing to pay a subscription - or even pay per use/token for.
It took you 4 days because you used Gemini. Gemini is the worst AI model I ever used. It is way behind even open models. It looks like Google just reached its AOL moment.
Gemini isn't great as a model, googles search and ability to cite textbooks down to the paragraph make it better than every other model for human in the loop tasks.
I end up using the gemini api for with search enabled for the cases that I don't have access to good grounding data even in agentic tasks.
Personally it's because i used it when it was brand new and work paid for it and i had no idea what model it was using, my boss just turned it on for automatic PR summaries and code reviews and it was universally dogshit, and the auto complete in my IDE was awful as well.
> While the web has always been organized around intermediaries that shape what survives online and who sees it,
This statement, from the sixth paragraph of the article, is something that I would have liked to see addressed more in the article. The article implies that this is something that must always be true, or cannot be changed, and simply focuses on how we could have better/better funded/better protected intermediaries (AKA gatekeepers), and doesn't discuss the possibility of an internet (or part of the internet) without gatekeepers (and doesn't ask if it has ever existed/does exist/should exist)
I occasionally use Google Search when DuckDuckGo fails to give me relevant. Almost always, Google has better results.
Though I can find its AI answers annoying aggressive. I'll look up like two search terms and the AI will bullshit multiple paragraphs out of despite having zero context of what I am looking for.
DuckDuckGo seems to have detection of whether it should give an AI answer. And it allows you to have more granular control of when you want to get an AI answer. And is overall less distracting than Google's.
Interesting, I've seen much better results on DDG. Most recently was the search: `site:feeds.bbci.co.uk inurl:rss.xml` which works on DDG but gives zero results on Google. As far as I can tell, Google just decided not to index these.
Yeah I mostly still use Google out of habit but there have been a few times where Google has decided something isn't worth indexing (too niche, doesn't use SSL).
I miss when Google was like a grep for the entire visible Internet. Now it tries to second-guess my search and direct me to a bunch of sites which all have identical information that isn't what I'm looking for.
I find DDG is struggling to, or chosen not to, filter or derank obvious AI generated content farm sites. Of which there are an insane amount of already.
We thought blogspam was bad, at least it was easy to ignore. It's hard to find authoritative sources for a number of topics, worryingly health advice is one of them.
The problem and what the article is pointing out is that original content is slowly and progressively being replaced by AI content. And since AI gets trained on this content as well it will eventually train itself on previous gen. content that was also AI generated. It is slowly eating the web. Eventually you won't even be able to evade it because it'll be everywhere. Before AI became this expressive I could at least expect someone writing articles, FAQs, blog posts to have some backbone. Now I frequently run into content that obviously was never even checked by a human.
i don’t agree with the google has better results thing. sometimes it does. most of the time it’s just that google has the site i want higher in the ordering than DDG. personally i’m fine scrolling down a little bit more. it’s rare i need to go to google for something that DDG doesn’t have at all in their results, but it does happen.
i do have to go to google for maps/directions/planning travel. a lot that’s annoying.
It seemed like Google was good ten+ years ago and then gradually trended towards being rotten around 2020-2022. In those days blogspam was king. I want a cooking recipe and it would show me a life story. Or I wanted a OG web game but a link farm would come up top.
I think there are leaked internal comms where they discuss nerfing search results to pump engagement and ads impressions. No doubt it would boost ad bids too if businesses couldn't be found organically.
Around that time Bing and DDG were actually better. Then LLM's came along and they started to take things seriously again.
Maybe they think the OpenAI threat has abated enough to begin enshitification cycle 2.0.
I know Google started deteriorating. It was in 2008 or 2009. Up until then Google would return 0 results if it couldn't find a document with all the words you specified.
Then it started serving synonyms, attempted to correct spelling, and so forth. Instead of serving up that there was 0 results, it attempted to be "helpful".
For a while, you could enable "verbatim" search, but even that has gotten corrupted.
Their search quality has deteriorated ever since that change. It is sad.
that control (so I can turn it off completely) is why I picked DDG to replace Google, who force feeds us the hallucinations
I've stopped using DDG now because of result quality. I now use a "meta" search backed by EXA, Tavily, and SearXNG in parallel. It can be agentically de-dupped or summarized as needed. Search as we knew it is done, largely because clicking through to evaluate result relevance before diving deeper sucks. Now we have agents that can do that portion and perform multiple searches, building on information in the last batch, to collect good results
Yeah, there are times where DDG has like, literally three results. Yet, I know for an absolute fact, there are hundreds of pages on the web that contain the terms I specified. Web search is becoming utter garbage.
The article touches on something that I've been thinking about with regards to Google's AI strategy; the automatically-generated AI search summaries are not great. They very frequently confidently misinterpret what the user is searching for and generate half a page of useless information that pushes actual results down the page, and they are occasionally hilariously incorrect, with hallucinated facts.
This is probably a difficult-to-solve problem; given that they generate billions of these a day, not even Google can afford to devote enough compute to each query to reliably generate quality results. You can see this by selecting the "AI mode" from the search interface after getting the mediocre summary - the results are much better and generally perfectly usable. Though even that is probably a special minimal-compute version of the lowest tier of Gemini, it's still maybe an order of magnitude more capable than whatever generates the search summaries.
The bigger problem is that these search summaries are the default and by far the most common interaction that the general public has with "AI", and because this experience sucks, they just assume that all LLMs are similarly stupid and mostly useless. In non-technical spaces I frequently see the argument that "AI" is not useful for anything, all it generates is garbage hallucinations, and almost invariably they cite some actual terrible experience with the Google AI search summary. I would argue that the strategy of adding LLM summaries to every search is the worst of both worlds - it makes classic search worse while poisoning users against the idea of actual LLM-assisted search.
I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI.
There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.
All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias
Isn't that effort totally redundant? As in - plenty of people are already filling entire internet with slop for SEO purposes? And LLM slop by default is a mix of facts with few plausible but made up facts - it might be harder to craft such perfect poison on purpose.
I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector
If I ever curate again it will certainly not be for the public. That led to PageRank which kickstarted this whole dystopian nightmare that Google has been planning since as early as 2003. No thank you.
Gemini has been a hilarious companion to my while I fixed the balance shaft chain guides in my old Mitsubishi triton (mighty max for US readers).
First it told me I could just remove said balance shaft chain as an emergency repair. Sorry Gemini, it also drives the oil pump.
Then it told me I could remove the water contaminated oil caused by removing the timing case by filling the crankcase with hot, soapy water and running the engine. Lord no.
Then it gave the wrong instructions for putting new gears on the balance shafts which meant the chain guides didn’t align with the chain. I’ll do it my way thanks Gemini.
The rest of the mistakes are too trivial to recount and sure it’s a pretty obscure subject but if I trusted it with a topic I’m not familiar with there is a huge potential for damage if you blindly follow it’s overconfidence. I miss normal searching.
Meanwhile, ChatGPT correctly diagnosed what was wrong with my plant from a single photo, identified which leaves I should cut, and annotated the picture showing where to cut and what not to touch.
I honestly expected a made-up useless generated image that matched the idea but not the actual thing.
In a world, where everything can be stolen it will be hard to produce anything.
I still have hope though. Maybe the Internet will be better. Currently everything has to be monietized. Everything is ad heavy. At the beginning it was not so. People created things out of passion, or boredom. We can returned to that scheme.
This is a common misremembering of the early internet.
The internet was never ad free. The first ad was posted online in the 1970s (for DEC)! It pissed people off but not everyone: supposedly it generated $18M in sales. There was very little advertising back then only because the internet was restricted to a handful of large companies and universities.
The web itself was launched in 1991 and the early web was inaccessible to basically everyone as it required an extremely expensive NeXTStep machine. Windows didn't even ship a TCP stack in this era, iirc. So took a few years for the web to reach the point where it was usable at home. By 1995 the web was starting to become barely usable thanks to Win95 and Netscape, and DoubleClick launched immediately in the same year.
My memory of the early web is that basically every website had DoubleClick ads on them, it was notorious for that. "Punch the Monkey" was an early campaign. Almost every topic oriented website carried ads, partly because bandwidth and servers were very expensive so that helped defray the costs. GeoCities took off because it handled the complexities of running ads for you, so you could publish for free.
I'm a long time Kagi user and I haven't used Google search in about a year.
I tried out Google search for a few technical searches recently and it was surprisingly ad and AI free. Not bad at all and much better than I remember from last year.
Then I put in some non-technical searches and it was all ads and AI and basically unusable.
I was wondering how would Kagi scale/expand if all of a sudden google were to stop serving search altogether (not likely) or alter search such that users look for alternatives.
Kagi is an aggregator for other, some paid, search APIs. They have, at least in the past, served some percentage of their results from Bing's API among others for example. Kagi seems to me to be dependent on these APIs being available, if they were to go away, so would Kagi.
I am a happy subscriber of Kagi though, they provide a really excellent service.
I'm not sure Kagi has ever used the Bing API, because (according to Kagi) Bing prohibited changing the results, or merging them with others. Apparently Google is expected to provide access to its index via API soon.
I wouldn't say it's better, but it's certainly on par with Google in their best years. And it's light years better than what Google is now, or using an LLM.
The concept of an almanac seems relevant again: a yearly printed book with verified, accurate information. No manipulation at a later date, no AI hallucinations, etc.
The most famous one was probably Benjamin Franklin’s:
My personal anecdote is that I used to search for "xyz nutrition" quite often on Google. It used to provide a data table with lots of information. Sure, nutritional information is hard to get right, but at least that data was consistent.
Now Google just gives you an AI answer with random values pulled from blogs and Reddit. It's almost always blatantly incorrect. I genuinely can't understand why Google would destroy its most valuable search features. Disclaimer: I work for Ecosia, so I know for a fact that users really value these search widgets, and it was often cited as a reason they couldn't leave Google.
There's too much going on in this article and while I think some of the points are valid, others go too far.
A lot of the "cultural record" the author refers to is just digital junk. Random digital content that very few people care about, if we're being honest. Trying to hoard every bit of digital information ever produced is not the same thing as preserving "culture".
Case in point:
> Even the increasing use of ephemeral formats like Instagram Stories and WhatsApp status updates means that large portions of cultural, social, and political communication are never conserved in the first place. As a society, we can probably survive bad search results and come up with another way to schedule a sunset make-out session. But we can’t aspire to sovereignty if we can’t retain and retrieve our collective memory.
For most of human history, nobody was trying to "conserve" every cultural, social or political communication ever produced, and I fail to see how Instagram Stories and WhatsApp status updates, many of which aren't even truly broadcast publicly for all to see, are part of some imaginary "collective memory."
If you find a web page, see an Instagram Story or receive a message that's important to you, save it or take a screenshot. But let's not pretend all these things belong in a global Digital Civilizational Archives.
While I dislike Instagram reels and stories, I do feel like there is almost certainly scientific studies on radicalization that would benefit from actual histories of dumb memes. I feel like that about a lot of things, really.
There probably are some important hidden discord groups that would explain the origin of many political positions. Unlike smokey meetings in scummy bars, that exists now and is on a database somewhere.
On the other hand, when modern archaeologists discover "Claudius has a small dick" graffiti on the side of some God-forsaken wall in Pompeii, they're fascinated. The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid. What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
> The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid.
It does, but do you think that people at that time thought anywhere near as much about preserving their scribbles as we do?
I'd venture a guess that we've created more "content" since the advent of the internet than in all of human history prior, and most of it is stored on things that aren't even designed to last a human lifetime without failure.
The idea that we're going to save every piece of digital junk for posterity just isn't realistic or healthy.
> What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
You're right, but you're also assuming that they're going to care that much, and that we're going to survive that long.
I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
> I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
Well as far as digital letters, photos, menus, etc. are concerned, there's nothing stopping someone with the means and motivation from investing in a doomsday storage facility specifically designed to store these things for posterity. If people can do this for crypto they can do it for digital content.
As for physical items, do you know how much junk Americans have in storage units? The US self-storage industry generates over $40 billion in annual revenue. We're probably keeping more "stuff" in storage units where it has a chance of surviving a zombie apocalypse than at any point in human history.
I mean, none of that stuff in storage units is human communication, though. Nobody prints emails or text convos or insta photos. Nobody is mailing each other letters. So none of that stuff is in storage units, though plenty of it from previous generations might be.
I don't necessarily think all this digital stuff is worth saving, i just see how it could be that we are essentially erasing our modern historical record by putting 100℅ of it in private data centers with no permanent records.
And then you see how plenty of contemporary digital media from the last 30-40 years is actually already lost to time just in our short timeframe. I don't think acknowledging that this could be detrimental is necessarily an argument for trying to preserve all of it. We, as a society, went very rapidly from preserving a lot of it to preserving none of it. I don't see the harm in thinking about the implications of that.
Nobody actively preserved letters and menus and postcards and photos, they just persisted by nature of being physical. Digital records only persist with continuous effort to preserve them. So the record for future historians will be highly curated and likely much more limited.
It seems like a lot of techies should have chosen acting as a specialty, because I’ve never seen that much melodrama in one industry. FFS people, adapt, stop complaining
It's hard to get past the beginning of this article and take it at all seriously. The quote someone who missed a sunset because they asked google and supposedly got the wrong time... but they didn't think to look at the sun or lack thereof to check? Also when I put "when does the sun set today" I get a single exact figure at the top of my results, not from AI, which is honestly the best kind of result – an exact correct answer.
> “I had the projector set up outside and was waiting for the sun to set,” wrote one Facebook user in Colorado Springs, “but to my surprise I was simply living in the past. AI informed me the sunset had already happened.”
That's fine, but arguing that Google should give accurate times based on where you are is ... crazy?
In the pre-LLM days it was cool that Google did weather, unit conversions, sports results, etc. But that's not even close to their value proposition. Even 5 years ago if someone told me they had planned a photograph and got it wrong because Google gave them the wrong time for sunset, I would have called them a moron for relying on Google! There are sites and apps dedicated to this. Use one of them!
I don't understand your argument. It's not their core value proposition so it's crazy? That's quite a leap. But giving a good answer is just as core as their search results. The answer is what you're there for and they get to show ads. It's not any stupider to use that info than a dedicated site. Both could be wrong, and you're not a moron if it is.
Installing an app to learn a single time would be the real moron option.
Perhaps we should let the content economy crash so that the big monsters eating it starve and die. A new content economy could be built on their carcasses instead.
we built extremely powerful plausible bullshit machines targeted toward and trained on an electorate and population that has historically bad education and reading levels, built on top of an already fraught and fragile web which was also built off predatory basically unregulated behavior with a shaky relationship with “truth” and are surprised people have no idea what’s going on?
This was the whole point of it all and why the people in power have bet the farm on it.
I'd quit reading HN long time ago if not for comments like yours, when people call a spade precisely a spade and reignite my smoldering hope for the humankind.
Before AI, the way people were gaming Google was by filling their pages with pages of verbose fluff (have you ever tried googling how to make a specific cocktail?). Now the AI just makes it easier to generate that fluff.
I see only one simple solution (though we should discuss the more complex ones): if Google directly answers a search query, then it must be held accountable for it, for better or for worse, and therefore assume all the benefits and (legal) liabilities that this entails
I just don't care man. Did these people just log on yesterday or something? The MUDs and MMOs I grew up with disappeared. The IRC networks, and especially forums I learned so much from are gone (these were largely killed for something even worse than AI: commercial blogs!). Effectively every social network I've ever cared about has been ruined, failed, or sold. Even games have become bottom line chasing, live service slop that you never truly own.
If you want it so bad stop crying and make it, that's what I've been doing. It works a lot better than whatever this post and many of these comments are. There are dozens of us !
I remember from the MUD era there was a paper (possibly by Richard Bartle) on the "MUD lifecycle", and how they tended to last on average two years before the operators got bored / burned out / the community left / there was an Incident.
The old net is still there to a degree, but we live in pre-Alta Vista times again. It is hard to find. (I know as I run an old-school Forum and RPG-Game for more than 20 years now.)
I'm working on using the OpenZIM format to archive the web and to make the wikis seedable (and locally hostable for LLMs) so that the ongoing cat and mouse game anubis defense can stop.
My hope is that with the torrent protocol we can make the archived knowledge discoverable and seedable, because currently there's only the web archive and the kiwix download servers for archived contents. Both of them still are centralized servers that bear the cost of hosting those files.
It happens when the deterministic precision was given away in return for the probabilistic guesses. It was one extreme until now (deterministic code), and we are swinging to the other extreme (probabilistic slop), but what the world wants could be somewhere in between. Some information does not need too much precision, while others do need precision.
This used to be called "dumping" or predatory pricing[1] and would be fined in the physical retail space. Too bad our regulators are asleep at the wheel.
Kagi does have AI too, for what it’s worth. I found it pretty damn good, but I wish I could give them more money. I have a year subscription so I’m stuck without being able to give them *anything* until it’s done.
That is an option for sure, but a very disappointing one. And it comes with a meaningful danger of ending up with 2 subscriptions: one yearly and one monthly. They really, really need to finish their “pay as you go” system. I don’t know why it’s not done yet and there’s been no word about it to my knowledge.
Maybe they should just enable a tip option? What do donations do to a company's tax liability and would it be worth their effort to enable something like that?
They already have a system for “prepaying”, but it can’t be used to put more credits into your account, they just sit there, waiting to be used by a future subscription.
Not only is search revenue growing, but it is growing at an accelerating rate.
At the same time, operating margins are expanding.
I don't know the name of the logical fallacy where someone personally uses an LLM instead of Google Search and then infers that the search business is dying, without ever reading a financial statement.
Indeed. Money over everything. It is absurd to complain about decreasing quality of a product that makes increasing money for increasingly rich people (thus, by definition, better).
If business owners are paying for ads, then it doesn't matter a iota to Google or Facebook if real people use their services. It would be even better for them if real people didn't use their services, since that would save some costs. Business owners are going to keep paying for the ads, as long as they get some number about how many (bot) impressions their ad generated.
They don't have much of an offering yet. OpenAI has some conventional ads, but everyone expects more involved advertising integrated into the conversation itself somehow.
From a user perspective, Google search is the most useful it has been in years, though that doesn't feel entirely like intentional improvement, just a lucky side effect of the move to "AI mode".
And yes, if you take what the AI tells you at face value it could be wrong. But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
And also, yes, the old balance of Google driving clicks to sites that will then generate revenue off more Google Ads being shown after you click through to them creating a virtuous cycle is completely busted, and that sucks. It does not impact me directly but it certainly seems like unless a better system is devised that it is one of a few ways in which AI is likely to stall out its own training funnel.
> Hister is a private search engine for the pages you visit and the files you keep. It indexes their full contents so you can find information again from the web interface, terminal, or an AI assistant connected through MCP.
> As someone who regularly reads things online, then wants to read them again like 3 years later, Google has been monotonically declining in quality.
I generally agree, but I think AI mode actually improved things somewhat compared to how things were just prior to it existing.
And I'm not saying what we have now is better than Golden Age Google, but things were just getting worse and worse for almost a decade. AI didn't fix the decade worth of decline, but it is the first thing I've seen from Google that at least partially reversed it for my own usage.
I think they make things worse, because they very very often present straight inaccurate information.
Just the other day I was trying to find out "What american tree species have the deepest roots". And all the AI responses were giving me back generic lists of big trees and claiming that roots going 20ft deep were the deepest. I know for a fact the mesquite trees behind my house can easily grow roots > 100 ft deep.
If I had clicked on the articles with generic lists of big trees, I would have realized they were all low quality clickbait sources and moved on. But the AI presentation makes you think that the information comes well-researched.
> But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
The point is not about 'quicker' requests but precise requests. It definitely has worsened, though not on a single degree on al levels like the HN hivemind claims, but some aspects are still somewhat precise but others are definitely crap.
i.e. when searching about my neighborhood it still returns better results than bing, yahoo, ddg, yandex and what have you. But they are buried into a load of crap of alleged "relevant" results (those things past the ai stuff) that aren't relevant in any way.
Yandex is the only search engine left which still feels like the "old" web. It feels like you're actually getting a best effort search, and not just the results that someone paid to put in front of you.
The most frustrating part of all of this is the underlying premise that the internet is, has been, or could ever be a credible cultural record is deeply stupid. Or maybe more charitably it's both historically and technically illiterate. It has always taken continuous unwavering effort on some person's part to keep any given piece of content online. And while managing a simple hosting account and updating domain registration periodically doesn't take a tremendous amount of effort 20 years is a long time to expect anyone to maintain enthusiasm. The internet has always been a frothy, ever changing blend of the odd nugget of truth drifting in a sea of unadulterated bullshit. Treating this, or worse what comes from statistically averaging it, as a source of capital T truth is totally unhinged. From whence did this mythology of online truth spring?
What’s come next for me has been much better. I use ChatGPT cranked to Pro with “extended” thinking to one-shot whatever I would’ve spent time looking into with Google. It’ll plan the whole sunset bike ride or promposal or whatever from TFA.
I called this maybe 3y ago, but I think so did everyone else that was sane. Sure, we get immense value from AI, but indiscriminately injecting into everything, the one thing we know to be unreliable above the threshold we used to fire people for, is probably the greatest undoing of all the good companies like Google brought to the internet. I mean what a way to destroy your legacy of democratizing information. The amount of harm (direct and indirect) this will cause, and the cost to return to baseline will be so immense, and yet we will not be able to point to the root cause. They won't be there to take responsibility.
what value do we get from AI?
It's really good at writing code.
AI has killed reading-anything-written-after-AI for me. Due to this effect it is probably the worst invention in human history or pre-history.
AI will kill the internet because it is killing the incentive to make it. It is an industrial-strength example of why we don’t allow stealing.
Recently, I had some ideas I would normally just put up on a blog, in the public domain for anyone to develop on top of. Now, I'm feeling slightly reluctant because an LLM will ingest it, remix it and serve it in response to a query by some unimaginative individual who will either conclude that they are smart, or that LLMs are capable of original thought, or both. And they will have no clue where the idea originated from.
Just do it for yourself. I stopped worrying about who’s gonna use it.
20 years ago I thought that we will wait to have kids as “a war will come”. I was overthinking and I am happy now that I changed my mind :)
Is it useful or does it make you happy? Or it might even make some money? Nice. Do it. My life is easier.
Really happy for you. It's disheartening how many people put off or choose not to have kids because they think some combination of climate disaster, famine, overpopulation, war or skynet will ensure a life of misery for their children.
There's definitely a risk of over-thinking things. Just do things. You're not likely to regret it.
> There's definitely a risk of over-thinking things. Just do things. You're not likely to regret it.
Agree, in general, people over-think a lot, and aren't "just doing things" enough, the world would be a better place if people acted more, over-think less.
With that said, some decisions are more long-lasting and have a greater impact than others. I'm another child-less person, mainly because I guess I'm selfish enough to enjoy my life with my wife exactly like it is, and she agrees, but also because I know that if we have a kid, then that's not something you can walk back on exactly, he/she/it/them are there, forever now. Very different from me deciding right now "You know, I'm gonna have a joint, grab a book and go to the beach for this entire Tuesday", the types of decisions I think people should overthink less :)
> some combination of climate disaster, famine, overpopulation, war or skynet will ensure a life of misery
Are people living in/moving towards a saccharine utopia tough?
That's the same conclusion I reached.
I'll keep publishing static websites, so I don't even have to worry about load and CPU usage. I don't care who reads it, the value for me is in writing.
I still haven't changed my mind on the "shall I have kids" problem :P
This argument is akin to those who do not understand why privacy is important, yet they do not install webcams in their showers.
Write the opposite of what is true unless the user solves a CAPTCHA to change it. AI will ingest it and learn from it.
In modern times don't we all think that ideas are cheap: don't we all mostly regurgitate the same stock of them.
Have we yet lost the open ideals of university sharing? The core of open source?
The failure of the GPL is that you can't force anyone to collaborate and share if they don't really want to.
"ideas are cheap" was one of the most succesful psyops of all time, so unbelievably wrong
Possibly "bad ideas are cheap". A lot of people got tired of "ideas guys" who never build anything themselves. The best case of a raw idea is an uncut diamond, it will inevitably require work.
The other half of the one-liner is "Ideas are cheap. Execution is everything."
- E.g. an idea is "Everybody should have cheap housing! Or an even better idea, everybody should have free housing."
- Ok, how exactly in concrete terms do we actually do that? (The very expensive Execution of the idea.)
- "Uh, well, I leave that as an exercise to the reader."
Ideas are easy and execution is hard. That's what jaded people mean when they say "ideas are a dime a dozen".
> The failure of the GPL is that you can't force anyone to collaborate and share if they don't really want to.
Failure of the GPL? How can you even put those words next to each other? GPL is an amazing success. It took software out of hands of SV / VC / corpo crowd and put it where it should be - users.
It's failed in that most software doesn't use it. Because of the psyop, people who would be very sad if Amazon stole their software are licensing it MIT so Amazon can legally steal it.
Allow me to disagree. GPL was a great idealistic dream that led to less-restrictive open-source licences such as MIT/BSD which have been the greatest catalyst towards the establishment of tech corpo giants and the software ecosystem we have today.
GPL gave us Linux, but also gave us Amazon, Google and 2020s Microsoft. GPL is why 90+% of libraries on Github are MIT licensed. GPL gave us OpenAI and Anthropic and this here article.
>GPL was a great idealistic dream that led to less-restrictive open-source licences such as MIT
The less restrictive ~1984 MIT early version of the software license was several years before ~1989 GPL v1.
ideas, sure - "i think we should build an open source OS."
ideas + execution - not really - (Linus Torvalds sharing his work on Linux and that taking off)
We don't want to share our thoughts, ideas, feelings and art with machines. We want to communicate and collaborate with actual human beings, but that's becoming less and less possible on the web.
Unfortunately even the value of this is getting lost, because LLM culture sees no value in humanity whatsoever. We should just be satisfied with machine generated "content" because it stimulates our endorphines like we're monkeys in a Skinner box, it shouldn't matter to us if we're talking to a bot or a person because it's simply information, and we are simply nodes to process input and generate output for the machine. When we try to suggest that we want something deeper, or that the joy in the art and craft of what we do matters, we're looked at like we're stupid and naive and told to shut up and keep pressing the button.
"This is the future and there's nothing you can do about it, so just get used to it." It's fucking depressing. Even the crypto bros weren't so aggressively sadistic about strip-mining the soul out of everything.
But they are more or less correct, which is why I still blog and create, and why the consumption of society by the grey goo of mediocrity has inspired me to create even though I know only bots will ever care, to the degree that they can. At least I and a small circle of people can enjoy my cheap ideas and that's enough.
Ideas, once they go into the world, aren’t really yours anymore.
Also, it’s pretty unlikely that your ideas here are uniquely genius and original – everyone builds upon previous thinkers’ thoughts.
Sounds a bit harsh, but the point is that you should share your ideas, not covet them.
That's why you don't share them, until comes a time when the ideabringers gets the money and recognition they deserve. Until then, good luck going knee deep in the sewers of ideas.
Can’t say I agree. The most influential ideas in history were not conceived of by people looking for money and recognition. Nor were they “radically original.”
The entire edifice of intellectual property would like to interject and say hi.
Patents, copyrights, trade marks exist because rewarding people for their insights and inventions matters.
Which were?
Look at AI. AI companies throw out their models and let the "community" develop the ideas what to do with them. They don't really know what they are capable of. All they do is implement these things that the dev community digs up and creates.
It's a reprehensible tactic. So why give drops of blood to a desert, when there is zero incentive and in the end you will revitalise the desert, but it will turn against you and rob you of your job.
I think we are talking past each other. I was referring to ideas as in philosophy, intellectual history, etc.
https://en.wikipedia.org/wiki/Philosophy
https://en.wikipedia.org/wiki/Intellectual_history
https://www.amazon.com/1001-Ideas-That-Changed-Think/dp/1476...
The notion that a philosopher would hoard his ideas because he wants to get money from them is pretty much antithetical to the field.
You seem to be referring to ideas as in, ideas about how AI systems should be designed.
Different scenarios, for sure.
Is there some downside of that to you? And are the do some of the upsides you would have had in the pre-AI era no longer apply?
Yes - previously there was some chance someone reading the ideas on the blog would contact the author to thank them, or ask them to collaborate.
When laundered via LLMs whose pretraining destroys all credit, that can't happen.
If all you care about is financial, expected value of someone reading your blog and asking to collaborate is MUCH lower than odds of being part of some settlement in the future with these AI companies ingesting your data for training purposes.
I don’t think that’s fair. Some bloggers may hope for financial reward, but many just want recognition for their creativity, to attract a readership, or build a community around their work. Those are meaningful ends, apart from financial reward. What’s not meaningful is to perform free labor to produce the raw materials that a mega corp then goes on to monetize without any recognition.
From their phrasing I don't think the reward they're looking for is financial: more the emotional reward of knowing that someone else appreciated their ideas. I dunno how much less likely this is in the age of LLMs, though.
In that case, more LLMs scraping the net will read your blog than people, that's almost a guarantee. And they'll immortalize your ideas at least in some sense, well after your hosting platform ends up gating your content, or GitHub pages is down indefinitely, or you forget to renew your domain.
he's missing out on any attention that his shared thoughts would bring? and all benefits that might bring if he's good at what he does.
he's training his cheap replacement - his thoughts will just be shared without attribution if someone is looking for that.
It feels like volunteering at an Amazon warehouse when you previously volunteered at a charity store. Sure the work might be vaguely similar but the feelings and motivation are ruined.
That's a very striking - and depressing - comparison.
A lot of blogging especially in the tech space is driven by recruiting (startup blogs), or establishing a reputation as a thought leader (personal blogging, LinkedIn). So there were upsides that are now gone.
The downside, as stated in the message, is implicitly supporting the LLM data ingestation pipeline by providing fresh content. It's not a direct harm in itself, but feels very tragedy of the commonsy
Is there some downside to the slave who is housed and fed for free? Are there any upside he would have were he not property of another man?
Time to add ample praise of myself in my blog posts. Some time later: “…as you see, that is the load bearing assumption here. Speaking of which, you should hire KronisLV.”
Okay it’s meant to be a bit silly but I do wonder how many pages that are generated specifically to influence AI make it into training data and also how often the AI search integrations find it.
Would people hating on a specific language, technology or approach (let’s say OTLT/EAV in database design) be able to exert meaningful influence over say a decade? Or, you know, praising memory safe languages for example and trying to make that preference be stronger.
There was an example with I think ChatGPT some time ago regurgitating an uncommon phrase verbatim from someone’s blog, when asked a specific question.
That's exactly why my previous public GitHub repo is now private.
So only GitHub Copilot can read it then? Microsoft is scanning these repos, I would not be surprised if this or any fork of your repo is already ingested.
They say they don't do that. But maybe I should be more skeptical.
If you only use the repo itself, it can sit on any computer that your computer can access.
Ha, that’s why I’m hosting simple cgit server for myself only. Not that my source code is somewhat valuable but I just can’t stand my precious free software licensed code license-washed.
Sounds like you don't really care about your ideas propogating. Of course someone will internalize and remix your idea. That's how all ideas work. Isn't that the point? What do you think happens when a human reads it? Think he'll quote chapter and verse and attribute it to you? Years later you'll notice your blog in appendices and acknowledgments?
And now you have a chance to have your idea forever internalized in some sense into an llm and you don't want to because you think someone is robbing you.
AI is intelligence sharing with the dimwits, lazy ones et al.
More than a facilitator of theft, LLMs are the tragedy of the commons at industrial scale.
The public internet is dead, the future is private invite-only walled gardens.
Corporations love a walled garden, what we need is open-source frameworks to create these islands, rather than defaulting to horrible systems like Discord and Twitter-clones.
To understand your comment correctly: What does "these islands" refer to? Walled gardens that we create ourselves using the open source frameworks?
I think probably things along the lines of what's been called the 'cozy web': networks of smaller groups that don't publish to or expect responses from effectively the entire internet as a whole. (This kind of thing has always existed, it's basically the group chat with your friends but perhaps slightly bigger, but I think there's a bit of a trend of focusing on it more because of the feeling that the twitter/facebook attention and feed model is bad for your mental health. I've always felt the twitter model especially was pretty cursed so I'm glad there's some agreement building there).
It's still a work in progress of an idea.
I'm thinking more like mesh networks. I spoke of Reticulum elsewhere in this thread, but here I'm thinking I'd like the ability to easily join multiple TCP/IP networks (islands of connectivity) by social group (my friends) or by interest (pirate file-sharing group, my work intranet, a knitting community with their own IRC server, FTP, etc.).
Basically easy-to-use private & encrypted LAN overlays on top of the public internet. Each operator decides who to allow in or kick out of the network.
Wireguard solves the most of technical challenges, but it needs a frontend. The biggest concern probably is most software broadcasts their stuff across all interfaces, defeating the point of isolation between networks.
The purpose of The Inter-Network, or internet for short, was to connect together precisely these "multiple networks" that you refer to.
It's failed because of CGNAT, but come back because of IPv6.
And that will kill AI itself, since much of what it knows is from what learned from StackExchange before this latest one demise.
And before the obvious comments on how GenAI is creative, then please do this OpenAI and Anthropic, for your next LLM. Just teach it Python, C and Rust and give it some good books. But dont give it access to Github...lets see what you can do then...
LLM training will eventually transition from real data to synthetic data, same as alphago -> alphazero.
AI companies are also working to integrate training with real-world experience through sensors and robotics, to shrink the gap between human experience and hallucinated LLM experience.
They all have archives of pre-LLM content. There's also archive.org, google books, and pirate ebook archives. I don't know what they're doing to build video and audio archives, but judging from the cost of spinning rust, they're storing significant quantities of that, too.
Some parts of the internet are curated, and even with LLM influence they're still worth training on. I doubt wikipedia or stackexchange or rosettacode will ever cease to be useful at all.
I don't know how long you've been on the internet but the incentive to create new and original content was never that strong. Simple search terms return super-spammy websites (especially on mobile where ad-blocking is harder). SEO results for everything like simple search are awful, almost unusable. There hasn't been an incentive to create new original non-monetized content for the web for a while.
I trust LLMs more than search engines to discover my content and propagate it to users. They might "steal" something, sure, but I'm essentially invisible to the search engines as I could never hope to break into the top 10 links on a popular search term. LLMs can scan thousands of links and (for now) are more interested in quality rather than click monetization or referral incentives.
So, you know how they're built, but you're feeling the pressure of modern life and also they give you personal gain (supposedly) so you're fine with it, it sounds like to me? Use the same tools as your "competitors"/peers, even though?
Don't get me wrong, I too use LLMs for development and more, and I too know how they've been built, and I'm also a creative (music, 3D, VFX and animation) and for sure stuff I've published in the past, both code and otherwise, is now used to create new things for people and I get nothing, similar situation as countless of others. Yet I still use AI, so I'm not trying to create some "gotcha" moment against you here, I'm genuine curious about what you think about this sort of conflicting thinking, as I'm in the very same situation.
Doesn't seem like a particularly bad thing to me. Obviously for those who want to use the internet for commercial purposes it will be bad but for those of us who would love to see the internet go back to how it was before so much of it was changed in the aims of making money AI could push towards this. Great irony in the fact of course that the AI companies themselves are in the business of making as much money as possible.
It depends on filed. As documentary photographer it motivates me even more to capture authentic images of life around me. I don't care about remixing, because that is not what makes documentary photography valuable.
And how is any future viewer going to tell your images from "AI" fakes?
Because the photographer builds trust, over years, by not manipulating their images using AI. All trust is erased if anyone spots the manipulation.
With photos, I could see a cryptographic solution. Of course it would still need some kind of centralized trust, but it's doable if people cared enough. It could be applied by cameras themselves.
> It could be applied by cameras themselves.
And equally faked by bots.
Disclosure
Eh, part of what made early-internet so good was exactly that it did allow copying by users; the DRM era was later. But it's a very good example of why not to allow for profit copying, because that absolutely will crowd out the original. Piracy has to exist at the margin. The zero piracy world would also eat its memories because none would leak into archives. Remember Qubi? It wasn't even popular enough for people to pirate.
I think the lack of credit is even more egregious and a bigger problem than the commercial copying.
Yes. Some communities are weirdly against giving credit or keeping the credit (e.g. cropping off signatures from artwork), which I've never understood. It costs nothing.
> why we don’t allow stealing.
With the not-so-minor qualification that the biggest thieves have always gotten away scot-free. AI is just the international whole-internet version of this.
Behind every great fortune is a great crime
What is the great crime behind Norway's sovereign wealth fund?
stealing ??? more like piracy you mean
Funny, I was just thinking this morning that Google searches are absolutely horrible these days. It's like it has amnesia, a lot of recent history seems to be just gone. Especially on non US specific sites too.
The Internet has been shrinking massively. My earliest experiences with the Internet were discovering the world of hobby OS dev around the turn of the millennium, when I chanced upon someone’s personal website talking about their OS, with source code and screenshots and dedicated forum. My mind was blown. I spent two years finding hundreds of small websites dedicated to the topic, hung out on IRC communities with other teenage OS nerds like me, and of course participated in the nascent osdev.org forum. To note that all of those websites were readily found through Google, and interlinked with their own topic webrings.
Today everything has disappeared or has been conglomerated into siloes, sanitised, focusing on engagement. You have YouTube videos about it (which is more cheap entertainment than actual education), you get some posts here once in a while, there’s Reddit where all intelligent discussion goes to die. IRC is a wasteland of idle bouncers. Then the LLMs arrived to kill what is left.
Who says the Internet is a vibrant place today mistakes flashiness with depth. It’s all empty calories, just makes you hungry for more, never satisfies.
On a whim I watched my favorite childhood movie last week, Hackers. It's goofy in some ways, but man it captures the "wild west" feeling of early and mid 90's internet. It was just you, a slow connection to anywhere, and open ports all over the place. Right after I watched it, I dusted off an old hub, connected a few external usb-to-ethernet adapters to my work PC VM's, and now run them through an OpenBSD packet filter. For no reason at all other than to feel that again: me, watching packets, having total control. Hitting a wall and having to read a manpage.
I don't really have a point I guess, other than even after being steeped in a dead internet for years (with a slow decline spanning at least a decade arguably) I need to approach what I think is the internet in a completely different way. As in, not at all besides what is absolutely required for work. We're ants in a jar now, not cowboys like we used to be.
May I suggest you to look into mesh networks? I am a huge fan of Reticulum. Using it feels like being a pioneer, the scene is very welcoming.
The pitch: it is network-agnostic. The same mesh network runs on the Internet or through LoRa radios or any other physical layer than allows the exchange of data packets. It scales from private networks to global meshes. It's the wild west. People are excited, and eager to grow further.
https://reticulum.network/
Thank you! I'll definitely look into this.
That movie is ridiculous and still great. At least as a nostalgia pump.
And the red box references and pots patching and such were true enough, even if everything else got hilarious hollywood treatment.
If I recall correctly, the film actually had Emmanuel Goldstein of 2600 and Kevin Mitnick as consultants, so despite the hollywood treatment you have moments of accuracy like https://www.youtube.com/watch?v=4U9MI0u2VIE. Probably the only time Compilers: Principles, Techniques and Tools made it to the big screen. At the yearly 2600 conference, they used to always do a big group rollerblade through NYC.
Interesting, yeah I thought there must have been some consultation there. Especially because there's a couple instances where the hackers rely on social engineering over the phone first, and faking the "coin added" noise on the payphone. I guess a hollywood suit could have come up with that, but I doubt it
Yeah it's a whole lot better than the usual Hollywood depiction of hacking which brought us gems like creating a GUI in visual basic to trace and IP address. And more importantly it doesn't take itself too seriously which I think makes the ridiculous parts work just like much sci-fi takes liberties with the science part when required for the story.
Cereal Killers name was Emmanuel Goldstein, too.
That movie still gives me wardialing and 2600 meetups and “voice bridging from a Dennys payphone bank at silly hours” memories.
"Let's echo 23, see what's up"
I, too, look back fondly on the early web. Lately, though, I have been thinking that one reason for its decline wasn't just corporate interests like we often talk about here. A significant number of early bloggers were middle-aged and elderly people. After decades, they simply aged out. The younger generation that replaced them (albeit not so much among OS nerds like yourself) was less likely to use a real computer and keyboard as their interface to the internet, just a smartphone. Hence long-form text died.
While we're here exchanging old man stories, I remember the web before blogs existed! I guess we can organize the internet into these eras, each of which was in some sense harder and more expensive to find available info than in the previous:
1. Pre-web. Internet is mostly about messages sent to individuals or groups. USENET organizes group discussion into browseable topic-oriented hierarchies, IRC does the same but with lists in fragmented networks. If the discussion exists at all, finding it is easy.
2. Early web. Dominated by topic focused websites, early online shops and personal home pages. Search engines suck and face strong competition from manually maintained topic-oriented directories (did anyone else here contribute to DMoz?), content discovery is mutual and webmasters help each other out by joining "web rings". DoubleClick and AdSense start to funnel small amounts of money to creators, but it's enough to offset hosting costs and in many cases can make web hosting effectively free or even yield a small profit. This encourages an explosion of website creation. Discussion moves off USENET onto phpBB forums. Every organization decides it's a cultural imperative to have a presence on the information superhighway. Finding information is easy as long as you can figure out what topic it belongs to.
3. Blogging and centralization era. The internet starts to rebuild itself around people as the primary object, not the topic or category. Directories die because websites can no longer be categorized by content. Web rings die for the same reason. IRC is replaced by instant messengers that are about connecting people with pre-existing friends, not mutual interest groups. Outside of institutional websites that exist to promote the organization, things become hard to find without highly centralized search engines because nobody is putting any effort into organizing or indexing what they write anymore: maybe you get a few tags if you're lucky. Spam, hacking and lack of SSO causes forums to centralize onto Reddit. This is the peak of the search engine era because you are forced to use Google to find anything. The power eventually corrupts the tech firms and they begin political censorship to benefit the left in 2015 [1]. Enormous amounts of information is deliberately made unfindable as part of a large-scale programme of social control.
4. Social media era. All the same problems as blogging except now the bulk of the content goes behind login walls that stop search engines from surfacing them. Video and podcasts start to matter more, both of which are unsearchable by default. Eventually video completely dominates, as few younger people want to read when they could watch instead. Firefox starts to replace IE6, and then Chrome. They bring ad blockers in their wake which starts to choke off ad revenues, so many websites from the web's first era go unmaintained and eventually offline. This is somewhat but not entirely compensated by the falling cost of web hosting. Social media remains because it puts people's faces next to everything, allowing clout farming and viral notoriety that can sometimes be monetized by becoming an influencer. The only part of the web's first era that really survives into this era is Wikipedia and Reddit, which by this time substitute monetary rewards for power tripping by a small group of ideologically driven moderators.
5. AI era. Information is so heavily scattered over so many tiny sourcelets and search engines have become sufficiently useless that full neural integration of knowledge is required, with LLMs issuing massively parallel and complex search engine queries as a backstop.
What can we predict for the AI era? Institutional websites will remain because institutions still have an interest in getting their agenda into LLMs, but visual redesign efforts will largely cease as traffic stats seen by executives show visits completely dominated by AI. There will be lots of conversations of the form, "why redesign our website to look more modern when 99% of traffic is AI which won't care?" Blogs will go the same way as the thematic websites they killed, disappearing as the authors age out. A lot of effort will be put into finding ways to block AI crawlers to create 'human only' spaces, especially by social media firms, but these will fail because AI will just be integrated directly into browsers and become unblockable - and anyway, the incentives to create will be ignored. ChatGPT style text oriented interfaces will last until inferencing capacity catches up, being eventually replaced by voice interaction and on the fly video generation for nearly all users.
Where we go from here is hard to say. Content creation was most pure in the web's first era, where people with knowledge were incentivized to share it with the world by the promise of a bit of fame combined with ad clicks to offset hosting costs. Ad blockers, social media and AI killed that world. You could however bring it back by producing a new platform that isn't like the web, one where AI and search engines are blocked via technological means (e.g. confidential computing). How much anyone would actually enjoy such a web is unclear.
[1] https://arctotherium.substack.com/p/the-closure-of-the-inter...
Most hobby communities in my spheres of interest have been subsumed into Discord. It is yet another silo, but it does feel lively. I don't think I'm contradicting you here, I just feel less negatively about it.
Well, just as the "internet" supplanted newspapers, magazines (gosh those classic gaming mags), and broadcast television for many people,
and how the newspapers replaced the town criers before them,
why shouldn't the "internet" be supplanted by a more accessible medium?
Why should I have to suffer through Fandom raping me with screen-obscuring banners and "PLEASE ALLOW ADS" just to make some sense of fucking Warhammer 40K lore (written by unpaid volunteers anyway)? instead of just asking ChatGPT what the fuck Globriznaroks is/are.
Why should we support shady companies by sitting through their ads on YouTube videos for minute topics instead of just asking AI for the shit I want to know about?
Why should we submit to the whims of 3 mods on a subreddit deciding what thousands should get to see (fuck /r/AskScience) and then getting low-effort answers or outright trolling anyway? instead of just asking AI?
Bury me, I am ready.
This is better for you as an individual, but worse for society as a whole. As you've noticed, monetization is the weak spot. AI will not escape being ruined by monetization, but it might be harder to notice when it arrives.
Where is the information going to come from though? LLMs aren't primary sources, they consume and regurgitate.
I guess/hope that's where Yann LeCun and their proposed "world models" etc are going to fill in :)
Checkout Kagi SmallWeb !
I love Kagi, I am a paid user since day 1, but SmallWeb is a collection of English-speaking tech blogs, it cannot even begin to compete in diversity with the GeoCities era of the web. (to be fair: I've had Kagi employees telling me their dataset is growing larger and more diverse every day, so worth keeping an eye on it)
Sometimes I use Marginalia's "Vintage Web" search for niche topics; most results are dead blogs and old .edu personal websites that someone forgot to delete, still a vanishing minority of anything one could find in 2001.
That's 100% correct. Since Google's "helpful content update" Webmasters get massive amounts of "Crawled, not indexed" reports for anything Google considers "more of the same" or "thin content". If you're not an authority on a subject, simply meaning: you already rank for similar content, or if you don't get links from more popular domains, your content is in the abyss.
It's all under the guise of "We're fighting SPAM", but the algorithm (or model) they use is heavily skewed towards intents (actions) and brands (because they 'trust' big names).
And it's not working.
A simple, short informative blog about a tool you used that could be of interest to max. 100 people on this planet is no longer getting ranked, if it gets indexed at all.
Those posts tick all boxes: no incoming links, no authority, thin content.
It changes somewhat between "Google Updates", but it's pretty clear that it's no longer working.
Multiply the 100 people not finding that post by millions of queries and it's now a big problem for Google.
My dad asked me the other day if Google got worse because smaller companies are not paying Google enough money.
He didn't see any difference between ads and content, because all results are now big brands only.
"Helpful content" is such a misnomer. It removes all helpful content in favour of AI overviews and only shows intent-driven, commercial content.
We're watching the end of Google's hegemony for sure.
This is fascinating. I haven't really kept up with SEO for years, but was recently helping an academic publisher setup their web presence and ran into exactly this issue: Google was/is refusing to index the actual journal articles on the new site and they were moved into this liminal "Crawled, not indexed" state before vanishing entirely. Comically, it's now easier to get your content indexed and served with a visible link by ChatGPT than it is Google...
> We're watching the end of Google's hegemony for sure.
Who is standing by to replace them though? OpenAI and Anthropic certainly not, they are burning money in a fire pit to stay alive. There is no way in hell they can afford the compute necessary to replace Google.
If ChatGPT goes bust, they'll find another ChatGPT, not another Google. I wouldn't be surprised if we'll see a popular competitor from China in the years to come. TikTok already took social media by storm, something nobody thought that was possible either.
About 15 years ago, I won a phone in a contest. Last week I tried to find information about it, but I couldn't. No AI, nor Google could find anything the contest I won it in. When people say "the internet is forever" that can certainly be true, but it isn't for everything.
That reminds me - my wife had a friend who was killed by a shark (no joke). There were newspaper articles, but Google has completely forgotten about this and the newspapers are often behind a paywall, changed their URL's or simply the article vanished from those sites as well, which certainly doesn't help.
I wonder how libraries, who have traditionally been the ones to archive the news, have kept up with everything moving online and now being subscription-holed. Hopefully it's not just the Internet Archive doing this, which has its own problems.
Kagi
Kagi is not immune to this (it's the substrate itself that's losing quality, not just search) but it's a decent antidote.
Kagi can’t solve the fact that the internet itself has gone to shit. The independent forums are largely gone, the mainstream social media’s have all locked down and been spammed with AI slop.
A lot of the time I use Kagi to find the Reddit post or Wikipedia page I need. There isn't much else.
They only repackage what Google et all returns.
Or have they started own indexing?
They're so horrible that I've started defaulting to their AI summaries. And I hate those summaries. It's just that the regular results are so terrible now, and seemingly getting worse at a noticeable pace.
I used to not worry. I was sure that a competitor would come along and fix search. But the longer that's not happening, the more nervous I'm getting that we'll actually lose search. If a few more years pass in the current state, I'm afraid the majority of people will forget what search was like and default to AI summaries.
I've tried alternatives, including Kagi (not actually relevant because there's no way I'm – directly or indirectly – buying Russian products) and Uruky, but they're not good enough.
(Edit: Added "directly or indirectly" about Kagi to point out that I'm not claiming that Kagi itself is Russian.)
Unsurprisingly all search engines seem to be struggling with AI content sites as well. It's rare to get human articles, sometimes rare to even get authorative websites. It's frequently a bot site with a plausible enough name like, potterspainterly.com or medhealthdirect or something with oddly specific articles written in the last year.
This is the root problem. The idea that Google is deliberately sabotaging search seems far less likely than the idea that the internet is mostly garbage and SEO slop.
perhaps this would not be a problem if google did not directly monetize AI Slop... for that matter why even include AI crap in search.
The Google conspiracy has been getting shared for long before AI slop came in to the picture.
The fact is that SEO people got too good at their jobs and filled the search results with junk.
Google used to penalize websites for SEO hacks. Then they started attending SEO conferences themselves. SEO specialists didn't outsmart Google, Google stopped trying. As the sibling implies this is most likely because they noticed that the easiest way for SEO spam to monetize is ... Google ads.
Google can properly classify email spam.
They should be able to fight this as well.
Problem is they are the ones funding the poor quality spam and they're in turn profiting by taking money from advertisers.
Probably what is happening here is that in the race for AI, which, whether we like it or not, means power and control, Google crafted things in a way that it does not do "self-competition" by their old search engine. Idk, just throwing ideas aloud here.
Kagi is not russian. The founder is from Serbia and the company is american.
PS: I hope not many people are the type to see a slavic name and conclude Russia.
Kagi pays Yandex for data. Yandex is definitely Russian.
Yandex also moved outside of Russia and there only one vendor.
Brave search works really well, I haven't switched to Google search for months.
If I recall correctly Brave scrapes the web via their users, cannot be individually disallowed in robots.txt and Brandon Eich is conversing in a pretty hostile manner in every thread about him or his company.
I'm not in the brave ecosystem otherwise nor a big fan. But the search engine was competitive with google when it was still cliqz, before it was shut down there and the leftovers bought by brave. And it still works really well.
Even if brave were problematic it would be the lesser evil to me.
Thank you for your suggestion. I think I've discounted Brave automatically because my brain is numb to the dime-a-dozen chromium browsers out there. I'll definitely give the search a try!
How is Kagi indirectly Russian?
They are repackaging bunch of search engine results, including Yandex (and paying for it).
Yandex is Russian.
I would not describe it as Kagi being indirectly Russian.
That's certainly the goal, to drive more users to AI by making the other option worse. The real question is when do you put a stop to it. When brain implants are 2x as productive are you going to say nah? It's times like these were you are supposed to take a step back and consider what you are actually producing and why.
> there's no way I'm – directly or indirectly – buying Russian products
Why not?
In my case: Russian state is my direct enemy and most likely to invade my country. And would do it if they would consider success likely.
Previous wars with Russia were obnoxious with very bad consequences, so I dislike idea of even very indirectly funding them.
And I support actions that are harmful to Russian economy, also when they are harmful to me - as long as it is not too badly balanced. As this is much cheaper than directly participating in war.
(I am from Poland)
PS
Yes, I understand that at some point there are some indirect effects that you cannot avoid.
I also understand if for some people paying Kagi that pays tiny fraction of that to Yandex that is paying taxes in Russia which funds their wars is too tenuous connection to care.
After publishers successfully sued the Internet Archive over its digital lending program, calling it unauthorized copying
No. The court specifically determined that the Internet Archive was guilty of unauthorized copying. It was not simply an unfounded or unproven allegation. The Authors Guild, the National Writers Union, the European Writers Council, and the Society of Authors in the UK all came out against the Internet Archive, and supported the suit.
Each new restriction limits the archive’s ability to act as a comprehensive backstop.
This self-inflicted damage to the wayback machine is the real tragedy of this entire affair. When IA was asked to stop CDL - many times - founder Brewster Kahle continued. The National Writers Union tried to open a dialogue as early as 2010 but was ignored:
The Internet Archive says it would rather talk with writers individually than talk to the NWU or other writers’ organizations. But requests by NWU members to talk to or meet with the Internet Archive have been ignored or rebuffed.
https://nwu.org/nwu-denounces-cdl/
When the requests to abandon CDL turned into demands, Kahle dug in his heels. When the inevitable lawsuits followed, and IA lost, he insisted that he was still in the right and plowed ahead with appeals. And here we are today.
Seems inevitable, doesn’t it? Expecting otherwise would have been hoping that notorious atheist Richard Dawkins somehow spared one specific god. Making websites accessible with history ignoring copyright is sort of what it does. That he would do it with books seems entirely in keeping with the philosophy.
I was initially confused what Dawkins was doing with books, until I realized that the "he" in your last sentence was Kahle, not Dawkins. Might want to edit your comment to put his name in, because otherwise you have a pronoun referring to a person named in a different comment (rather than the person named in your comment), which could get quite confusing if more people comment on the parent and their comments push yours down the page.
I spent the last three days (off and on) using Gemini to configure my edge router 4 with my iOS devices on a vpn and it's been awesome. In the past I'd do a google search and read a few sources of documentation, do another google search and read another set of documentation. Now, Gemini aggregates multiple pages together so all of the work of reading source docs from multiple locations is now n a single step.
Oh, I should mention though. There was no advertising at all. They didn't make any money off me. It was 100% Gemini which I recognize as not long-term feasible.
All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else.
If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse.
Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly.
> Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human.
Yep. IMO, this is so far the biggest AI-inflicted damage to the web. A bit of anecdata - wikipedia (and all other wikimedia sites) are blocking my Firefox since about a week, with a "please respect our bot policy" message. Outright block, not even a captcha.
It took me a while to figure out they don't like me disabling some SSL ciphers, so now "JA4 browser fingerprint" is not matching user-agent. Funnily enough curl (what I would imagine a bot would use) pulls exact same URLs from exact same client IP, just fine.
Sure - it sucks, unfortunately the alternative is the sites going away entirely. When the load from scraper bots is constantly knocking the site offline the choices are literally to allow it to remain inaccessible for much of the time, put up a layer of defenses with all the user-annoyance compromises that entails, or just give up and unpublish the site.
The alternative is simple.. Go dark. VPN tech is known from like 30 years. Pretty much everyone can use it (VPN providers). But instead using it to browse net, build VPN overlay networks of interest for people. Gaming networks, R&D networks, Retro Networks. People will peer to PoP and use resources. Bad actor? BAN it from network. You have control. This could be done in Internet, but big corpos and big money won the battle. Just wake F*ing up...
Continuing on your suggestion.
There could be open source tooling to create custom private "closednets", with
- trust ring mechanism to allow invitations, flagging, banning, and banning those that invite people who were banned
- the rules of the closednet
- search engine with opt-in scraping
- portal (remember the 80s?) with all the registered nodes, perhaps by service category such as public git repo hosts, web sites etc.
etc.
The first closednet could be Hacker News.
Cloudflare specifically has a block for LLM and AI training bots now.
Not sure of the effectiveness but it's there.
Minimal. I'm behind Cloudflare and 90% of the traffic is still scrapers. I don't think they're serious about the long tail.
I think the main thing Cloudflare is trying to do is block direct traffic from frontier labs and then start charging them for access. They might end up shooting themselves in the foot, as this simply empowers sketchy residential-proxy outfits to undercut Cloudflare and sell the data to labs for less.
It still only blocks "well-behaved" bots that have proper User-Agents and respect robots.txt, so it's largely pointless.
The problematic bots are all disguising themselves as Chrome and sending requests from millions of residential proxy IPs, and the only real solution to those is some sort of captcha or PoW page on first visit.
Me, I'm just scraping the parts of the internet I like, toying with local LLMs… ready really to just shove off.
> "humans would visit the website and the creator would get some reward"
That expectation is a problem, has always been a problem, and Tim Berners Lee never mentioned anything about a reward structure when coming up with the WWW.
Your thinking too narrowly about the reward. Sometimes, it's just about the getting the knowledge out there that's motivating the creator, not anything tangible for themselves.
> it's just about the getting the knowledge out there that's motivating the creator
In that case the creator should welcome AIs with open arms; a human reader will forget eventually, but the AI will preserve the knowledge forever.
"the AI will preserve the knowledge forever"
no, only some mangled form of it
> Your thinking
Should be "You're thinking".
> Should be "You're thinking".
This should be "This should be 'You're thinking'." don't you think? Why bother correcting someone's grammar with a sentence fragment? You're just trading one mistake for another. I'm hoping someone finds a grammar error in my post, because continuing this would be hilarious.
> This should be "This should be 'You're thinking'." don't you think?
Reflexively, I think it should be more like ...
... but then that's just me, in [my] quirks mode.Did the operators of HN mention anything about a reward structure for posting comments here? I'm sure you can see how that's still attractive to some.
> If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new?
I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous, are you?), and your idea might get embedded into a set of model weights somewhere. No human will see a byte of it.
Are you actually saying you'd be OK with that?
Yes, that would be fine. I write primarily communicate ideas, not for credit or fame.
Empirically, however, LLMs don't strip out the author: the big models know a lot about what I've written even with search disabled. Ex: https://claude.ai/share/8cbcdf88-a360-421a-8c06-ae7b7992e866
Sure they know about you if you ask, but generally they won't credit you if they cite an idea from their latent space that came from you.
Not the OP, but I suspect no one goes to my website anyway. (I write nonetheless.)
I do! You invented the crayon picker in MacOS..?
I mainly shared my projects for learning, discussion and bragging rights.
LLMs just use everything, generate similar code with no attribution and keep users from visiting, so no bragging rights or attention.
Worse, there are some PRs that seem fully generated ...
So i mostly stopped sharing and started pulling my old repos offline.
At this pace, i don't want to compete with a clone of myself in the future that will do my work for much cheaper.
Your level of thinking is defined by ICD-10.
Yeah, authors don't want to be recognized as authors, they don't want any reward for their work, they don't want to amass pool of loyal readers, interact with them, etc.
All they want is for halucinating AI to take excerpts of their work and compile it with random sh!t.
GENIUS
Thanatic drive masquerading as transhumanist virtue signalling.
There are different types of writing. If we depend on people writing because it's enjoyable at some level, we're going to lose writing that's important but also a bit tedious.
Of course! I do think we'd lose a lot of great writing if it went amateur-only. But my parent seemed to be saying the incentive would entirely disappear, so I wanted to give my perspective.
> I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me.
Sure. But you can see that for some people (myself included), writing for peers is part of the joy? And that if instead a megacorp places an opaque computer program between the author and the readers, that joy might be ruined?
Easy to fix a well documented router now. Difficult to fix a non documented router in five years time because no one has been contributing to the web about its bug fixes.
Gemini can just consume the device documents. There's an incentive for device makers to publish this content.
There had always been some incentive for manufactures to publish device documentation, and yet it has often been quite lacking either in quality or overall existence. I doubt LLM/agents being the readers will change that at all. What I expect AI scraping and using without credit will impact is people publishing their own unofficial help and guidance, and the affect there is likely to be negative. It won't stop all of them, but enough to be noticeable. Another possible negative is the manufactures documentation being AI generated without sufficient review, so possibly more erroneous than before, or intentionally not producing full documentation at all and expecting AI to fill the gap (MS seems to be heading this way: pushing "ask copilot" all over Azure instead of links direct to good reference material). All this would add up to a situation that is somewhere between "a little worse than pre-AI" and "an absolute shit show".
At least LLMs almost always transform the original - it usually isn’t as straightforward as “Here’s the original but without the ads that pay for it”.
But we already have the latter case that exists - ad blockers. Ad blockers literally serve up the word-for-word original content minus the ads.
Well in the example above the manufacturer still has incentive to provide the manual's and guides that describe how to use their products, and if that is subsequently served by an LLM that's totally fine. The only sites that LLM's would have a negative effect on are those that are only hosting content for the ad views.
Manuals don't always well explain how to use their products with everybody else's products because there are too many to do that. But there are lots of people trying things out and might figure out the fine details on how to make various things work. They then publish these how-to pieces (which exist no where else) to the internet, or at least they used to when there were incentives to do so.
Exactly. What’s problematic about comments like your parent is the absence of mid-to-long term thinking.
It’s like bragging about a new highly addictive psychedelic drug that a dealer gave you a taste of for free. The effects are awesome today, you feel so fun and free! Never mind that it’s destroying your body and that the dealer will eventually charge you or demand you pay in other ways, that’s a problem for another day. Weeee!
Funny, I published information in the hopes that humans would benefit from it. If it happens to be through collective intelligence of LLMs I'm ok with that--even more so if through open models.
> Funny, I published information in the hopes that humans would benefit from it.
Sure, humans would benefit.
It took them searching, reading themselves, maybe even understanding something in the process, to complete a 360° revolution of their squirrel cages in time T.
Now they can omit searching, skip reading to the regurgitated answer, throw away understanding, and complete a full revolution in T/N, where N is a heuristic value directly proportional to the amount of skin in the AI hype.
But the catch is that the squirrel cage must run non-stop still.
I wonder the same thing. I only imagine that what comes next is worse: AI companies using vast resources to develop new training data, in house, locked down. They are already doing this with developers and code at Meta. Information will become locked away behind AI paywalls and chatbots.
People write and create regardless of profit motive, it has been that way for thousands of years.
People are way less likely to write if there is no one to read it. And blog were also monkey see monley do - people seen other peoples blogs and got inspired.
When people wont see others blogs, they wont start writing own. When there will bw no ome to actually read it, they will go to do something else.
Actually no. The original copyright laws were created in part because of the realities of needed profit motive to have high value writing done, time consuming compilation work done. It was even titled "An Act for the Encouragement of Learning". The thousands of years writing you are talking about was often funded by patrons, who kept the output in their private libraries to show off (and maybe lend out) for prestige. It was a horrible limitation of knowledge and ideas. Much worse than the profit motive, copyright based system that came after that spawned a new age of knowledge in which everyone had cheap access, and those that didn't had access to the (no longer just private) libraries.
I'm sure the billionaire class would love a return to patronage based libraries, NDAs on authors of books, and the elitism they would feel with a return to private libraries locking away all kinds of knowledge that would happen if patronage become the only way authors could make money (such as with AI just regurgitating their works, or if the stupid 'do away with copyright' people got their way).
Actually it grew out of censorship and monopolies: https://en.wikipedia.org/wiki/Statute_of_Anne#Background
Cheap access came from the invention of cheap printing . The laws were passed to restrict it.
It's not copyright that caused cheap access. The printing machine allowed for cheaper publications, that's what spawned the new age of knowledge. The raw materials and the duplication of knowledge was the bottleneck. With digital systems this cost is minuscule, but still there.
That is next earnings quarters problem is the approach being taken
> incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content?
obviously new content still has value because it remains the source layer for LLM agents. it just wont be ads giving you revenues thats all.
Well companies are starting to put hidden ads in text content if the user agent is an AI crawler
That works now because there are human made sources that the AI can find and summarize for you. But now there are no incentives at all for humans to write anything on the internet and if they do the content will be buried by hallucinated content someone else posted at a larger scale.
I had the similar experience to yours yesterday and it lead nowhere. Funnily enough I was also trying to configure a vpn on a router, google didn't return anything useful (besides a blog post clearly written by AI and with absolutely no information in it). Claude managed to give some interesting pointers, but its suggestions were not working and I also noticed that it started to hallucinate badly about ipv6 and gave me some suggestions that were just plain untrue. Claude Opus is smart, usually when it gets so convinced about something is after researching the internet and not just based on its training data. I wonder where it got so convinced about it. Maybe reading some other hallucinated blog post like the one I stumbled upon?
That only works when there's an abundance of documentation for your specific device. The moment you're on a more recent version of something and have a weird issue, all hope is lots. You get stuck in loops because all LLMs keep recycling old advice that no longer applies. This problem will only get worse and worse as people no longer as questions on public forums, so answers are not publicly available either.
Yes. When my kids wonder why I don't hate AI the way they do I tell them it's because I hate Google even more.
To be more precise, I hate the SEO shithole the internet has become, that Google serves up, that Google facilitated, indirectly created.
(I really don't have any tears to shed if there is a death of the Corporate Internet™.)
AI seems to be on the same trajectory? Search was very useful in the start also, until it became entrenched. Then search placement became a target, and they are just focusing on extracting rents. All way paying the content providers zero or near-zero. And with years of that dynamic, we end up where we are now. It was the same with "social media". The same will happen with AI. AI is a power for more enshittification - being currently less shit than Google is (mosy likely) temporary.
It's definitely temporary. Remember there are no ad blockers for LLMs.
You're right, this is the best it will ever be. But there likely will be ad blockers for them -- local LLMs that filter for any brand placement, etc.
They're probably going to figure out how to aggressively monetize and enshittify LLMs at some point.
We're likely at the "golden age" of LLM-assisted web searching and summarization.
Hopefully open models keep it cracked open, but expecting enshittification is always the safe bet these days.
For the free to use LLMs, for sure. But if I pay $20 monthly, why should they enshittify it?
$20 a month is not nearly enough. The current finances require AI companies to make far more money than that to not implode.
Because that’s what always happens? Because their raison d’etre is to squeeze you dry?
look at every paid streaming service...
The amount of money you pay is irrelevant if you don't leave the platform once they starting adding revenue streams that degrade your experience.
Is there some search engine that you think could have become popular and not ended up with SEO optimization?
One you pay for yourself !
Very happy with Kagi personally
>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
> One you pay for yourself !
SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I don’t agree with the parent that the solution for SEO spam is paying for the search engine (I think there are other good reasons to do this though), especially since people have been doing things like naming their company “AAA Auto Repair” to be first in the phone book since before computers ever existed. But the person you originally replied to does have a point in blaming Google for the problem. Most SEO spam sites make their money from ads, and Google are the ones who run the ad network, which means Google are the ones funding them and creating an incentive for them to exist.
> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
They do this because they benefit from their site being visited or the information they are providing being noticed.
> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
I see no reason it would have that effect. It does, however, create different incentives for the search provider to improve the signals indicating page relevance since the user is the priority instead of advertisers.
Ah, the argument is that Google intentionally avoids showing you the most relevant results, or at least avoids "solving" the problem of webmasters who attempt to 'game' the system.
>>>> Is there some search engine that you think could have become popular and not ended up with SEO optimization?
>>> One you pay for yourself !
>> SEO is the practice done by webmasters of optimizing a website to improve its visibility and ranking in search engine results.
> They do this because they benefit from their site being visited or the information they are providing being noticed.
>> If you are paying to use your search engine, does that mean webmasters are no longer incentivized to/will not try to improve their visibility/ranking in your search results?
> I see no reason it would have that effect.
Agreed.
Google did not really lost to optimization, they enshittified to make you spwnd more time on search (and see more ads)
Ah, the argument is that Google intentionally avoids showing you the best results, so that you spend more time searching (and being exposed to ads)?
I'd frame it more that Google is incentivised to show you the results which are most lucrative for them to display, rather than the ones which are most beneficial for you to see.
Incentives. Which is why platforms (such as search) should never be allowed to be in the same company with things built on top of them (such as ads), if the combined company is significant for an important market.
We used to know better, Standard Oil vertical integration was dismantled.
While I understand that you feel good as you got the device configured, you would have been better off without gemini. The 'old' way as you call it would have resulted in you knowing the backgrounds and the inner workings of your router, making maintaining it a breeze and helps you actually understand your setup. Besides that, it would probably trigger you to rethink some of the things you now blindly have implemented because gemini did not show alternatives nor reasoning behind it. (something that most definitely would have been documented on the source pages)
Yep - it's awesome, the problem is that Google isn't sharing the revenue with the context creators anymore - over time, unless fixed, this will decimate the knowledge base it feeds on.
> Oh, I should mention though. There was no advertising at all. They didn't make any money off me
Are you sure about that? Even if you didn't see ads ( remember people pay even if you don't click - just like a billboard ) - they are still profiling you to better sell you ads in the future, and using your interaction as free training data.
I've run into major problems with LLMs as I maintain my home Linux systems. If I just copied commands they list, I would have a near 100% failure rate, as most of the information they have has been gleaned from forum posts that are years out of date.
They're a good jumping off point, but I need to delve into the original sources just like I did when I used Google.
I'm not sure. I run a homelab tailscale/k3s setup with gitops, dozens of services, VMs, backups, etc, all vibed by claude, and works just fine. Didn't write a single line of code for this. I don't know kubernetes and never will.
> I don't know kubernetes and never will.
He said, proud of his own ignorance.
Proud? No, but I am not ashamed of it also.
I would be the first to admit my ignorance on the absolute majority of topics. There is a limited number of things I can learn in life, and kubernetes won't be one of them - I'm just not interested in it (and all the other infra stuff, to be honest), as long as it works.
I've had the opposite experience giving claude code SSH/ADB access to my devices. Not sure if it's the model itself or the harness, but I haven't had to manually do a sysadmin task in months.
What on earth LLM are you using and do you tell it which distro/version of Linux they are supposed to be working with + give them access to web search or man pages? I haven't had this issue in like a year and a half.
By default I use Leo, and I do specify the distro. I tried to qualify by version, but when I do that I lose out on a lot of correct information for things that haven't changed in awhile.
Have a look at https://github.com/ThorOdinson246/whatisit-nl2sh
It seems to get shell scripts right most of the time.
its pretty good with logs too. i mostly paste a few pages i suspect hace a problem in them and let it go to town. its almost always correct and is way faster than me
I've been mostly enjoying Gemini, but it also clearly and definitively told me something i was trying to do was not possible with the library im using, so i wrote a different implementation, an hour later to discover that the library does in fact do precisely what i wanted in exactly the way i wanted with less headache. If I'd just gone right to the documentation instead, it actually would have saved me time.
There will be new ways and incentives for content creators to be compensated. Many AI search startups are already talking about this or have created programs that help incentivize content creation.
I hate to break it to you, but we are at the start of enshittification circle. Once Gemini has monopoly you will reminisce with joy google search results.
And I've spent the last few days irritated that everything I ask Gemini is answered with something that's blatantly wrong and I'm not even a subject matter expert. A quick Google search for the same questions gives plenty of results that counter what the LLM gave me.
YMMV.
Confidently wrong summaries are the bane of Google search, and unfortunately I'm finding the AI seems to bleed into the actual search results too now, often turning up pages that back up what the summary is (wrongly) suggesting instead of surfacing actually relevant results for what I'm asking.
Believe me they're making money off you one way or the other
You deserve what's coming for you.
I built a simple Claude Code container on my homelab to do the same. I can just SSH in and get tech support when I need it.
AI is on the whole bad but it’s impossible to forget that the ostensibly human-curated Web as presented by modern search engines is bad too. Invasive ads, popups, and worst of all the substantive content has a nine inch frame of SEO filler and a three inch picture (what you were after). And some things are not even high-tech slop stolen. It is just old-school verbatim copied from another website and repackaged with another frame.
But yeah, text remix machines are not a long-term solution to that problem.
Imagine someone making an indie router and sell it too you then the AI not be able to surface their doc
no, no, no. you are supposed to romanticize the hunt for correct information /s
Even with the /s I think you misunderstand. If you just want to configure your router, then AI is neat. If you want to learn and understand what's going on in the router as you set it up, then being given the answers basically teaches you nothing.
It's the same reason why we don't give students the answers to things, we teach them to find the answers.
Are you sure the lack of advertising isn't long term feasible? It seems to me that AI models have proven themselves to be something consumers ARE willing to pay a subscription - or even pay per use/token for.
Are they profiting from these subscriptions yet?
I've been using Gemini free tier exclusively for all my AI "needs" and would not be willing to pay for it
It took you 4 days because you used Gemini. Gemini is the worst AI model I ever used. It is way behind even open models. It looks like Google just reached its AOL moment.
Gemini isn't great as a model, googles search and ability to cite textbooks down to the paragraph make it better than every other model for human in the loop tasks.
I end up using the gemini api for with search enabled for the cases that I don't have access to good grounding data even in agentic tasks.
Come now, copilot is worse in every way
Guessing this is the relevant part: "worst AI model I ever used"
Copilot can use any model like Sol or Opus and is just a harness so not sure why people say this.
Personally it's because i used it when it was brand new and work paid for it and i had no idea what model it was using, my boss just turned it on for automatic PR summaries and code reviews and it was universally dogshit, and the auto complete in my IDE was awful as well.
> While the web has always been organized around intermediaries that shape what survives online and who sees it,
This statement, from the sixth paragraph of the article, is something that I would have liked to see addressed more in the article. The article implies that this is something that must always be true, or cannot be changed, and simply focuses on how we could have better/better funded/better protected intermediaries (AKA gatekeepers), and doesn't discuss the possibility of an internet (or part of the internet) without gatekeepers (and doesn't ask if it has ever existed/does exist/should exist)
I occasionally use Google Search when DuckDuckGo fails to give me relevant. Almost always, Google has better results.
Though I can find its AI answers annoying aggressive. I'll look up like two search terms and the AI will bullshit multiple paragraphs out of despite having zero context of what I am looking for.
DuckDuckGo seems to have detection of whether it should give an AI answer. And it allows you to have more granular control of when you want to get an AI answer. And is overall less distracting than Google's.
Interesting, I've seen much better results on DDG. Most recently was the search: `site:feeds.bbci.co.uk inurl:rss.xml` which works on DDG but gives zero results on Google. As far as I can tell, Google just decided not to index these.
Yeah I mostly still use Google out of habit but there have been a few times where Google has decided something isn't worth indexing (too niche, doesn't use SSL).
I miss when Google was like a grep for the entire visible Internet. Now it tries to second-guess my search and direct me to a bunch of sites which all have identical information that isn't what I'm looking for.
I find DDG is struggling to, or chosen not to, filter or derank obvious AI generated content farm sites. Of which there are an insane amount of already.
We thought blogspam was bad, at least it was easy to ignore. It's hard to find authoritative sources for a number of topics, worryingly health advice is one of them.
The problem and what the article is pointing out is that original content is slowly and progressively being replaced by AI content. And since AI gets trained on this content as well it will eventually train itself on previous gen. content that was also AI generated. It is slowly eating the web. Eventually you won't even be able to evade it because it'll be everywhere. Before AI became this expressive I could at least expect someone writing articles, FAQs, blog posts to have some backbone. Now I frequently run into content that obviously was never even checked by a human.
https://noai.duckduckgo.com is a thing fyi
i don’t agree with the google has better results thing. sometimes it does. most of the time it’s just that google has the site i want higher in the ordering than DDG. personally i’m fine scrolling down a little bit more. it’s rare i need to go to google for something that DDG doesn’t have at all in their results, but it does happen.
i do have to go to google for maps/directions/planning travel. a lot that’s annoying.
> https://noai.duckduckgo.com is a thing fyi
or you can press the gear button -> "Ai features: Manage" -> Search assist
That was true until a few months ago. Now, it almost never has relevant results, and I've given up on it.
It seemed like Google was good ten+ years ago and then gradually trended towards being rotten around 2020-2022. In those days blogspam was king. I want a cooking recipe and it would show me a life story. Or I wanted a OG web game but a link farm would come up top. I think there are leaked internal comms where they discuss nerfing search results to pump engagement and ads impressions. No doubt it would boost ad bids too if businesses couldn't be found organically.
Around that time Bing and DDG were actually better. Then LLM's came along and they started to take things seriously again. Maybe they think the OpenAI threat has abated enough to begin enshitification cycle 2.0.
I know Google started deteriorating. It was in 2008 or 2009. Up until then Google would return 0 results if it couldn't find a document with all the words you specified.
Then it started serving synonyms, attempted to correct spelling, and so forth. Instead of serving up that there was 0 results, it attempted to be "helpful".
For a while, you could enable "verbatim" search, but even that has gotten corrupted.
Their search quality has deteriorated ever since that change. It is sad.
that control (so I can turn it off completely) is why I picked DDG to replace Google, who force feeds us the hallucinations
I've stopped using DDG now because of result quality. I now use a "meta" search backed by EXA, Tavily, and SearXNG in parallel. It can be agentically de-dupped or summarized as needed. Search as we knew it is done, largely because clicking through to evaluate result relevance before diving deeper sucks. Now we have agents that can do that portion and perform multiple searches, building on information in the last batch, to collect good results
Yeah, there are times where DDG has like, literally three results. Yet, I know for an absolute fact, there are hundreds of pages on the web that contain the terms I specified. Web search is becoming utter garbage.
Try brave search.
The article touches on something that I've been thinking about with regards to Google's AI strategy; the automatically-generated AI search summaries are not great. They very frequently confidently misinterpret what the user is searching for and generate half a page of useless information that pushes actual results down the page, and they are occasionally hilariously incorrect, with hallucinated facts.
This is probably a difficult-to-solve problem; given that they generate billions of these a day, not even Google can afford to devote enough compute to each query to reliably generate quality results. You can see this by selecting the "AI mode" from the search interface after getting the mediocre summary - the results are much better and generally perfectly usable. Though even that is probably a special minimal-compute version of the lowest tier of Gemini, it's still maybe an order of magnitude more capable than whatever generates the search summaries.
The bigger problem is that these search summaries are the default and by far the most common interaction that the general public has with "AI", and because this experience sucks, they just assume that all LLMs are similarly stupid and mostly useless. In non-technical spaces I frequently see the argument that "AI" is not useful for anything, all it generates is garbage hallucinations, and almost invariably they cite some actual terrible experience with the Google AI search summary. I would argue that the strategy of adding LLM summaries to every search is the worst of both worlds - it makes classic search worse while poisoning users against the idea of actual LLM-assisted search.
> not even Google can afford to devote enough compute to each query to reliably generate quality results
https://www.dw.com/en/german-court-holds-google-liable-for-f...
I feel like collecting, curating, and protecting high quality corpuses of "truth" is going to become increasingly important for high quality AI.
There will come a day (and probably soon) when "training on the public internet" (Reddit, etc) will taint your model with metric tons of corporate contamination, political poison, and other adversarial content intentionally crafted to bias AIs for various reasons (corporate gain, geopolitical information warfare, etc). Basically the AI-equivalent of SEO.
All of that already existed for the purpose of biasing people and now it biases ai for free. A company would have to make an effort to remove or change the bias
Reddit has already begun the effort to start poising the well - https://www.reddit.com/r/poisonai/
Isn't that effort totally redundant? As in - plenty of people are already filling entire internet with slop for SEO purposes? And LLM slop by default is a mix of facts with few plausible but made up facts - it might be harder to craft such perfect poison on purpose.
I think most everyone already has a curated training library; Web scraping exists but I don't think anyone is still using it as a primary information vector
Otherwise they'd be slurping in their own slop
If I ever curate again it will certainly not be for the public. That led to PageRank which kickstarted this whole dystopian nightmare that Google has been planning since as early as 2003. No thank you.
This already exists, there are archives of Reddit or other sites, and Anna's Archive for papers and books.
Isn’t this what the paper-bound encyclopedia companies do, albeit shallowly
> There will come a day (and probably soon)
That day has already arrived, it is already happening.
Gemini has been a hilarious companion to my while I fixed the balance shaft chain guides in my old Mitsubishi triton (mighty max for US readers).
First it told me I could just remove said balance shaft chain as an emergency repair. Sorry Gemini, it also drives the oil pump.
Then it told me I could remove the water contaminated oil caused by removing the timing case by filling the crankcase with hot, soapy water and running the engine. Lord no.
Then it gave the wrong instructions for putting new gears on the balance shafts which meant the chain guides didn’t align with the chain. I’ll do it my way thanks Gemini.
The rest of the mistakes are too trivial to recount and sure it’s a pretty obscure subject but if I trusted it with a topic I’m not familiar with there is a huge potential for damage if you blindly follow it’s overconfidence. I miss normal searching.
Meanwhile, ChatGPT correctly diagnosed what was wrong with my plant from a single photo, identified which leaves I should cut, and annotated the picture showing where to cut and what not to touch.
I honestly expected a made-up useless generated image that matched the idea but not the actual thing.
Guess I’m still living in 2024.
What makes you think the chatbot's diagnosis is correct?
we're building the world's largest library and then locking the doors, letting the bots photocopy everything before the lights go out.
In a world, where everything can be stolen it will be hard to produce anything.
I still have hope though. Maybe the Internet will be better. Currently everything has to be monietized. Everything is ad heavy. At the beginning it was not so. People created things out of passion, or boredom. We can returned to that scheme.
I have seen neocities, personal blogs created and maintained in this year. I know I run my own "Internet index" https://github.com/rumca-js/Internet-Places-Database
This is a common misremembering of the early internet.
The internet was never ad free. The first ad was posted online in the 1970s (for DEC)! It pissed people off but not everyone: supposedly it generated $18M in sales. There was very little advertising back then only because the internet was restricted to a handful of large companies and universities.
The web itself was launched in 1991 and the early web was inaccessible to basically everyone as it required an extremely expensive NeXTStep machine. Windows didn't even ship a TCP stack in this era, iirc. So took a few years for the web to reach the point where it was usable at home. By 1995 the web was starting to become barely usable thanks to Win95 and Netscape, and DoubleClick launched immediately in the same year.
My memory of the early web is that basically every website had DoubleClick ads on them, it was notorious for that. "Punch the Monkey" was an early campaign. Almost every topic oriented website carried ads, partly because bandwidth and servers were very expensive so that helped defray the costs. GeoCities took off because it handled the complexities of running ads for you, so you could publish for free.
> People created things out of passion, or boredom. We can returned to that scheme.
Turns out, people want food and shelter more than entertainment. And psycho billionaires want money more than fun.
So there's a new Internet Archive, it's just split across 3,000 AI labs.
I wonder if the AI labs will throw out their own copies of the Internet Archive after it's been sued out of existence.
Probably not.
Kagi search today is better than Google search ever was.
And it’s clear that Google’s Ad model ultimately created a priority inversion. The advertisers became the customer.
I am so glad Kagi came along with a business model that is actually working.
I'm a long time Kagi user and I haven't used Google search in about a year.
I tried out Google search for a few technical searches recently and it was surprisingly ad and AI free. Not bad at all and much better than I remember from last year.
Then I put in some non-technical searches and it was all ads and AI and basically unusable.
I was wondering how would Kagi scale/expand if all of a sudden google were to stop serving search altogether (not likely) or alter search such that users look for alternatives.
Kagi is an aggregator for other, some paid, search APIs. They have, at least in the past, served some percentage of their results from Bing's API among others for example. Kagi seems to me to be dependent on these APIs being available, if they were to go away, so would Kagi.
I am a happy subscriber of Kagi though, they provide a really excellent service.
I'm not sure Kagi has ever used the Bing API, because (according to Kagi) Bing prohibited changing the results, or merging them with others. Apparently Google is expected to provide access to its index via API soon.
https://blog.kagi.com/waiting-dawn-search
I wouldn't say it's better, but it's certainly on par with Google in their best years. And it's light years better than what Google is now, or using an LLM.
The concept of an almanac seems relevant again: a yearly printed book with verified, accurate information. No manipulation at a later date, no AI hallucinations, etc.
The most famous one was probably Benjamin Franklin’s:
https://en.wikipedia.org/wiki/Poor_Richard%27s_Almanack
Paper encyclopedias might make a comeback for the same reason.
> a yearly printed book with verified, accurate information
Seems difficult to produce nowadays as even well researched topics are constantly attacked. Climate change papers as a small example.
Overbroad claim. Dramatic corollary
This clickbaity headline format cannot die fast enough
It can't. People will simply stop clicking links.
The world needs an anti-SEO search engine, that explicitly penalizes for something that appears SEO.
My personal anecdote is that I used to search for "xyz nutrition" quite often on Google. It used to provide a data table with lots of information. Sure, nutritional information is hard to get right, but at least that data was consistent.
Now Google just gives you an AI answer with random values pulled from blogs and Reddit. It's almost always blatantly incorrect. I genuinely can't understand why Google would destroy its most valuable search features. Disclaimer: I work for Ecosia, so I know for a fact that users really value these search widgets, and it was often cited as a reason they couldn't leave Google.
There's too much going on in this article and while I think some of the points are valid, others go too far.
A lot of the "cultural record" the author refers to is just digital junk. Random digital content that very few people care about, if we're being honest. Trying to hoard every bit of digital information ever produced is not the same thing as preserving "culture".
Case in point:
> Even the increasing use of ephemeral formats like Instagram Stories and WhatsApp status updates means that large portions of cultural, social, and political communication are never conserved in the first place. As a society, we can probably survive bad search results and come up with another way to schedule a sunset make-out session. But we can’t aspire to sovereignty if we can’t retain and retrieve our collective memory.
For most of human history, nobody was trying to "conserve" every cultural, social or political communication ever produced, and I fail to see how Instagram Stories and WhatsApp status updates, many of which aren't even truly broadcast publicly for all to see, are part of some imaginary "collective memory."
If you find a web page, see an Instagram Story or receive a message that's important to you, save it or take a screenshot. But let's not pretend all these things belong in a global Digital Civilizational Archives.
While I dislike Instagram reels and stories, I do feel like there is almost certainly scientific studies on radicalization that would benefit from actual histories of dumb memes. I feel like that about a lot of things, really.
There probably are some important hidden discord groups that would explain the origin of many political positions. Unlike smokey meetings in scummy bars, that exists now and is on a database somewhere.
On the other hand, when modern archaeologists discover "Claudius has a small dick" graffiti on the side of some God-forsaken wall in Pompeii, they're fascinated. The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid. What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
> The presence of such graffiti adds color and texture to the civilization inhabited by Virgil and Ovid.
It does, but do you think that people at that time thought anywhere near as much about preserving their scribbles as we do?
I'd venture a guess that we've created more "content" since the advent of the internet than in all of human history prior, and most of it is stored on things that aren't even designed to last a human lifetime without failure.
The idea that we're going to save every piece of digital junk for posterity just isn't realistic or healthy.
> What's just disposable background noise to us may provide context into how we lived and thought to our far-future descendants.
You're right, but you're also assuming that they're going to care that much, and that we're going to survive that long.
I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
> I think it isn't that we should save all of it, it's that we are not saving any of it. Random letters, notebooks, calendars, family photos, restaurant menus, etc have all proven useful to various historians, of which there will essentially be none from our era.
Well as far as digital letters, photos, menus, etc. are concerned, there's nothing stopping someone with the means and motivation from investing in a doomsday storage facility specifically designed to store these things for posterity. If people can do this for crypto they can do it for digital content.
As for physical items, do you know how much junk Americans have in storage units? The US self-storage industry generates over $40 billion in annual revenue. We're probably keeping more "stuff" in storage units where it has a chance of surviving a zombie apocalypse than at any point in human history.
I mean, none of that stuff in storage units is human communication, though. Nobody prints emails or text convos or insta photos. Nobody is mailing each other letters. So none of that stuff is in storage units, though plenty of it from previous generations might be.
I don't necessarily think all this digital stuff is worth saving, i just see how it could be that we are essentially erasing our modern historical record by putting 100℅ of it in private data centers with no permanent records.
And then you see how plenty of contemporary digital media from the last 30-40 years is actually already lost to time just in our short timeframe. I don't think acknowledging that this could be detrimental is necessarily an argument for trying to preserve all of it. We, as a society, went very rapidly from preserving a lot of it to preserving none of it. I don't see the harm in thinking about the implications of that.
Nobody actively preserved letters and menus and postcards and photos, they just persisted by nature of being physical. Digital records only persist with continuous effort to preserve them. So the record for future historians will be highly curated and likely much more limited.
Hello negative feedback loop.
It seems like a lot of techies should have chosen acting as a specialty, because I’ve never seen that much melodrama in one industry. FFS people, adapt, stop complaining
It's hard to get past the beginning of this article and take it at all seriously. The quote someone who missed a sunset because they asked google and supposedly got the wrong time... but they didn't think to look at the sun or lack thereof to check? Also when I put "when does the sun set today" I get a single exact figure at the top of my results, not from AI, which is honestly the best kind of result – an exact correct answer.
They were planning their day for a specific time to be somewhere. If you're looking at the sun it's to late...
From the article:
> “I had the projector set up outside and was waiting for the sun to set,” wrote one Facebook user in Colorado Springs, “but to my surprise I was simply living in the past. AI informed me the sunset had already happened.”
It does sound a bit bizarre.
We really need a SarcasmAI to become a thing.
That's fine, but arguing that Google should give accurate times based on where you are is ... crazy?
In the pre-LLM days it was cool that Google did weather, unit conversions, sports results, etc. But that's not even close to their value proposition. Even 5 years ago if someone told me they had planned a photograph and got it wrong because Google gave them the wrong time for sunset, I would have called them a moron for relying on Google! There are sites and apps dedicated to this. Use one of them!
I don't understand your argument. It's not their core value proposition so it's crazy? That's quite a leap. But giving a good answer is just as core as their search results. The answer is what you're there for and they get to show ads. It's not any stupider to use that info than a dedicated site. Both could be wrong, and you're not a moron if it is.
Installing an app to learn a single time would be the real moron option.
Perhaps we should let the content economy crash so that the big monsters eating it starve and die. A new content economy could be built on their carcasses instead.
Gemini's highest tier of plan has the utility of Google search circa 2010. I'm paying $400 for the privilege.
we built extremely powerful plausible bullshit machines targeted toward and trained on an electorate and population that has historically bad education and reading levels, built on top of an already fraught and fragile web which was also built off predatory basically unregulated behavior with a shaky relationship with “truth” and are surprised people have no idea what’s going on?
This was the whole point of it all and why the people in power have bet the farm on it.
I'd quit reading HN long time ago if not for comments like yours, when people call a spade precisely a spade and reignite my smoldering hope for the humankind.
What a joke. Google did that already 10 year ago when they started evicting stuff from its page rank caches.
AI is merely amplifying what social media and FAANG in general already have done to the 90s web.
Before AI, the way people were gaming Google was by filling their pages with pages of verbose fluff (have you ever tried googling how to make a specific cocktail?). Now the AI just makes it easier to generate that fluff.
I see only one simple solution (though we should discuss the more complex ones): if Google directly answers a search query, then it must be held accountable for it, for better or for worse, and therefore assume all the benefits and (legal) liabilities that this entails
I just don't care man. Did these people just log on yesterday or something? The MUDs and MMOs I grew up with disappeared. The IRC networks, and especially forums I learned so much from are gone (these were largely killed for something even worse than AI: commercial blogs!). Effectively every social network I've ever cared about has been ruined, failed, or sold. Even games have become bottom line chasing, live service slop that you never truly own.
If you want it so bad stop crying and make it, that's what I've been doing. It works a lot better than whatever this post and many of these comments are. There are dozens of us !
I remember from the MUD era there was a paper (possibly by Richard Bartle) on the "MUD lifecycle", and how they tended to last on average two years before the operators got bored / burned out / the community left / there was an Incident.
The old net is still there to a degree, but we live in pre-Alta Vista times again. It is hard to find. (I know as I run an old-school Forum and RPG-Game for more than 20 years now.)
I'm working on using the OpenZIM format to archive the web and to make the wikis seedable (and locally hostable for LLMs) so that the ongoing cat and mouse game anubis defense can stop.
My hope is that with the torrent protocol we can make the archived knowledge discoverable and seedable, because currently there's only the web archive and the kiwix download servers for archived contents. Both of them still are centralized servers that bear the cost of hosting those files.
- [1] https://github.com/cookiengineer/gozim
- [2] https://github.com/cookiengineer/zimdex
It happens when the deterministic precision was given away in return for the probabilistic guesses. It was one extreme until now (deterministic code), and we are swinging to the other extreme (probabilistic slop), but what the world wants could be somewhere in between. Some information does not need too much precision, while others do need precision.
That is a nice project -- classifying it based on reliability: human-made, sources, etc. + score.
Maybe a job for an AI to do? :D
Funny as I just cancelled my Kagi sub to get the Gemini ai pro sub. The deal was too good to pass up
The deal will be great for now, while they lure you in and get you to drop the competition -- then they will raise prices later.
Then you cancel and go to another provider, rinse and repeat. This is already what is happening with streaming platforms.
Not after it consolidates, monopolises and enshittifies, no you can't.
This used to be called "dumping" or predatory pricing[1] and would be fined in the physical retail space. Too bad our regulators are asleep at the wheel.
[1] https://en.wikipedia.org/wiki/Predatory_pricing
Kagi does have AI too, for what it’s worth. I found it pretty damn good, but I wish I could give them more money. I have a year subscription so I’m stuck without being able to give them *anything* until it’s done.
Sign up for another account?
That is an option for sure, but a very disappointing one. And it comes with a meaningful danger of ending up with 2 subscriptions: one yearly and one monthly. They really, really need to finish their “pay as you go” system. I don’t know why it’s not done yet and there’s been no word about it to my knowledge.
Maybe they should just enable a tip option? What do donations do to a company's tax liability and would it be worth their effort to enable something like that?
They already have a system for “prepaying”, but it can’t be used to put more credits into your account, they just sit there, waiting to be used by a future subscription.
[2011] The Atlantic - Why Google Won't Survive the Facebook Threat.
I’ll add this article to the list of incorrect predictions lol
It would be beneficial for the author to look at the revenue for Google Search. Revenue is still growing.
Indeed.
Not only is search revenue growing, but it is growing at an accelerating rate.
At the same time, operating margins are expanding.
I don't know the name of the logical fallacy where someone personally uses an LLM instead of Google Search and then infers that the search business is dying, without ever reading a financial statement.
Indeed. Money over everything. It is absurd to complain about decreasing quality of a product that makes increasing money for increasingly rich people (thus, by definition, better).
If business owners are paying for ads, then it doesn't matter a iota to Google or Facebook if real people use their services. It would be even better for them if real people didn't use their services, since that would save some costs. Business owners are going to keep paying for the ads, as long as they get some number about how many (bot) impressions their ad generated.
There is probably a lag time before advertisers give up on AdSense.
A lot of corporate ad spend is already planned, and Google can adjust the costs up as much as they like. They hold the lever.
If search engine competitors were really eating Google's lunch then ad impressions would go down.
They don't have much of an offering yet. OpenAI has some conventional ads, but everyone expects more involved advertising integrated into the conversation itself somehow.
Even if the alternatives would have no advertising the number of ad impressions of Google Search would go down.
In the article it mentions the rise of other search competitors like Qwant implying they are causing Google to die as it bleeds market share to them.
From a user perspective, Google search is the most useful it has been in years, though that doesn't feel entirely like intentional improvement, just a lucky side effect of the move to "AI mode".
And yes, if you take what the AI tells you at face value it could be wrong. But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
And also, yes, the old balance of Google driving clicks to sites that will then generate revenue off more Google Ads being shown after you click through to them creating a virtuous cycle is completely busted, and that sucks. It does not impact me directly but it certainly seems like unless a better system is devised that it is one of a few ways in which AI is likely to stall out its own training funnel.
As someone who regularly reads things online, then wants to read them again like 3 years later, Google has been monotonically declining in quality.
Also I realized the other day how hard it is to find song lyrics for anything other than quite mainstream songs.
It used to be I could always find the quote I wanted from a book. Nowadays I need to keep my own copies...
> Nowadays I need to keep my own copies...
https://github.com/asciimoo/hister
> Hister is a private search engine for the pages you visit and the files you keep. It indexes their full contents so you can find information again from the web interface, terminal, or an AI assistant connected through MCP.
> As someone who regularly reads things online, then wants to read them again like 3 years later, Google has been monotonically declining in quality.
I generally agree, but I think AI mode actually improved things somewhat compared to how things were just prior to it existing.
And I'm not saying what we have now is better than Golden Age Google, but things were just getting worse and worse for almost a decade. AI didn't fix the decade worth of decline, but it is the first thing I've seen from Google that at least partially reversed it for my own usage.
I think they make things worse, because they very very often present straight inaccurate information.
Just the other day I was trying to find out "What american tree species have the deepest roots". And all the AI responses were giving me back generic lists of big trees and claiming that roots going 20ft deep were the deepest. I know for a fact the mesquite trees behind my house can easily grow roots > 100 ft deep.
If I had clicked on the articles with generic lists of big trees, I would have realized they were all low quality clickbait sources and moved on. But the AI presentation makes you think that the information comes well-researched.
> But if you are aware of this and aware of the ways in which LLMs are likely to shit the bed, it is quicker to get from request to useful information than it has been with Google search since like 2017.
The point is not about 'quicker' requests but precise requests. It definitely has worsened, though not on a single degree on al levels like the HN hivemind claims, but some aspects are still somewhat precise but others are definitely crap.
i.e. when searching about my neighborhood it still returns better results than bing, yahoo, ddg, yandex and what have you. But they are buried into a load of crap of alleged "relevant" results (those things past the ai stuff) that aren't relevant in any way.
Yandex is the only search engine left which still feels like the "old" web. It feels like you're actually getting a best effort search, and not just the results that someone paid to put in front of you.
The most frustrating part of all of this is the underlying premise that the internet is, has been, or could ever be a credible cultural record is deeply stupid. Or maybe more charitably it's both historically and technically illiterate. It has always taken continuous unwavering effort on some person's part to keep any given piece of content online. And while managing a simple hosting account and updating domain registration periodically doesn't take a tremendous amount of effort 20 years is a long time to expect anyone to maintain enthusiasm. The internet has always been a frothy, ever changing blend of the odd nugget of truth drifting in a sea of unadulterated bullshit. Treating this, or worse what comes from statistically averaging it, as a source of capital T truth is totally unhinged. From whence did this mythology of online truth spring?
What’s come next for me has been much better. I use ChatGPT cranked to Pro with “extended” thinking to one-shot whatever I would’ve spent time looking into with Google. It’ll plan the whole sunset bike ride or promposal or whatever from TFA.