SEO and then AI
Ask Lemmy
A Fediverse community for open-ended, thought provoking questions
Rules: (interactive)
1) Be nice and; have fun
Doxxing, trolling, sealioning, racism, toxicity and dog-whistling are not welcomed in AskLemmy. Remember what your mother said: if you can't say something nice, don't say anything at all. In addition, the site-wide Lemmy.world terms of service also apply here. Please familiarize yourself with them
2) All posts must end with a '?'
This is sort of like Jeopardy. Please phrase all post titles in the form of a proper question ending with ?
3) No spam
Please do not flood the community with nonsense. Actual suspected spammers will be banned on site. No astroturfing.
4) NSFW is okay, within reason
Just remember to tag posts with either a content warning or a [NSFW] tag. Overtly sexual posts are not allowed, please direct them to either !asklemmyafterdark@lemmy.world or !asklemmynsfw@lemmynsfw.com.
NSFW comments should be restricted to posts tagged [NSFW].
5) This is not a support community.
It is not a place for 'how do I?', type questions.
If you have any questions regarding the site itself or would like to report a community, please direct them to Lemmy.world Support or email info@lemmy.world. For other questions check our partnered communities list, or use the search function.
6) No US Politics.
Please don't post about current US Politics. If you need to do this, try !politicaldiscussion@lemmy.world or !uspolitics@lemmy.world
7) No Hit-and-Run questions.
Please don't delete your post for no apparent reason. If you plan on deleting a question later, say so in the post, or if you feel that you have a good reason to remove it, message a mod beforehand. It's not fair to the ones who took their time to answer, and it's not in the spirit of the community.
8) No Bots.
Posts or comments from bots, LLM's, AIs, Neural Networks, Transformers, or Marvin the Paranoid Android are not welcome in AskLemmy. Real humans only please.
Reminder: The terms of service apply here too.
Partnered Communities:
Logo design credit goes to: tubbadu
This. Right before AI became public, I remember Google searches were getting more and more crammed with SEO bullshit. Now we have AI powered SEO and searching for anything is super difficult.
There was a piece on npr a few days back, all things considered, that talked to someone who was talking about search and seo and ai and such.
The tldr was that seo tools before ai gave you a lot of information that you could use to grow a website. Ai search on the other hand is a black box that may or may not do anything at all.
Are you familiar with the concept of signal-to-noise ratio?
Good information is probably still out there somewhere, but sadly it is bound to get increasingly lost in an endless sea of stupid bullshit and generative AI slop. The internet will probably never go back to the way it was before, and will instead becoming increasingly polluted with useless recycled data.
It's up to the rest of us to begin collecting and curating good data.
I don’t doubt that at all. But I hoped that searching with a date range limited to 2002 and earlier would turn up an article written at the time… but most all of that seems to be gone or in un-indexed archives.
I'd be so down to code a basic ass website from scratch, hell, it can even be a static website where we host it on either some Asian servers where the west has no jurisdiction to send DMCA requests, or on a decentralized network chain, and fucking moderate it, search for the news people can not find, publish it, and all that
It all boils down to unfortunately, you still need some kinda money to operate this shit. We have a problem in this modern age with unification, and being the voice for others. Financially it's so easy to develop this, if people could give a shit and throw a few buxes onto the people who would take the responsibility of such website.
I've had okay luck with the internet archive for news articles.
News sites don't keep things online forever, even if they are still operating. I wish I knew that there was an industry standard and what that standard was. I'd save more things for myself knowing there was an expiry date.
I think that would be nice.
I think it'd be cool if there were a common standard for a "Paper of Record". A news outlet is considered a Paper of Record if it makes its publications public within a certain amount of time (say, 60 or 90 days after initial publication), keeps it available for a minimum amount of time so that archivists can copy it (ideally, provide a torrent), and don't modify the contents after they're published (appending updates is OK. Encouraged even, in my opinion). Maybe some other criteria I haven't thought of in these few minutes too.
There could be a torrent network of archivists, each hosting copies of news articles they consider important. anyone could join the network at any time and get their own copies of articles to host, maybe because a certain article is down to a few copies and they want it preserved. It would be a huge boon to historians I bet.
I just tried it, it has really spotty coverage from 25 years ago. I was only able to find a very few pages and even what it does have is just the front page kind of view. Nothing specific.
Took 3 minutes ...
It was an amazing phone call. 40 Wall Street actually was the second tallest building in downtown Manhattan and it was, actually before the World Trade Center, the tallest. And when they built the World Trade Center it became known as the second tallest and now it's the tallest," the future president told WWOR at the time.
Even has the audio, he was calling in to Howard Stern
I wanted the original, day-of commentary. The actual article about what he said, the day he said it.
I have been downvoted so many times when I've responded to claims that "nothing ever dies on the internet" or like information once on the internet never goes away.
sure, a lot of it is still around, but a shit ton is NOT. Case in point.
But sorry, I have no topical answers to help :(
If it's useful information, it's likely to fucking disappear.
If it's embarrassing toxic shit that will haunt you forever, that is now a meme and everyone will have a copy.
They’re actively killing it.
Who’d have ever thought that microfiche would be the pinnacle of our archival technology.
Librarians
It only exists if the search engines say it exists.
Only about 15% of world texts, documentary, and archival materials have been digitized. But, since two generations now have no idea how to access data off line, it might as well be on the moon.
Imagine the power that comes from controlling what everyone can find on the internet.
Pretty bloody easy to find, I just searched for "trump claims he has biggest building in Manhattan after 9/11" https://www.snopes.com/fact-check/trump-bragged-tallest-building/
Tldr it wasn't a callous lie, and it was pretty close to being entirely factual
it was callous and it wasn’t true, but it was not a lie used to any particular end.
it was callous because a bunch of people died, but he doesn’t have any room in his tiny head to fit empathy
It's the tallest by number of floors, but technically 20 foot shorter by absolute height, I think we can let this slide
I think as news sites update they lose some of their old stories. Internet archive may have what you want.
it seems pretty difficult to find ways to bracket dates on search engines
Are you searching on your phone? I haven't used Google in a while, but it used to be that you could enter specific dates for custom search ranges on news articles - but only in desktop mode. If you're searching on your phone, try switching to desktop mode, then entering your custom dates.
I did it on phone and desktop. The search modifiers are the same.
Search engines are fucked and will get worse, we need open source search engines desperately.
But as to your question, I saw an article just this morning with prez' statements on 9/11, in the guardian I think. I didn't read it because I try to ignore things prez says while following actual events.
https://www.theguardian.com/us-news/2026/sep/09/trump-911-falsehoods
I guess it's only the top 6 sorry.
I went to google.com and typed in "trump tallest building in Manhattan" and there it was.
This extension helps a lot - https://addons.mozilla.org/en-US/firefox/addon/nogoogleai/
All the articles I found of that type were current articles referencing what he said…IOW “Did trump actually say thing about 9/11?” “Video resurfaces about trump…”
I wanted the original text in situ from the day of. But maybe these will have to do.
Oh, I see!
Yeah the internet was a different world in 2001, hardly any of the websites from that time are still running and those that are will have rebuilt themselves multiple times over, often changing url structure in the process.
It's still there, but Google has no incentive to rank it particularly high due to SEO and how they make money.
There are indexes that are better at searching the small web, instead of funneling everything to Facereddittube l. Kagi is a my recommendation, but there are many others.
Honestly I don’t know if anybody was talking about trump’s dumb comment at that point except the live news show he called in to
On sep 16 the LA Times posted an article about “NYCs tallest building, again” (The Empire State Building), and if they had caught wind of trumps comments they might have included it as being incorrect.
https://www.latimes.com/archives/la-xpm-2001-sep-16-mn-46445-story.html
My first guess for something like this would be your local library's databases, or Archive.org. 🤔
It's crazy how poisoned search has gotten over the last like ten years.
From what I can tell, the search engines now heavily favor recently updated and new content and push old content way down in the results or as it seems more recently will quite often not even show it at all.
Have you tried old sites that actually scan/preserve old newspapers? The moguls cannot scrub them words away as easily.
This kind of thing is just about my most frequent use-case in GPT. I find it generally very good for crafting complex searches, and it can go in steps, offering to search various types of archives as you direct it and redirect it.
The great thing is that you don't have to fish around for just the right words, as with traditional search engines. Just explain in plain language, and away it goes.
EDIT: Rock on with the paranoid, anti-AI downvotes! 🙃
I was searching for who said a specific comment about a book that Churchill wrote. I remembered the comment correctly but not who said it and where. So I asked the AI to see if it can find it for me. It did, right away. I was satisfied. But then I started to doubt myself a bit, tried to find the source but I couldn't. I asked the AI to make sure that it was suggesting the correct person, turns out, it wasn't. It even apologized, said "sorry, I was mistaken, it wasn't X who said it, it was actually Y!" and then I was again happy... For a second. I Asked it to make sure, perhaps find a source for it.. It couldn't. So I gave up on that, took out my copy of "Walking With Destiny" and flipped through the pages for like 4 hours and found the quote and who said it and where. The AI didn't suggest the correct answer at any point.
So yeah... AI might be good at times, but sometimes it just makes shit up for you. I still use it to find stuff, if I cant find it myself, but I ask it to find the actual source and give that to me, instead of just trusting it.
Yes, good points. Hallucination is common. LLM's have a bias towards hallucination and taking shortcuts. Part of why I always try to err on the side of simple queries, as well as ask for sources on anything that looks promising.
Also, different LLM's and different flavors of LLM's can fetch considerably different results. Google's built-in Gemini is still hilariously bad at times. Personally I usually use GPT, partly because it lets me program it in ways that stay global between sessions. Very helpful for getting it to stay on point and sober with the facts.
It's not unlike curating one's FV or Reddit feed-- you have to put in the work over time setting it up to best answer your particular needs.
Does it not bother you, search engines make their ai agent at the top of the page, then remove applicable results, so you have to use that agent to get that information?
You can't say they need the ai to find the results, because we saw search engines work better without them for 20 years.
I've only seen that with Google. If the result is useful, then I'll consider it. If not, then I'll scroll past to the search results.
Google's pure convenience, though. It's been compromised for years, but does have a huge database. If I'm not finding what I want, then I'll try something else.
Problem?
All the search engines became shit at once. Also I was forced to use bing the other day, shich ddg also is a sort of instance of, and both also have the ai agent lording over the results they no longer offer you.
All the search engines are engaged in an illegal shittrust to maximize revenue, and they've all removed many many applicable results, from links to streaming and downloading, to news articles, to everything.
Pure convenience my foot, Google is a shit company actively becoming worse, and only a frog in a boiling pot of google's shit doesn't see how bad it's gotten since 2021.
*ribbit!* *ribbit!*
You know there's an extension you can add that blocks Gemini from search results, right?
However, if you're truly that hand-wringingly concerned over this stuff, and don't like DuckDuckGo (or whatever), there's also a P2P search engine project, SearXNG: https://searx.space/
Thanks for the tip, I tried those other alternates and none worked well so far, the pay one is supposed to be good, maybe I will give that one another chance I do appreciate it.
But that's neither here nor there. They removed results we wanted, to force us to use their AI. That's the big problem here.
And google's full of shit about removing gemini. I followed multiple recent articles about removing it from my phone and none of the instructions worked, not even the ones published by google.
Idk about google search specifically because I use startpage, a proxy server that runs google through it, based in the netherlands, started in the wake of the Snowden revelations, that keeps your browsing history anonymous. But you can also view pages anonymously from it, though not all sites work with that.