LLMs and their precursors were always how search engines compiled and ranked their search results. Google has been using some form of AI to build up its base of information for decades.
The absurd turn of the last decade was Google C-level pivot from engagement through useful results to engagement through bad results that force you to click around on the site longer. “How do we get more ad interaction?” became the question that dominated the development of their technology.
given that by LLMs, people generally refer to the transformer architecture, it definitely wasn’t how google always did search. they famously started out by using pagerank. now, back then this was definitely seen as AI, but what gets called ai by the general public has changed over the last couple of decades. now people use it almost exclusively to refer to deep learning models or even just llms.
so yes, in a strict sense google has always used ai for search. but when the general public says AI now, they usually mean llms specifically and using that to replace actually going to the source websites is new
Pagerank was built on the Markov Model, which became the core of modern LLM graph design.
It was a precursor technology, not an independent system, which used Hyperlinks rather than word tokens to form associations.
they usually mean llms specifically and using that to replace actually going to the source websites is new
So much of the complaint around AI isn’t even the LLM part. It’s the weighting and censorship of certain sources and data, convinced with the masking of results behind natural language output templates.
You had the first problem under Pagerank already (Google Bombing and “shadow banning” of sources the company didn’t find reliable). The second is entirely a product of the interface and has nothing to do with the heuristics of the program.
sure both pagerank and llms can modelled as marlov chains, but the working and output of these models is very different. it’s disingenuous to state that google used llms and their precursors since the beginning. their precursors maybe but they are very different beasts.
as for the complaints, for me a big problem of their use in search engines is that they often hide the sources of the summary they provide and disincentivize people from clicking through to the source sites. This is very much a new problem that didn’t exist before google included output from generative models
Yeah. It’s almost impossible to create a system grounded in ethics in a hyper-capitalist architecture. Misinformation is by its very nature more entertaining and easier to make, so the unavoidable market incentive is to eventually shift the revenue base away from a broad, but low capitalization user base toward the handful of high capitalization snake-oil mongers. It’s a direct product of market concentration that’s allowed a handful of talented liars to accumulate more wealth than the majority of the rest of humanity.
I’m Canadian and their prices aren’t adjusted per market, so it’s pretty expensive for us. It’s about the price of YouTube Premium for a base plan. (I’ll go double check)
Edit: nope, YT Premium is a wild $27 CAD now, far from the $14 CAD for the cheapest unlimited search Kagi plan lol
$5 a month is really not expensive, though. That’s less than the price of a AAA game for a full year. Google is only free because they make their money selling user data. How cheap does a good search engine need to be?
That’s wild to me; I feel like I search a lot, and I only go over the 300 searches rarely. Those few times I do, it’s painless to just start the cycle a few days early.
I’m not who you replied to, but in the same boat as him. I burn through 300 in a week, let alone a month. In my kagi settings I can see I make around 1100 searches each month. That 300 limit is just not possible for me
Interesting! I guess I must search a lot less than I feel like I do. I do some things like type in URLs rather than searching for a known URL, and going to Wikipedia directly and then searching there; maybe that stuff helps more than I’d think? Or maybe we just use the Internet differently, which is fine too.
ah, yeah, I am not sure I have ever seen a website have decent search of its own so I use Kagi for basically every search lol. I just enter my query then add “site:site.com” to the end to limit it to results from that site
Impossible. The problem is not in the search process, but in data availability. LLM gives you a concrete concise answer to your question while the search engine gives you dead links or spam.
A concrete, concise answer that’s often confidently incorrect.
I’ve had MS Copilot first tell me to use a Microsoft product that was retired nearly a decade ago, and when called out it told me to use one that didn’t exist.
Also, the garbage web results are because of LLMs.
LLM gives you a concrete concise answer to your question while the search engine gives you dead links or spam.
LLM also very confidently gives you a wrong answer way too often. Just over the weekend I tinkered with diy-home automation gadget and fed a bunch of documents to duck.ai in hope that it could dig up protocol and electrical behavior of HAN port on our electric meter. The documentation is there, but it’s recursively referenced to a crapload of different standards so the actual answer is pretty labour intensive to dig up manually.
LLM spat out an answer which seemed to make sense, but the solution it offered didn’t work and when I asked for sources for spesific facts it just said that it doesn’t have any and that it shouldn’t have recommended such a solution in the first place, with apologies of course.
I’d much prefer the old school way of searching things, but that doesn’t really exist anymore either. Search engines, google specially, are far less useful than they used to be. Part of the problem is LLM slop filling the whole internet but somehow I feel that’s not the only reason.
Here’s an example of a time where a single letter spelling error affected the outcome even though Google knew it was a typo. Yea, LLMs will give an answer but anecdotally it’s correct less than half the time on any question of substance and at least 10%+ miss rate on simple questions.
Ditching Google is good, but not for AI. We should take the ditching part and redirect the young people to use other search engines.
LLMs and their precursors were always how search engines compiled and ranked their search results. Google has been using some form of AI to build up its base of information for decades.
The absurd turn of the last decade was Google C-level pivot from engagement through useful results to engagement through bad results that force you to click around on the site longer. “How do we get more ad interaction?” became the question that dominated the development of their technology.
That’s what enahittified their website. Not “AI”.
given that by LLMs, people generally refer to the transformer architecture, it definitely wasn’t how google always did search. they famously started out by using pagerank. now, back then this was definitely seen as AI, but what gets called ai by the general public has changed over the last couple of decades. now people use it almost exclusively to refer to deep learning models or even just llms.
so yes, in a strict sense google has always used ai for search. but when the general public says AI now, they usually mean llms specifically and using that to replace actually going to the source websites is new
Pagerank was built on the Markov Model, which became the core of modern LLM graph design.
It was a precursor technology, not an independent system, which used Hyperlinks rather than word tokens to form associations.
So much of the complaint around AI isn’t even the LLM part. It’s the weighting and censorship of certain sources and data, convinced with the masking of results behind natural language output templates.
You had the first problem under Pagerank already (Google Bombing and “shadow banning” of sources the company didn’t find reliable). The second is entirely a product of the interface and has nothing to do with the heuristics of the program.
sure both pagerank and llms can modelled as marlov chains, but the working and output of these models is very different. it’s disingenuous to state that google used llms and their precursors since the beginning. their precursors maybe but they are very different beasts.
as for the complaints, for me a big problem of their use in search engines is that they often hide the sources of the summary they provide and disincentivize people from clicking through to the source sites. This is very much a new problem that didn’t exist before google included output from generative models
Search is probably one of the few “ok” uses for LLMs if you find a way to reduce the carbon footprint and train on ethical material
Which, up until maybe eight or nine years ago, Google was pretty good at.
But then we went into this insane build-out cycle, where you were seemingly incentivized to be extra sloppy with every aspect of the business.
I think we need more solutions like Kagi, but less expensive.
Since you’re the customer, no ads are served to you, and the search results are in your best interest and can be fully customized.
I remember the streaming services making this claim when they first took off.
Oof :(
Yeah. It’s almost impossible to create a system grounded in ethics in a hyper-capitalist architecture. Misinformation is by its very nature more entertaining and easier to make, so the unavoidable market incentive is to eventually shift the revenue base away from a broad, but low capitalization user base toward the handful of high capitalization snake-oil mongers. It’s a direct product of market concentration that’s allowed a handful of talented liars to accumulate more wealth than the majority of the rest of humanity.
I thought Kagi was expensive until I used it for a year. Worth every penny.
I’m Canadian and their prices aren’t adjusted per market, so it’s pretty expensive for us. It’s about the price of YouTube Premium for a base plan. (I’ll go double check)
Edit: nope, YT Premium is a wild $27 CAD now, far from the $14 CAD for the cheapest unlimited search Kagi plan lol
$5 a month is really not expensive, though. That’s less than the price of a AAA game for a full year. Google is only free because they make their money selling user data. How cheap does a good search engine need to be?
I’m not considering the 300 searches plan, I bust that basically instantly. The minimum plan I can use is the $10 USD, which is about $14 CAD.
That’s wild to me; I feel like I search a lot, and I only go over the 300 searches rarely. Those few times I do, it’s painless to just start the cycle a few days early.
I’m not who you replied to, but in the same boat as him. I burn through 300 in a week, let alone a month. In my kagi settings I can see I make around 1100 searches each month. That 300 limit is just not possible for me
Interesting! I guess I must search a lot less than I feel like I do. I do some things like type in URLs rather than searching for a known URL, and going to Wikipedia directly and then searching there; maybe that stuff helps more than I’d think? Or maybe we just use the Internet differently, which is fine too.
You also probably don’t use it for spell check when a random system doesn’t have it built in. I mean, I totally don’t do that… 🤣
Haha if I did I’d just go to wkitionary.com and search there or something
ah, yeah, I am not sure I have ever seen a website have decent search of its own so I use Kagi for basically every search lol. I just enter my query then add “site:site.com” to the end to limit it to results from that site
I’ve been enjoying using https://marginalia-search.com/ for various topics.
Impossible. The problem is not in the search process, but in data availability. LLM gives you a concrete concise answer to your question while the search engine gives you dead links or spam.
A concrete, concise answer that’s often confidently incorrect.
I’ve had MS Copilot first tell me to use a Microsoft product that was retired nearly a decade ago, and when called out it told me to use one that didn’t exist.
Also, the garbage web results are because of LLMs.
LLM also very confidently gives you a wrong answer way too often. Just over the weekend I tinkered with diy-home automation gadget and fed a bunch of documents to duck.ai in hope that it could dig up protocol and electrical behavior of HAN port on our electric meter. The documentation is there, but it’s recursively referenced to a crapload of different standards so the actual answer is pretty labour intensive to dig up manually.
LLM spat out an answer which seemed to make sense, but the solution it offered didn’t work and when I asked for sources for spesific facts it just said that it doesn’t have any and that it shouldn’t have recommended such a solution in the first place, with apologies of course.
I’d much prefer the old school way of searching things, but that doesn’t really exist anymore either. Search engines, google specially, are far less useful than they used to be. Part of the problem is LLM slop filling the whole internet but somehow I feel that’s not the only reason.
LLMs give you sycophantic bullshit. Tell it it’s wrong even if the response is correct and see what it does.
That may be true for basic, simple questions, but it is not true for medium or hard questions.
They will give an answer, yes, but often it’s not answering what you asked, or giving the wrong answer.
Here’s an example of a time where a single letter spelling error affected the outcome even though Google knew it was a typo. Yea, LLMs will give an answer but anecdotally it’s correct less than half the time on any question of substance and at least 10%+ miss rate on simple questions.
Then you use a better search engine, not flock to an AI that has no real concept of what’s true or false so just feeds you whatever looks right