I've been going mad with facebook search. It ignores every search parameter. Geofence? Here's results from 500mi away.
Length of item? nah here's soemthing 50% shorter
Price? Merely a construct.
They are obviously and purposefully keeping the feed full to get as many ads in front of you as possible, not to give you accurate results.
Contrast this with craigslist which actually gives you a "zero results" found if you are too restrictive. Hell, even the autotrader sites are going the facebook way now. It's maddening.
Yeah, but as you're "going mad" at Facebook, you keep engaging with the site.
You try to peer through the noisy results to see if any of them are actually a match, you try again with different terms, you give up and engage with a result outside your parameters, or (best of all) you give up and engage with an ad or get distracted by some other Facebook feature.
As far as they're concerned, engagement is the key factor and satisfaction is only a tiny little bonus when things work out. Engagement is what earns them money and makes them make bigger claims in investment statements and is what keeps eyeballs away from their competitors.
Everything is working perfectly for them... just not for you, and they don't really care. That was the big business insight of the 2010's and there's still very little reason for anyone in online media to stop aiming for it.
I have the same complaint but with Facebook Dating (yes, of course that's a thing). The app itself is less enshittified (so far) than many other dating apps (thanks, Match Group) but the filtering is just dumb.
"Hey FB I live in a major city and don't want to match with anyone further than 25 miles away" okay here are a bunch of matches from cities 200 miles away
"Hey FB I don't want to match 'as friends' because what even is that" okay cool we'll make sure to ignore that
Pretty sure this was written by an intern for their summer project. At least it's free.
And I guess I'm still on there, so maybe FB knows what they're doing
Maybe I'm just dense and there's some weird psychology thing FB is doing here. but I can't think of any reason why ignoring search parameters would be intentional, especially for a classifieds search! Even if it's for sponsored listings, they're just bogging down the conversion rate.
If FB marketplace is showing me a thing I can't buy because it's too far away to pick up, doesn't fit in the space I'm specifying, and over my stated price range, I'm just going to assume there were no results and look elsewhere. I guess FB gets paid for an impression, but an advertiser that sees impressions going up and conversions not changing is a pretty unhappy advertiser.
Yeah, I'm sure the newsfeed is an example where ignoring settings/search params like this works, because there's no real hard-lines to what posts someone will be somehow able to read other than maybe the language it's in, and even that's flexible. You can just blast users with whatever slop pays FB the most. But some sponsored decor tchotchke being 500 miles away might as well be on the moon.
They don’t actually give a shit whether you find what you were looking for, they heavily optimize for their time-on-site metric. If the search results were terrible, and you had to refine your terms and try again, that’s more time on site! More time is more better!
I feel like amazon just appends the values you select on the left sidebar to your search. There's no way they're actually properties of the listings in any way.
Give me search like McMaster-Carr. Everything is an instance of a category class with madatory properties. If your product doesn't fit in any existing category, tough shit. Go convince someone at McBezos-Carr that whatever premade-landfill product you're selling is a category, along with providing dimensions/materials/properties that category varies by, and maybe they'll consider adding it.
Google search still has that feature where the little preview blurb will bolden where it matches parts of your query so you can see what context the pattern appeared in. Always very useful but now it gives you a pretty good look at what google replaced your actual query with. Since I started paying attention to this, I've realized that most of the cases where I need to add a bunch of excluded keywords, I'm actually just gradually twisting googles arm into searching what I wanted to search. It takes a lot of excluded keywords sometimes until it finally runs out of ideas for stuff to replace my entire query with.
It’s wild how broken search has become, LLM made it worse but it was already downhill prior to that. Even Reddit forces you login to search within a subreddit and displays the annoying “Best place on internet” overlay after staying on the site for 2 min. I just use AnythingLLM with DuckDuckGo + local LLM model to do searches now.
> Searching on YouTube means scrolling over two entire rows of Shorts,"People also watched", "Explore more", with your actual results somewhere in between
Google ruins all its search interface. I never understood why I would need "people also watched". That made never any sense - neither on youtube, but even less so on google search. Of what interest to me is it when Charly Joker searches for hentai cyborgs? Even IF I were interested in that, I could search for this ON MY OWN. I don't need mental handholding, so this ALWAYS wasted space I could use otherwise. Thanks to ublock origin I block such UI downgrades anyway, but this is not the only example. Google abuses people here. Google wants you to use a privated web that Google controls. This madness has to stop. Google will not budge, so Google has to be ended by mankind.
I'm increasingly thinking that the only solution may be to have a search index with human editorial control. As much as I think Google deserves scorn for how badly they have made the search experience on the web, we have to remember that they're fighting a battle against SEO and spam. It feels like they have mostly given up. Any automatic solution will always be an arms race where they will be one step behind. The only way to get ahead is to ensure that all content has to pass through human moderation before it can be included -- that's the only way I can think of to break the cat and mouse game dynamics at play here.
Google, Facebook and others will never do this because it's expensive. Their cost savings is our toxic waste.
So what does a viable solution look like? I'm not sure. Maybe a system with distributed, volunteer moderators like Wikipedia. It's not perfect, but that model has held up remarkably well over its lifespan and has seemed to mostly hold its own against increasingly sophisticated attacks.
Figuring out how to scale it may be a challenge, but it may be worth considering something along the lines of the directory sites that used to be the main method of discovery on the web prior to the rise of search engines, which itself can be searched.
I would personally find this useful at least, particularly when diving into unfamiliar topics. In a lot of those situations, a handful of high quality links can preempt a litany of more granular searches.
This might be an area where a traditional, respectful search engine style crawler backed by a small cheap LLM paired with human curation would work well. The crawler+LLM finds sites, attempts to compute some quality score, figures out where they'd fit in the directory's hierarchy, and then submits them for human review and approval. If a poor quality link makes its way through, users can flag it for review and if curation agrees it gets booted.
While the article is more about specialized search (videos, shopping) for general web search, Kagi has to be mentioned. It's been so long since I've used another search engine (years), I'm generally shocked to see what they look like and how they behave now. It makes web search just work, including things like exact phrases with quotes, logical operators, etc. which are mentioned as generally rotting capabilities.
Honestly I haven't found kagi that useful in the last 12 months. I used to be a paying member but then dropped it because the results didn't feel more useful than just asking the free versions of Gemini, claude, etc
I had a huge rant a couple weeks ago about how I couldn't find anything in discord's search, it used to be pretty decent, but it's been getting worse as well...
Paraphrasing, but is it too hard to "SELECT FROM 'posts' WHERE 'text' LIKE %search%"?
The current state of things makes the problem seem worse than it is, but it is not a simple problem. Should a post about "L'acteur Gérard Philipe" be found by a search for "gerard", without the accent, through collation? Probably, but should quoted/verbatim searches also do collation? How about the non-verbatim but misspelled "gérard philippe"? What if the search has a correct spelling, but not the page? What if the search is for "ジェラール", or if the page only has the name "ジェラール", especially if the user specifies that they can read Japanese? Should "acteur" find "l'acteur", and vice-versa? How about "comédien gérard", even though "comédien" does not appear on the page? Should a search for a "theater" find "theatre", and vice-versa?
More importantly, how do you rank the thousands of hits?
These were all issues mostly resolved by 2000 or so. I remember integrating a search index into our company app in 1998-1999 and it handled most of these cases. The current madness is mostly deliberate bait and switch.
I've stopped waiting for sites to do what I want in terms of features and functionality -- their incentives and mine very rarely align.
I think the solution here eventually will be a search engine that lives on your personal computer (it's also what I'm building currently). Ingest content from the web -> sort/filter/view it locally as you choose.
Tbh I too have experienced the same but I think that is just what we are moving towards. I mean everything is now shifting to semantic search and it is better for natural language yeah it sucks that literals dont work as good now but "Optimization".
Yesterday I was doing an automotive auction photography gig with a 2006-ish convertible Chevrolet Corvette. I received the car with the top down, and in the midst of a 104 degree day needed to raise the top. No obvious buttons or latches were present and the owners manual, naturally, was on vacation as well. Googling "C6 Corvette top down button" showed where a button should be, but wasn't, so I concluded this car had a manual top instead of power. Thus began an infuriating cycle where Google ignored all attempts to search for manual Corvette top lowering instructions and instead simply INSISTED that what I was actually looking for was the instructions to manually lower a power top.
Agitation multiplied when I finally found the answer, which was reaching blindly under a plastic trim panel for a hidden downward-facing button to open the panel, lift the top up, and then balance the top midway between open and closed while I finesse the plastic tonneau cover panel back down under the rear part of the top. Garbage UX on both the part of Chevrolet and Google. I wish to become a hermit in the woods.
> what if we just show people the results we think they want
Author is on the right track but it's much more actively evil than this. The real algorithm is "what if we just show people the results we want to show them, because showing them those things makes us more money*?
* Except for services like Spotify, which shows you the songs that cost them the least.
It's obvious that the companies exposing search features for their own service data don't actually want you to be able to search the data with precision. They 1. think they know better than you, and so want to show you stuff besides "what you think you want", because their metrics say you might click it anyway; and 2. they want big long listings for you to scroll through, so they have more opportunities to inject paid promotional elements into those listings. It's just another kind of enshittification.
That being said... the first wave of web search engines weren't built by the companies hosting the content. Search engines as tools became popular because, even in 1998, browsing and navigating sites (esp. corporate sites) to surface content, was already becoming an increasingly adversarial experience.
Search engines were built by companies spidering other sites' content. This was content that was often — if you were navigating along these sites' happy paths — found five to ten links deep through a confusing warren of subtle and unintuitive click targets. This content was the original "deep web." And search engines made it shallow... often against the spidered sites' consent.
Companies at the time much preferred "directory" sites that would only ever link to their landing pages. To companies of that era, sites linking directly to specific URLs of your site would be like if the Yellow Pages listed specific directory-extensions of your company's phone number! But this first battle in the "war against deep-linking" was a losing one from the moment it began, since early websites were almost inherently bot-accessible.
The second battle, in the early 2000s, where sites were built as [non-deep-linkable] Flash apps, took longer to determine, only finally resolving (again in favor of an open web, and thus third-party search) when laws and regulatory compliances both started to force companies back toward UA accessibility.
But for going on a decade now, we've been embroiled in a seemingly-indefinite third battle in the "war against deep-linking" — this time with companies tucking every possible type of user-generated content behind a login wall of some walled-garden everything [web]apps. Permalink URLs for individual content-items still exist in these webapps; but you get redirected to a login interstitial if you visit them.
Where's the adversarial spirit of the first search-engine companies today? Where are the companies trying to index the modern "deep web" of login-walled content? Why can't any company give me a search box that searches "into" YouTube video transcripts, "into" Facebook Marketplace, "into" public Slack and Discord and Telegram communities (probably requiring archiving, ala how Deja News/Google Groups surfaced Usenet), etc?
Sure, this sort of thing is intensely adversarial (but see https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn — while it might be against a site's ToS, but it's not against the law.) And sure, this sort of thing requires per-service connectors (but any thought of this business model therefore being "unscalable" comes from a pre-coding-agent era.)
But isn't it also, very clearly, "what people want"?
They are obviously and purposefully keeping the feed full to get as many ads in front of you as possible, not to give you accurate results.
Contrast this with craigslist which actually gives you a "zero results" found if you are too restrictive. Hell, even the autotrader sites are going the facebook way now. It's maddening.
You try to peer through the noisy results to see if any of them are actually a match, you try again with different terms, you give up and engage with a result outside your parameters, or (best of all) you give up and engage with an ad or get distracted by some other Facebook feature.
As far as they're concerned, engagement is the key factor and satisfaction is only a tiny little bonus when things work out. Engagement is what earns them money and makes them make bigger claims in investment statements and is what keeps eyeballs away from their competitors.
Everything is working perfectly for them... just not for you, and they don't really care. That was the big business insight of the 2010's and there's still very little reason for anyone in online media to stop aiming for it.
"Hey FB I live in a major city and don't want to match with anyone further than 25 miles away" okay here are a bunch of matches from cities 200 miles away
"Hey FB I don't want to match 'as friends' because what even is that" okay cool we'll make sure to ignore that
Pretty sure this was written by an intern for their summer project. At least it's free.
And I guess I'm still on there, so maybe FB knows what they're doing
If FB marketplace is showing me a thing I can't buy because it's too far away to pick up, doesn't fit in the space I'm specifying, and over my stated price range, I'm just going to assume there were no results and look elsewhere. I guess FB gets paid for an impression, but an advertiser that sees impressions going up and conversions not changing is a pretty unhappy advertiser.
Yeah, I'm sure the newsfeed is an example where ignoring settings/search params like this works, because there's no real hard-lines to what posts someone will be somehow able to read other than maybe the language it's in, and even that's flexible. You can just blast users with whatever slop pays FB the most. But some sponsored decor tchotchke being 500 miles away might as well be on the moon.
Yesterday I searched for "3/8 inch handles", and got several ball-shaped ones, plus some things that included (but were not) handles.
I added the term - "3/8 inch ball handles" - and only the not-actually-handles remained.
Give me search like McMaster-Carr. Everything is an instance of a category class with madatory properties. If your product doesn't fit in any existing category, tough shit. Go convince someone at McBezos-Carr that whatever premade-landfill product you're selling is a category, along with providing dimensions/materials/properties that category varies by, and maybe they'll consider adding it.
Google ruins all its search interface. I never understood why I would need "people also watched". That made never any sense - neither on youtube, but even less so on google search. Of what interest to me is it when Charly Joker searches for hentai cyborgs? Even IF I were interested in that, I could search for this ON MY OWN. I don't need mental handholding, so this ALWAYS wasted space I could use otherwise. Thanks to ublock origin I block such UI downgrades anyway, but this is not the only example. Google abuses people here. Google wants you to use a privated web that Google controls. This madness has to stop. Google will not budge, so Google has to be ended by mankind.
Google, Facebook and others will never do this because it's expensive. Their cost savings is our toxic waste.
So what does a viable solution look like? I'm not sure. Maybe a system with distributed, volunteer moderators like Wikipedia. It's not perfect, but that model has held up remarkably well over its lifespan and has seemed to mostly hold its own against increasingly sophisticated attacks.
I would personally find this useful at least, particularly when diving into unfamiliar topics. In a lot of those situations, a handful of high quality links can preempt a litany of more granular searches.
This might be an area where a traditional, respectful search engine style crawler backed by a small cheap LLM paired with human curation would work well. The crawler+LLM finds sites, attempts to compute some quality score, figures out where they'd fit in the directory's hierarchy, and then submits them for human review and approval. If a poor quality link makes its way through, users can flag it for review and if curation agrees it gets booted.
Where did you get this idea from? They actively promote it.
We have to remember that they decided to side with the SEOmongers rather than crush them, because it was better for Google's bottom line.
Paraphrasing, but is it too hard to "SELECT FROM 'posts' WHERE 'text' LIKE %search%"?
More importantly, how do you rank the thousands of hits?
I think the solution here eventually will be a search engine that lives on your personal computer (it's also what I'm building currently). Ingest content from the web -> sort/filter/view it locally as you choose.
The author is consciously trying to connect dots together via active searching through contents.
I can recommend him a puzzle as a pasttime (as he's bored after retiring in his 30s)
Agitation multiplied when I finally found the answer, which was reaching blindly under a plastic trim panel for a hidden downward-facing button to open the panel, lift the top up, and then balance the top midway between open and closed while I finesse the plastic tonneau cover panel back down under the rear part of the top. Garbage UX on both the part of Chevrolet and Google. I wish to become a hermit in the woods.
Author is on the right track but it's much more actively evil than this. The real algorithm is "what if we just show people the results we want to show them, because showing them those things makes us more money*?
* Except for services like Spotify, which shows you the songs that cost them the least.
That being said... the first wave of web search engines weren't built by the companies hosting the content. Search engines as tools became popular because, even in 1998, browsing and navigating sites (esp. corporate sites) to surface content, was already becoming an increasingly adversarial experience.
Search engines were built by companies spidering other sites' content. This was content that was often — if you were navigating along these sites' happy paths — found five to ten links deep through a confusing warren of subtle and unintuitive click targets. This content was the original "deep web." And search engines made it shallow... often against the spidered sites' consent.
Companies at the time much preferred "directory" sites that would only ever link to their landing pages. To companies of that era, sites linking directly to specific URLs of your site would be like if the Yellow Pages listed specific directory-extensions of your company's phone number! But this first battle in the "war against deep-linking" was a losing one from the moment it began, since early websites were almost inherently bot-accessible.
The second battle, in the early 2000s, where sites were built as [non-deep-linkable] Flash apps, took longer to determine, only finally resolving (again in favor of an open web, and thus third-party search) when laws and regulatory compliances both started to force companies back toward UA accessibility.
But for going on a decade now, we've been embroiled in a seemingly-indefinite third battle in the "war against deep-linking" — this time with companies tucking every possible type of user-generated content behind a login wall of some walled-garden everything [web]apps. Permalink URLs for individual content-items still exist in these webapps; but you get redirected to a login interstitial if you visit them.
Where's the adversarial spirit of the first search-engine companies today? Where are the companies trying to index the modern "deep web" of login-walled content? Why can't any company give me a search box that searches "into" YouTube video transcripts, "into" Facebook Marketplace, "into" public Slack and Discord and Telegram communities (probably requiring archiving, ala how Deja News/Google Groups surfaced Usenet), etc?
Sure, this sort of thing is intensely adversarial (but see https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn — while it might be against a site's ToS, but it's not against the law.) And sure, this sort of thing requires per-service connectors (but any thought of this business model therefore being "unscalable" comes from a pre-coding-agent era.)
But isn't it also, very clearly, "what people want"?