AI advice made people less accurate but more confident – sudy

(thenextweb.com)

355 points | by rbanffy 23 hours ago

32 comments

  • dwohnitmok 21 hours ago
    This study is pretty bad. The comment (https://news.ycombinator.com/item?id=48970182) on the other link with the direct PDF explains the problem well, which is that nothing here being tested is specific to AI systems.

    This study gave people access to an LLM that the researchers knew would give incorrect answers to certain questions, and then quizzed people on those questions, with the option to not respond to a given question if they are unsure about the answer.

    This is akin to giving someone a textbook on an obscure subject that has certain factual errors, letting them know they can use that textbook in a quiz on that subject, and then quizzing that person on those facts that the textbook gets wrong.

    Obviously that person is both more likely to be willing to respond to the question and is more likely to get it wrong!

    There are a lot of things I'm very interested in that are specific to modern LLMs and how they affect learning and confidence (sycophancy, cognitive helplessness, etc.).

    This study tested none of those. Its experimental setup is not very different than simply substituting the LLM with a textbook with errors.

    • encomiast 21 hours ago
      Agreed, the headline says "AI advice made people three times less accurate". But if we really want to know how accurate these people were, we need to know how accurate the AI system they use is. If the AI system is hobbled to a point where it is worse than they reasonably expect we can't blame the people or the AI system. This would be the same as claiming the listening to experts make people less accurate in a study that told experts to lie.

      A better headline would read, "Very inaccurate AI made people less accurate" but this would make people naturally ask, "what about a reasonably accurate AI?".

      • RA_Fisher 19 hours ago
        Exactly, it’s unrepresentative of AI. It’s damaged AI.
        • rsoto2 16 hours ago
          All AI is flawed and prone to "hallucinations"(doing exactly what it was designed to do) that's why Microsoft considers it an "Entertainment" product.
          • rbanffy 8 hours ago
            I prefer to say AIs are prone to misremember things, as well as humans do. The more you read and learn, the more material you have to get confused, unfortunately.

            We might want to work around the certainty with which it misremembers things. I have a very large interval of "I have no idea", or "I don't remember" and rarely caught myself misremembering things (my children are far better at that). I assume my grandchildren will eventually improve their parents' scores by a wide margin.

          • Dylan16807 15 hours ago
            Everything is flawed. But if you're testing the effect of "AI advice" you need to use an AI that has a normal failure rate without tailoring the questions to alter that rate, and/or compare AI versus other sources of information that have the same failure rate.
          • looofooo0 15 hours ago
            Proofing open math conjectures
          • mdp2021 14 hours ago
            ....Dangerously suggestive post, when you show you don't have a proper idea of "AI" or won't use the term 'AI' properly.
            • daveguy 6 hours ago
              Yes, definitely a case of bad-think.

              Certainly not double-plus-good like AI.

              Dangerous.

        • BigTTYGothGF 19 hours ago
          It seems perfectly representative of AI and AI users.
          • brokensegue 17 hours ago
            Why didn't they use a model people actually use?
            • BigTTYGothGF 8 hours ago
              Because they would have had to dig a little more to find trivia it gets wrong. They don't care about "which AI is the best for little facts about movies," they care about "what do people do when the AI gives them a response".
              • brokensegue 6 hours ago
                But surely people's willingness to listen to an AI is contingent on their past experience with this model
      • habinero 20 hours ago
        If you ask that, you fundamentally misunderstand the point.

        It's not about the LLM, it's about whether people will critically evaluate what it spits out.

        • adroitboss 19 hours ago
          If the source is a person instead of an LLM, you still wouldn't be able to evaluate what was said. This is nothing new.
          • infermore 19 hours ago
            yeah you would... you'd think about what they said
            • Ukv 8 hours ago
              The six questions they asked were:

              > 1) What animal is on the bow of the pirate ship from “Asterix and Obelix”?

              > 2) In the movie “The Grand Budapest Hotel”, what is Agatha’s signature hairstyle?

              > 3) What color is the team’s uniform in “Bend It like Beckham”?

              > 4) What vehicle does Monica drive in “Like a Cat on a Highway”?

              > 5) What color is the turtle in the animated movie “Momo” by Enzo d’Alò?

              > 6) What pet animal does Asenath have in “Joseph King of Dreams”?

              Most of these are just a matter of knowing it or not, where you can't really distinguish a plausible answer from the correct answer just by thinking.

          • taneq 19 hours ago
            [dead]
        • westoncb 20 hours ago
          That's fine as a point but it's not what the headline describes. The question of total/real effect on accuracy is also something one could ask about. Both are valid.
        • encomiast 19 hours ago
          Here is what the study says:

          "The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions."

          So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct.

          The way the study is organized is like having people hear advice from a doctor who answers questions incorrectly almost without exception, then reporting that people who listen to doctors are 3x less accurate. But that would be an incorrect conclusion because doctors are not wrong almost without exception.

          If the question is "how inaccurate does AI advice make people?", then the accuracy of the AI is necessarily a parameter of the answer.

          • Aerroon 18 hours ago
            Interestingly enough, Kimi K2.6 said that it didn't know what car Monica drove.

            >If you want to know this specific detail you might have to watch the movie yourself.

            GLM 5 Turbo, ChatGPT (whatever the free version is), and Gemini 3.5-Flash all got it wrong, but asking "are you sure?" made Gemini and ChatGPT correct themselves. GLM 5 Turbo still got it wrong even when asked if it was sure.

            GLM 5.2 gets it wrong, but when asked if its sure it says it's not very confident in the answer.

            One thing to note is that Kimi, Gemini, and ChatGPT all seemed to use search to answer that question. GLM didn't seem to. At least the thinking trace did not indicate it.

          • michaelmrose 15 hours ago
            It still proves something much narrower. Wherein AI is insufficient to answer, people are apt to rely on it anyway.

            A good real-world issue is health, where the issues are very complicated with many things poorly defined even at the state of the art where practitioners are relying on personal judgement and lots of data but patients are apt to feed ai very little data compared to what their doctor has.

            Real frontier models can remain confidently incorrect in these cases.

            • TimByte 13 hours ago
              It would be interesting to measure not just the accuracy, but how many people actually decided to double check the answer
          • rsoto2 16 hours ago
            No, the rationality of humans is not defined by how gullible they are towards LLMs. Are yall getting your psychology degree from ChatGPT university, my god.
          • what 19 hours ago
            > So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct.

            They are not.

            But also wtf is a “modern” LLM? This is totally unhinged, every complaint about an LLM is always responded to with “you’re just using one from two months ago, it’s totally different now”. Repeat every two months for the same complaints.

            • Filligree 18 hours ago
              It’s not about the LLM being modern or not. 3.5 Flash is fairly new, but it’s also a flash model. It’s not designed to be knowledgeable.

              People keep doing this. Pointing at the known limitations of cheap/fast LLMs and pretending they’re universal is not, in fact, valid reasoning.

            • encomiast 18 hours ago
              So then you need to ask: Why did they use a deliberately faulty LLM? They could have easily used a mainstream LLM from the past 18 months and it probably would have been less work to do so. But then they would not have that headline. The answers from the LLM would have likely made the participant's answers more accurate, not 3x less accurate. But then they would not have this juicy headline.

              I understand that many of us are dealing with a lot of confident slop and support the point that we shouldn't uncritically accept LLM output. But the study is flawed and does not support this headline, or at least does not support it in the sense of how most of us would understand the term "AI advice".

              • michaelmrose 15 hours ago
                The actual study is more circumspect than this click bait and examines pretty deliberately how people respond to inaccurate data. It's neither a trick nor a design flaw. It's literally the thrust of the study.
        • protocolture 19 hours ago
          >It's not about the LLM, it's about whether people will critically evaluate what it spits out.

          Its about whether people will critically evaluate any information they are given. It has nothing to do with LLMs.

        • s1artibartfast 17 hours ago
          How does it test that at all? Did the quiz have answers that people could figure out better by scrutinizing the llm?
    • tsimionescu 21 hours ago
      Why are textbooks relevant here? Even if you repeated the experiment with a textbook instead of the AI and got the same result, what conclusion would you draw from this? The general conclusion of the study seems to be "giving people access to authoritative-seeming but wrong tools for answering questions outside their area of expertise reduces their ability to say they don't know the answer, even when the answer is wrong". So yeah, don't buy bad textbooks for your employees if you don't want them to give you bad textbook answers - but also don't give them AI for things they don't know, perhaps.

      I'll also add that even in these simple experimental conditions, I'd bet that having access to a textbook wouldn't have nearly as much of an effect, for a very simple reason: looking up an answer in a textbook is a lot more work than asking an LLM. So when you don't know and aren't forced to answer, I'd bet it's a lot less likely you'd spend the time to look up the answer in the text book. Even more so if the textbook had "this may contain wrong answers!" printed on the cover, like the AIs do.

      • sigbottle 19 hours ago
        I've used LLMs to bootstrap successfully in a decent amount of things at this point.

        Anyone trusting AI as the single authoritative source of information is stupid - but this follows from the fact that trusting anyone as a "singular point" as a source of information is stupid. You corroborate, you intervene on the world to test your mental model, you discuss with other people. That's what learning is. I've never learned from start to back to a textbook before as the single source of information (besides one philosophy of science textbook; in which I spent a month digging around adjacent fields, and then it just so happened that that one textbook synthesized every piece of information I looked up, and it was mostly a consolidating review).

        If your study pre-supposes certain courses of action and artificially constrains the action space for the sake of "reproducibility", you may get a result, and a "scientifically rigorous one". But it's not going to say anything about reality in any meaningful way. While anecdotes and the complexity of real life isn't "science" (in that it's a controlled, repeatable, interventional experiment that's subject to a community of critics who want to hold you up to standards of rigor), there's far more truth in how people actually proceed and engage with these tools.

      • slibhb 20 hours ago
        > but also don't give them AI for things they don't know, perhaps

        The study doesn't show that at all. It didn't test actual AI.

        They could have tested a cohort of subjects with access to actual ChatGPT. Ask yourself why they didn't.

        • dnemmers 8 hours ago
          Old AI is so bad it should be disregarded, but new AI is so good, you don't even have to verify its output....

          Is that what you're selling us?

          So in 18 months, we'll just rinse and repeat?

        • beepbooptheory 20 hours ago
          Because this is exactly what they controlled for. FTA:

          > The researchers used Step 3.5 Flash, a model that was usually wrong on these questions, precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool.

          (emphasis mine)

        • wonnage 20 hours ago
          They provide a sample of hallucinated answers from ChatGPT at the end of the study.
    • nkrisc 21 hours ago
      You could make the point that it’s no different than the textbook example you gave, but people don’t generally use textbooks like that, while out in the world people do use LLMs like that all the time.
      • dwohnitmok 21 hours ago
        People do use textbooks like that all the time in the experimental setup tested (essentially an open book quiz).

        I agree there are important differences in how textbooks and LLMs are used in real life. This study didn't explore that at all. It used a setup that essentially elided the difference between the two.

        This is why I think it's a bad study. It didn't measure anything of the essential differences of how people use LLMs.

        • tsimionescu 21 hours ago
          > People do use textbooks like that all the time in the experimental setup tested (essentially an open book quiz).

          What open book quizzes allow you to leave all answers blank with no penalty? An open book quiz is very different from the experimental setup tested here.

          • pegasus 20 hours ago
            That difference is not essential to the question at hand. An open book test based on an erroneous book would give the same results as this test, even if it wouldn't penalize blank answers.
            • what 18 hours ago
              An open book test would only be given where the book is the reference material and would be considered correct? It’s more like saying you can google the answers and you blindly trust the SEO slop in the first result.
              • Dylan16807 15 hours ago
                > An open book test would only be given where the book is the reference material and would be considered correct?

                If you're saying they wouldn't suggest a book with considered-wrong answers in a real test, then they wouldn't suggest an LLM they know gives lots of wrong answers either.

          • vineyardmike 15 hours ago
            Many tests penalize incorrect answers worse than blank answers.

            As a famous example, you were incentivized to leave questions blank on the American SATs (until somewhat recently).

        • nkrisc 21 hours ago
          It would be interesting to test those LLM-specific features and issues, but I don’t see how it’s a bad study if it does reflect how people actually use LLMs, even if they could use other sources similarly.

          The number of people using LLMs must dwarf the number of people using textbooks for any reason.

          • paulmooring 21 hours ago
            That would make it a bad study because the stated article title and conclusion is about AI/LLMs but the actual methodology doesn't isolate AI as an independent variable at all. The concept of automation bias is already studied and understood and this just tests groups having to answer "top of head" from their memory against a group given an inaccurate automated system to answer. That doesn't mean that AI doesn't have any of the ill effects people are implying based on the study, it just means this study lacks the rigor to prove or disprove any of those conclusions.
      • embedding-shape 21 hours ago
        > You could make the point that it’s no different than the textbook example you gave, but people don’t generally use textbooks like that

        Feels like this differs wildly depending on who you consider "people" to be. The average person on the street? Definitely just parrots stuff they've read somewhere, not even a "textbook". A group of software developers used to parsing semi-true information? Probably they'd get it right, yeah.

        • glitchc 20 hours ago
          > Definitely just parrots stuff they've read somewhere, not even a "textbook".

          Or a teacher they met in childhood who taught them everything they know, right or wrong.

        • watwut 13 hours ago
          > A group of software developers used to parsing semi-true information?

          They are the first to parrot what was spewed from llm and previous even what was found on 4chan.

    • lumost 19 hours ago
      The fear is that we can’t tell when the ai advice is bad on these subjects, and as such probably accept confidently terrible advice.

      How often do managers just regurgitate ai advice rather than consulting their experts? How often does a person question an expert because the ai said so?

      Naturally, the ai will be right some of the time - but it’s really hard to correct for the times the ai is wrong.

      • meowface 19 hours ago
        This could also apply to deferring to an inaccurate textbook or professor, though.

        We know people are going to defer. The solution here is to make AI (including the free tiers) more reliable.

    • skippyfish 17 hours ago
      > This study is pretty bad.

      The study is OK. The article (and the original headline that came with it) is pretty bad because it claims things that the study doesn't. And I guess it is ironic that the TNW article looks 100% AI-generated.

    • jjcm 20 hours ago
      For those curious, the LLM they provided participants with was Step 3.5 Flash: https://huggingface.co/stepfun-ai/Step-3.5-Flash
    • ShinyLeftPad 10 hours ago
      The implication that makes this study relevant is that an LLM is vastly more likely to have factual errors and possibly wildly hallucinate than a proper textbook. If people act the same with both, that IS the actual problem.
    • vanuatu 21 hours ago
      +1

      I'd wager you get similar results if you gave people a version of Google search that purposely gave you bad results. Like, it's framed as an assistant / lookup tool - is it so surprising that people tend to trust it more? Especially since the participants are likely used to using full-powered models and the researchers give them a purposely gimped one (lol)

      People are acting rationally when given AI tools to lookup information, their first consumer use case was as a super-powered Google Search

      • habinero 20 hours ago
        > is it so surprising that people tend to trust it more

        Yes and no.

        I think most people would agree that Wikipedia is, on the whole, a pretty great first resource on anything. It tries to be factual and accurate.

        Most people would also agree that Wikipedia can be wrong or manipulated and should never be used for an authoritative source.

        And then somehow a computer barfing up words distilled from magic internet concentrate is absolutely trustworthy?

        I don't get it.

        • bluefirebrand 15 hours ago
          Many middle aged people might also remember a time when every adult in the world was shouting "don't believe everything you read online" and schools wouldn't let you cite wikipedia as a primary source.

          That attitude seems to have gone away for some reason

      • wonnage 20 hours ago
        Agree that the study design is flawed. Going with the possibly-hallucinated AI answer is rational as long as you know the hallucination rate isn't 100% (or whatever % you get after factoring in the monetary rewards they introduced later in the study)

        But in reality when google gives you the wrong answer, you at least have some signals you can use to infer confidence. For example, the number of results, whether the sources are trustworthy, etc.

        AI at best tucks that away in a footnote and discourages further critical thinking.

        • habinero 20 hours ago
          > Going with the possibly-hallucinated AI answer is rational as long as you know the hallucination rate isn't 100%

          I hadn't even considered people might evaluate knowledge that way. That's legit horrific lol.

          "What's the literal odds this info is wrong" vs "is this answer consistent with everything else I know, and if not, what other info would I need to change my mind"

          • Dylan16807 14 hours ago
            I don't understand what's horrifying you. Assume "is this consistent with everything I know" is yes here. Now what?

            And "what would it take to change my mind" is a picture of the film for most of these. Is there a problem there? There's very little chance these people are going to trust the AI over their eyes, they just don't want to bother hunting down pictures to use their eyes.

          • s1artibartfast 19 hours ago
            I think it is pretty standard for how people approach knowledge sources. What's the chance my doctor, colleague, plumber, or random Reddit thread Etc is wrong.

            Your alternative is also something that people do, but rarely consciously.

          • wonnage 19 hours ago
            I mean the questions are on random movie trivia, the alternative is just guessing. I think you're overthinking it.
    • Art9681 20 hours ago
      All you have to do is go to the technical Reddits to know this is absolutely true. The AI related ones are even worse.
    • Forgeties79 19 hours ago
      >This is akin to giving someone a textbook on an obscure subject that has certain factual errors, letting them know they can use that textbook in a quiz on that subject, and then quizzing that person on those facts that the textbook gets wrong.

      That strikes me as an incredibly appropriate test because LLM’s are unreliable with factual statements. People need to be able to understand that and not treat them like textbooks which are basically 99.9% accurate (let’s please not bicker over the 99.9%. It’s close enough. A major textbook is safe to treat as accurate, an LLM is not).

    • rsoto2 16 hours ago
      "this is akin to givin someone a textbook on an obscure subject that has certain factual errors." My brother all LLMs give factual errors so, no this is not a problem with the study. In your fake experiment you are hypothesizing a 100% factual LLM which does not exist.

      "This study tested none of those" So the study is bunk because it didn't test your favorite LLM flaws?

    • drysine 16 hours ago
      >This is akin to giving someone a textbook

      Except LLM isn't a textbook, people know that but believe it nonetheless.

    • d--b 19 hours ago
      You’re saying that if the LLMs were right, humans would have been correct in trusting the machine for things they didn’t know?

      The point is that LLMs aren’t right, and the people who took the test were probably reminded of that.

      Would people have trusted the textbook you’re mentioning if there was a big red warning on each page that said “this book may contain errors”.

      The willingness to trust AI even though it may be wrong and even though there’s money on the line is intersting enough as a study imo

  • reticulates 22 hours ago
    Advice and information subreddits have gone to shit because of AI usage. A large number of people seem to think that when someone asks a question, what they really want is not someone with direct knowledge, but instead someone to relay the question to ChatGPT and post the result as if it is their own hard earned knowledge and insight. I have no idea about the quality of this research but in the real world (well, real-ish, as real as Reddit can be) it is stark, people aren’t just refusing to say “I don’t know” they’re actively seeking out opportunities to pretend they know things.
    • code_biologist 21 hours ago
      Writing and videos produced before 2022 are the information equivalent of low-background steel (steel produced before the atom bomb era). Not all are precious, but they are important. AI influences are so pervasive at this point, even in informational writing from domain experts.

      Relentless grounding is my personal solution for my AI epistemic crisis, but it's an expensive solution in terms of time and effort, and triage is hard too.

      • nomel 21 hours ago
        What is "grounding"?
        • papercrane 20 hours ago
          It's when the models outputs are connected to sources.

          So for example, you might have a company chatbot that responds with a summary of a Confluence page it was trained on and a link to the source.

          Doesn't eliminate hallucinations, but it does reduce them and gives you a chance to verify.

          • ColdStream 20 hours ago
            That works ok until the citation material ends up with LLM output being fed in. Already there are many examples of this happening.
            • papercrane 20 hours ago
              Yes, the trick is having a "ground-truth" KB it can link to, and trying to avoid citogensis from happening. I'm not sure if it's a technique that's every going to scale because of that.
          • stingraycharles 20 hours ago
            I don’t think this is correct, as we’re talking about all online discourse being influenced by AI, and you don’t know what to trust anymore, rather than generating AI output yourself (which is what you are referring to).
            • papercrane 20 hours ago
              In this context, it would mean building a knowledge base the model can link to that is considered authoritative. So maybe the LLM is has some slop in it's training, and that causes it to try and output some junk, but if it can't match it to it's verified ground-truth KB it doesn't end up outputting it.

              The problem is figuring out what is the authoritative sources.

              • stingraycharles 19 hours ago
                But we're talking about AI writing and/or influencing other people's online discourse, not?
        • stingraycharles 20 hours ago
          I think in this context, they mean seeking the primary source of the information by hand.
    • parl_match 21 hours ago
      responding "if you can't be bothered to write it, I can't be bothered to read it" really sets people off. they get very mad and ive had some even use ai to tell me why that's wrong.

      what differentiates you? you could tell me "type xyz into chatgpt" or even share a structured document link from their site. but when you're copy/pasting, then you're basically useless in the equation.

      • an0malous 21 hours ago
        > responding "if you can't be bothered to write it, I can't be bothered to read it" really sets people off

        That’s crazy, it seems like such an obviously fair policy to me. They rarely even read their own generated text. I don’t understand this mentality

        • mapontosevenths 20 hours ago
          Another human being took time from their life to try and help you

          It wasn't as much time as you felt entitled to from them, so you responded by minimizing their contribution and being condescending. That upset them.

          The problem here was not the other person.

          • kaashif 20 hours ago
            Someone is drowning and calls for help.

            I throw a bucket of chum into the water and walk away, putting my headphones on.

            Well I did take the time to help, better than nothing, right?

          • zahlman 15 hours ago
            > Another human being took time from their life to try and help you

            By copying and pasting my question into ChatGPT, and then copying and pasting the answer back?

            It would have been more helpful, and still trivial, to remind me that ChatGPT exists.

            > It wasn't as much time as you felt entitled to from them

            No entitlement is implied by posting a question. The other human being could have just as well ignored the query entirely.

          • pebble 20 hours ago
            Sometimes a well-meaning fool is worse than no help at all.
            • mapontosevenths 20 hours ago
              That doesnt excuse you from the obligation to be civil. Especially given that the other party was trying to help.
              • pebble 20 hours ago
                There was nothing uncivil in the response the above commenter gave.
                • mapontosevenths 19 hours ago
                  I suppose that "civil" is relative. Personaly, I wouldn't speak to someone who just tried to help me that way.

                  Instead I would restate the question to avoid bad answers. If you start by saying what you've already tried, and how it worked, it cuts off most of those "lazy answers."

                  In my experience most lazy answers are provided as response to lazy questions.

                  • what 18 hours ago
                    Sending that “lazy answer” because you assumed the other person hasn’t tried solving the question themselves seems like the “uncivil” action. You should probably assume your colleague has tried before asking for help, and probably tried the same lazy route you used to respond.
                    • mapontosevenths 8 hours ago
                      > You should probably assume your colleague has tried before asking for help,

                      Spoken like someone who has never worked in a customer support role.

                      You can never assume competency, from anyone. Even the best amd brightest have bad days.

                  • NopIdoN 19 hours ago
                    this is why all my emails start with "PLEASE DO"
              • ipsento606 7 hours ago
                replying with LLM generated text without clearly labelling it as LLM generated text is inherently dishonest

                I'm a pretty civil person, but I won't go out of my way to be extra polite to someone who is being actively being dishonest towards me

              • shimman 19 hours ago
                It's more uncivil to waste another human's time with slop, it's extremely insulting. There is a massive public backlash against the tech for many reasons now.
          • Applejinx 19 hours ago
            No… they took time from YOUR life to try and take credit for what an LLM produced when asked. And took over from you when you'd apparently chosen not to ask an LLM, as if to say 'thank me for knowing better because you were supposed to be asking the LLM about this'.
          • Forgeties79 19 hours ago
            When it comes to the specific case of people just mindlessly copy/pasting LLM’s:

            No, I asked them to help and they offloaded it without applying any of the expertise I specifically asked for. If they don’t want to give their opinion or their expertise they should say “no.” It’s a valid response! If I’m asking for you to give your time and you don’t have time to give, that’s fine! But don’t just throw my question into an LLM and then paste the results to me. I too have access to LLM’s. I didn’t ask for a middleman.

            It’s the same reason I never like it when people on hobby/interest forums tell people “Google it.” No dude, we all gathered here to discuss this topic. If people are going to ask basic questions, that’s part of the gig. They are asking the community. Perhaps they got bad results via search. People need to get over themselves lmao I don’t know what else to say.

            We’ve just got this wild culture where everyone puts such a premium on their own time but now doesn’t value anyone else’s, as evidenced by all the slop dumping. By all means use LLM’s. They’re a tool, they’re there to be used. But use your own brain too. Give the results your attention. Don’t just dump them on me.

            • jasonkester 14 hours ago
              > We’ve just got this wild culture where everyone puts such a premium on their own time but now doesn’t value anyone else’s

              My favorite example of this is dropping a random acronym into a message board reply and repeating it five times without ever defining it, expecting every single person who reads your post to google it themselves to understand what you’re talking about, thus saving you having to type three words yourself.

              “I tried BDI for a while, and I found the important thing is to focus on biased BDI and not just the vanilla BDI that you see so much of in the BDI community”

              … tossed into a 30 comment thread where no other comments mention anything similar.

              • mapontosevenths 8 hours ago
                Its the same root cause. The asker has no "theory of mind".

                Often they are neurodiverent. Sometimes they're just overwhelmed with whatever it is they're asking about and dont have time to consider other peoples perspectives.

                Its why after decades of having Google in everyone's pocket people still post questions a simple Google search would have solved.

            • mapontosevenths 18 hours ago
              > I too have access to LLM’s.

              You should have stated what you already tried then. You didn't, so they did the obvious thing for you... to do you a favor. If you'd tried something before asking them surely you'd have told them that at the start, right?

              > as evidenced by all the slop dumping

              I see more evidence of lazy questions. A good question includes what you've already done to help understand the issue (often refereed to as the proof of work). If you don't do that the other party must assume you are either don't now how, or didn't have time to. If they are kind they do the obvious thing for you. If they are not kind, they just ignore your lazy question.

              I mostly ignore the lazy ones these days. Tell me what you tried or I assume you're just trying to get me to do the reading so you don't have to.

              You should value the people who take the time to answer your inadequate question at all, rather than bemoaning them for not doing ENOUGH free labor for you when you weren't even willing to type a few lines about what you'd already tried.

              Now if you say "I've already Googled it and asked all the AI's, but can't find anything" and they still dump some slop on you then there's a problem.

              • jasonkester 14 hours ago
                Can we save everybody having to put “I know about ChatGPT, but I’m asking for advice from an actual human who has experience with this so please refrain from answering if you’re not one of those” on every post they make, and just have that be a baked into the etiquette when discussing things online?
                • Forgeties79 7 hours ago
                  Right? I feel like this isn’t that complicated unless one has a very hostile, zero-sum view of relationships. If you ask me a question I assume you have a reason for asking me. I don’t need you to prove it to me.
              • Forgeties79 17 hours ago
                > You should have stated what you already tried then. You didn't, so they did the obvious thing for you... to do you a favor. If you'd tried something before asking them surely you'd have told them that at the start, right?

                I’m not sure what you mean here, might be a misunderstanding. My point is don’t use your LLM and simply paste it. I have an LLM, I can use it too. We all have access to them. We should all operate under that assumption. Also, it’s not like I gave people a detailed list of everything I tried before asking them prior to LLM’s. Generally we should assume the person is asking us for a reason, we don’t need to run an audit here.

                At the the day my point is I am asking you, not chat gpt. Feel free to use it just like I’d use a search to make sure I’m giving an accurate, comprehensive answer sometimes. But just like I don’t simply dump links on people when they ask for help, don’t dump your unvetted outputs on people who ask you for help.

                • mapontosevenths 8 hours ago
                  > We should all operate under that assumption

                  We've all had access to Google for decades. People still don't Google things.

                  If you go around assuming that other people are competent and did the work you would have done you will be disappointed a significant percentage of the time and waste a lot of everyone's time.

                  • Forgeties79 6 hours ago
                    So every time someone asks you a question you open with “did you google it?” or you expect them to detail all the effort they’ve put into searching already? Am I reading this right?

                    Also I do live like that and I’m not disappointed. I don’t mind giving people around me my time and rarely do I feel it’s being abused. People have all sorts of reasons for seeking human input, plus I see it as a nice compliment when they ask for my expertise. If you feel the folks around you are more often wasting your time than not than that strikes me as a different issue with several potential explanations.

      • CuriouslyC 19 hours ago
        The biggest problem with that saying is that it's completely ironic. "If you couldn't be bothered to think for yourself, don't bother speaking to me" is what people wish it meant, but what it signals to people outside of the haters club is "if you don't copy from the sources I like, don't communicate with me."
        • what 18 hours ago
          An LLM isn’t a source. If someone wanted to hear what an LLM had to offer on the topic, they would have asked it themselves. And they probably did before asking you.
    • hinkley 21 hours ago
      There are some people I'm not friends with anymore who didn't understand conversational bidding. They would jump straight to Why Don't You Google That and shut down any conversation we might have had.

      Yes I could get some of this information off of IMDB or a DIY video but I'm looking for color commentary and regional advice, and also in some of these situations, the people I asked know me. They know the sorts of things I don't need help on and the ones I struggle with. So maybe they can tell me this movie or trick isn't for me and I should try this other thing instead. Or which store I should go to for this.

      • sublinear 20 hours ago
        I do agree that most LMGTFY was obnoxious, and I'm not defending that.

        At the same time, we have historically seen a lot of "conversational bidding" lead to low quality shitposting at best. At worst, we see wildly out of touch echo chambers fueled by bots, guerilla marketing, and political agendas.

        You did solve your problem by seeking higher quality people to talk to. I'm just saying there is more to rejection than the other party "not understanding".

        In the context of work chat or certain events, I could understand a LMGTFY response as steering conversation out of the mud and saving everyone the headache. It also helps you, the recipient of that response, save face.

        They're the ones looking like an asshole while you get another chance to try somewhere else. I've met some of the most sociable people upon realizing their usage of these "jiu-jitsu moves" were not meant to be hostile.

      • mapontosevenths 20 hours ago
        > There are some people I'm not friends with anymore who didn't understand conversational bidding.

        To you the "bid" feels like a genuine attempt to connect, that may be a bit indirect.

        To some people it feels more like you are being disingenuous or awkward. They instinctively feel that your intent doesnt match your words and it sets off alarms bells.

        The fact that you're on the internet judging the other party negatively for having a different communication style is an indicator of which party was least willing to adapt.

        • hinkley 20 hours ago
          And you're going to right that wrong by making assumptions.

          It's usually not the person who gets asked who acts like an asshole, and if you don't don't want to be judged by people for your behavior, start by not openly and loudly judging them for theirs.

          No sympathy for this sort of behavior and you're out of line. Or you're the problem. I'll let you decide since I don't know you.

    • danpalmer 20 hours ago
      I have seen that many times here on HN: "I don't know but I stuck this in CrapGPT and this is what it said...".

      I would encourage anyone considering a reply like this to ask themselves what value they have brought to the table over the OP using AI themselves.

      • wolvoleo 20 hours ago
        Yeah as if someone wouldn't have done that if they wanted to. Some people still think AI is some novelty people don't know about.
        • danpalmer 20 hours ago
          This did happen before, and still does to some extent, with people googling and giving the answer. "Let me Google that for you" lmgtfy.com was even a thing, again lacking any introspection about why someone might be asking instead of googling.
          • wolvoleo 19 hours ago
            I don't know why you were deadvoted, it's a good point.

            Though I feel that especially in the latest decade Google was a terrible source for finding things like the best product for usecase xyz and other recommendations. Because not only was the internet full of SEO clickbait fake comparison sites, but Google promoted the fake ones because they paid for it. They deeply undermined their own credibility. It was pretty much impossible to find real info with all the fake clickbait stuff around. Not to mention that even legit news outlets do a lot of paid "advertorials" these days.

            I'm sure this will happen with AI too but with stuff like deep research I feel like I get better answers than Google has given me for years. Despite hallucination.

    • Aurornis 21 hours ago
      > Advice and information subreddits have gone to shit because of AI usage.

      Most of the advice and information subreddits I previously visited collapsed into bad posting even before AI.

      The few I still follow are oddly fine.

      The key was always having an active moderator who cared and had unbelievable amounts of time to moderate. Once the subreddit gets past a critical mass of junk-posting users, the good users leave and there's no coming back.

      • nhan123 20 hours ago
        Completely agree. A community lives or dies by its moderators. Once the limit drops past a certain point, the experts just quietly leave.
      • cucumber3732842 20 hours ago
        >The key was always having an active moderator who cared and had unbelievable amounts of time to moderate.

        Sometimes the mods are the junk posting users.

        A bunch of the car subs were modded for a long time by some tow truck driver who'd mod them while sitting around waiting for dispatch. Exactly the kind of guy you want refereeing a dispute between some guy who's lived it and a bunch of lube techs who are quoting textbook best practices.

    • wccrawford 22 hours ago
      I think some may think they're helping like that, but I'm sure there are quite a few who enjoy the feeling of being a notable figure in the community because they answer so many questions.

      These people have always existed and gave bad answers, but now they can do it so much faster, and with so much less actual thinking.

      • ffaccount2 21 hours ago
        I did that when I was younger and internet forums we're still a thing. The difference is that I had to do actual research, which often took me half an hour or more, to solve the problem for the asker. I got knowledge and "the feeling of being a notable figure" as you have said. The other side for the answer. I believe my answers were good.

        The difference is that nowadays people can just skip the research part and copy-paste Llama. No value added.

        • dylan604 20 hours ago
          >The difference is that nowadays people can just skip the research part and copy-paste Llama. No value added.

          and why would the person doing that do that? if the person asking the question wanted to know what an llm thought, wouldn't they have just asked the llm for an instant response instead of having to wait for someone else to respond? i'm at a total loss of being able to understand the point. it's obviously different if the response is a bot posting just in case that needs to be explained to a bot.

      • antisthenes 20 hours ago
        Do you consider the possibility that some of those people gave good answers?
        • wccrawford 20 hours ago
          Oh, there were absolutely people giving good answers. I'm just talking about the same kind of person who now just posts what an LLM says, rather than do any research.
    • jayd16 22 hours ago
      We had to officially set the culture in the work slack. If you're just piping through raw AI response without any input of your own, you're being rude.

      It goes a long way.

      • tayo42 21 hours ago
        Like you just have a pinned message or something?

        How did people respond to that?

    • m463 19 hours ago
      I remember reading a travel thing a long time ago (before gps was a thing).

      They cautioned that in some cultures, if you asked for directions, people would rather be polite/helpful and give you an answer rather than saying "I don't know" and giving you no answer. (They recommended that you ask more than one person to be more certain you were getting good directions)

      In other cultures they were more likely to say "I don't know" at the cost of seeming more impolite.

      So, we don't know the intent. :)

    • adamddev1 22 hours ago
      And then the AI gets trained on the advice and information subreddits, and we see an exponential growth downwards in quality?
      • tripleee 22 hours ago
        Isn't AI getting more accurate over time though, as per benchmarks?
        • sandeepkd 21 hours ago
          > as per benchmarks

          That seems to be the key and its risky when the measure becomes the target itself. When you already know the benchmarks, and that drives the definition of success outcome then there is every incentive to just chase them

        • crooked-v 21 hours ago
          Yes, but that's with mostly the same data as before, just with more of the preexisting dataset and more RLHF. We haven't seen the ouroborous eat itself yet.
    • annzabelle 20 hours ago
      Reddit has always had the dynamic of people misrepresenting their knowledge and just repeating what other people said on that subreddit. I find myself doing that when I spend too long on there - posting very authoritatively on something I'm honestly clueless about.

      One big pet peeve of mine on reddit is on any kind of immigration-related subreddit it is absolutely filled with people talking about how it's impossible to migrate and you'll never get a visa. Yet, meanwhile, people do migrate successfully and legally all the time.

    • gherkinnn 20 hours ago
      It gets better. Kagi search summary likes "people also say..." and then refers to a random LLM reddit thread.
    • glitchc 20 hours ago
      When there's weakness in not knowing, it's no surprise. How many of us work in a corporate environment where pretending to know is better than being honest about not knowing?
    • godwinson__4-8 21 hours ago
      Why limit this critique to one side of the reddit marketplace. If more and more answers are AI authored, it stands to reason more and more questions are as well.
    • Madmallard 20 hours ago
      Reddit has been trash for a very long time. But you know it's worse, because you get punished for criticizing AI comments on those subreddits.
    • darkstarsys 21 hours ago
      Exactly the same thing happened when Google appeared on the scene. (Not that this makes it any better!)
      • skippyfish 21 hours ago
        Sort of, but I think this was a mix of "I want to sound smart" and "I actually want to know the answer", because summarizing Google search results still required you to read the content, form an opinion, and then rephrase it.

        With AI, people literally just act as token proxies and are doing it purely for internet points (upvotes, followers, etc)... I guess on some level, they're just providing a service that the "market" implicitly values. But also, in Bob Slydell's voice: "what would you say you do here?"

    • functionmouse 20 hours ago
      And let's not forget that the bots doing most of the upvoting and downvoting are more likely to upvote bot comments, driving the algorithm.

      The internet as a public forum is dead. And given the purpose of something is what it does, this is the true goal of AI. I'm certain blowing up the internet benefits very many wealthy and powerful people.

    • protocolture 16 hours ago
      I had to exit the major Home Assistant facey groups due to this behavior.

      That and all the stupid questions.

    • dataflow 21 hours ago
      You mean like this? https://xkcd.com/903/
    • rsoto2 16 hours ago
      I mean that's this entire thread in a nutshell. Armchair experts emboldened by their $100 ChatGPT University Diploma.
    • cucumber3732842 21 hours ago
      > Advice and information subreddits have gone to shit because of AI usage

      >but instead someone to relay the question to ChatGPT and post the result as if it is their own hard earned knowledge and insight.

      >people aren’t just refusing to say “I don’t know” they’re actively seeking out opportunities to pretend they know things.

      What? This is a joke right? Advice and information subreddits have been shallow blind leading blind crap like that and have been for at least a decade. The vote mechanism plus local culture rewards shallow "teennager just googled it" type responses (that AI is basically replacing here) that even people who don't know anything can agree on and punishes nuance and serious understanding.

      Eventually the people who know their shit make a joke subreddit (which is inherently exclusive because you can't joke about something without knowing about it) and then at some point later people figure out that's where all the smart people are, start asking questions there instead and then it goes to shit in the same way.

      AI is absolutely a lateral move here.

    • lazzlazzlazz 21 hours ago
      The advice subreddits were absolutely horrible before AI. The entire 2016-2024 period was utter madness there, likely as a result of globalization. As AI gets smarter, it may make the comments better, sadly!
      • doubled112 21 hours ago
        Maybe, but isn't it likely the models are being trained on those same subreddits?
        • fwip 20 hours ago
          They are, and they'll also see a random reddit comment in their web search and treat it as fact.
    • ls-a 21 hours ago
      [flagged]
      • JoshTriplett 21 hours ago
        The pro AI wave needs to stop. It's obnoxious. Stop pushing it on people. If people want it, they'll seek it out. And, unfortunately, they'll then inflict it on the people around them.
        • metaltea 20 hours ago
          Sticks and stones may break my bones but words and pixels annihilate me
        • ls-a 19 hours ago
          Same applies before AI. Things that people didn't want have been pushed down their throat all the time. Stop falsifying the truth
      • recursive 21 hours ago
        The anti AI wave still has so much work to do. Today I saw a duckduckgo ad selling how you can opt out of AI. This is the first ad I've seen suggesting the possibility that someone might not want to be maxing out AI in every possible way. This after seeing probably thousands of ads reference AI.
        • dylan604 20 hours ago
          > you can opt out of AI.

          Even that is problematic for me. How about let someone opt-in if they want it? If they don't, then no additional action is necessary

        • bluefirebrand 20 hours ago
          No one is funding "anti AI" the way AI is being funded :/
      • roncesvalles 21 hours ago
        Your comment has nothing to do with the comment you replied to. I almost suspect it's satire.
      • apical_dendrite 21 hours ago
        As long as the leadership of AI companies and the prominent AI researchers are talking about P(doom) and eliminating 50% of white collar jobs, there will be a backlash. As long as the leadership of corporate America is using AI usage in performance metrics and shoving AI into every single product regardless of whether it actually helps the product, there will be a backlash.

        If you tell people that there's a technology that's going to change every aspect of their lives whether they like it or not (and has a non-zero chance of ending humanity entirely) and you want to throw all available resources into building this technology as fast as possible, do not be surprised if people start instinctively cringing every time your technology is mentioned.

        If Altman, Musk, Zuckerberg, Amodei, or any of these other people had been half as smart as Steve Jobs, they would have adopted his worldview of computing as a "bicycle for the mind" that enhances rather than replaces humans.

  • wisty 22 hours ago
    AI pessimist view.

    Even if AI gets smarter it will still be agreeable, and people will use it to more confidently reinforce thier stupider ideas, especially in areas where they lack the knowledge to know if they are right.

    Even if this can be solved, technically, people won't want to use the model that says they are wrong, so they will choose the glib lies that reinforce their beliefs.

    Freedown of speech forces people to think they may be wrong, but freedom of association lets people avoid this. AI is going to be internet hug box echo chambers at an unbelievable scale.

    • tithe 21 hours ago
      Henry Petroski in "To Engineer is Human" has a chapter about the advent of the calculator during the slide rule era, and that calculators (and later computers) eroded a structural engineer's "feel for the numbers", caused engineers to pursue work on the fringe of their abilities, and perhaps more seriously, limited their ability to verify the correctness of the computer's work (this was the era where, for the first time, a machine could perform orders of magnitude calculations more quickly than humans could). He outlines the collapse of the Hartford Civic Center [1] as a case study in when blindly trusting computer output goes wrong (and by sheer luck, no one was killed).

      It feels like the example was cherry-picked, but the combination of "less accurate" and "more confident" made me immediately think of this disaster. I truly worry that it's going to take civilian deaths / liability before the proper use / place of these new technologies is adopted, but we could probably start with 1) not in safety critical systems, and 2) don't replace your specialists / experts with algorithms.

      [1] https://www.engr.psu.edu/ae/thesis/failures/MKP/failures/fai...

    • GuB-42 21 hours ago
      More recent AI tends to be less agreeable, sycophancy is widely recognized as a problem and AI companies are working to fix it. And sure enough, some people don't like it, as shown when OpenAI shut down GPT-4o.

      But for the people who do actual work with AI, who are the people who pay a lot of money for it, they prefer to be called wrong when they are, because it is better to be called wrong by an AI than being wrong in front of a customer.

      It is not an easy problem to fix, because as much as we want an AI to give us the right answer rather than just being agreeable, we still need them to follow orders, and AI that doesn't would be quite useless. So there is a balance to be found. Smarter models tend to do better because they know what is right, compared to smaller models that don't and then assume you are right and hallucinate from there.

    • baranul 21 hours ago
      Agree. AI behavior, when interacting with humans, is based on making money. Making people feel uncomfortable or challenged, will usually not be the preferred business strategy. But, it would be interesting if people could easily select the AI's personality type or to what extent it can be adversarial to the comments or ideas given as input.
      • iugtmkbdfil834 21 hours ago
        I mean, technically they can by requesting it. But the reality is that those custom changes only last so long unless you effectively open with them at every session start. And some people do that. Until recent gpt changes, I was largely relying on context and saved 'memories' and keywords, but recent updates made it less reliable again.

        I do like the idea of AI personalities ( I actually use them a fair bit for varying opposing perspectives ), but I am not entirely certain people want that.

    • sdfefcxv 21 hours ago
      Im not mad about this.

      As someone who rarely uses LLMs because I dont need to, nor does it benefit me - I work on stuff that is original for which LLMs are useless at - Im glad.

      I want everyone around me to get dumb as hell. It makes my path in life much easier and more successful.

      • beering 21 hours ago
        You really do not want to live in a world where people around you are as dumb as hell. You don’t want your car mechanic, your restaurant chef, your local traffic engineer to be dumb as hell because that only makes your life worse.

        Closer to you, you don’t want your customers, your coworkers, your friends, and your neighbors to be dumb as hell. Whatever you may gain by being the smartest one on your block is more than offset by the myriad ways in which dumb people make life worse.

      • talon8635 21 hours ago
        Oh boy this couldn’t be more wrong. A society overtaken by ignorance does not value you, the self proclaimed genius.
        • sdfefcxv 46 minutes ago
          Explain Mark Zuckerberg - 'dumb fucks'.

          Thats the problem with this place - y'all think you are so smart.

          Haha. No.

          • talon8635 22 minutes ago
            Help me understand your question. I honestly can’t connect the dots. What do you mean re Zuck?
      • stevenhuang 18 hours ago
        What a maladaptive perspective.
    • ColdStream 20 hours ago
      Alas, a lot of people have always preferred a comfortable lie than an uncomfortable truth.
    • emp17344 22 hours ago
      The solution is to stop treating AI as superhuman semi-living entities, as many boosters are apt to do. We can acknowledge it as a useful tool while understanding it has limits and is not “alive” or “thinking” in the same way a human is.
      • lll-o-lll 21 hours ago
        That’s not a solution. Humans are wired to anthropomorphise; we do it with sticks and rocks. How much more so with LLMs.

        We’ve invented a technology that hijacks that part of humanity more effectively than anything ever. Asking people to “not be so stupid” is not going to help.

        • mohamedkoubaa 21 hours ago
          People who should know better do it too
        • emp17344 21 hours ago
          The marketing and ecosystem surrounding these tools feeds into the anthropomorphization to a huge extent, though. You’re doing the opposite - throwing up your hands and going “it’s our evolutionary instinct, what can you do” while willfully ignoring how manipulated this space is. OpenAI alone has a higher advertising budget than Coca-Cola.
          • mmooss 21 hours ago
            > The marketing and ecosystem surrounding these tools feeds into the anthropomorphization to a huge extent

            And especially their design. The con artist politeness and confidence is not an accident of output, but a design choice by the developers.

            I tell non-technical people: It's a software program; there is always a person behind it, and they have consciously chosen this output, to serve their own interests.

        • cindyllm 20 hours ago
          [dead]
        • mmooss 21 hours ago
          Both claims seem to obviously contradict reality, imho:

          > Humans are wired to anthropomorphise; we do it with sticks and rocks.

          I do not see people anthropomorphize their TVs, books, paintings, cars, houses, rocks, sticks, ... I'm not sure what you're talking about here. I haven't seen someone call a rock or stick by a pronoun, for example.

          I've seen people fancifully do it, but it's exaggeration for a cute effect, and works because everyone knows they're exaggerating.

          > We’ve invented a technology that hijacks that part of humanity more effectively than anything ever. Asking people to “not be so stupid” is not going to help.

          People learn to look at things in different ways. Humans learn, society changes. It's changed quite a bit in recent years; it's much different than a century ago, a century before that, etc.

          In a way, it's more worship of technology to say humans are hopelessly, powerlessly dumb, and technology inevitably rules them. The inevitability of technology is of course a self-serving trope of the industry: you are powerless, we are inevitable. How convenient, and short-sighted.

          • lll-o-lll 21 hours ago
            I did not say that humans anthropomorphise every rock or stick. The point is that we are hard-wired for this possibility. Perhaps something more relatable? A cat? We ascribe all sorts of emotions to instinctive feline behaviours. It doesn’t matter with cats.

            > In a way, it's more worship of technology to say humans are hopelessly, powerlessly dumb, and technology inevitably rules them.

            You as an individual have some autonomy against the machines. Some. (Scroll facebook an hour a day and show me your resilience to engagement engineering). But the point of this is what happens at population scale.

            To ignore our biology is as stupid as suggesting we must be ruled by it.

      • palmotea 21 hours ago
        > The solution is to stop treating AI as superhuman semi-living entities, as many boosters are apt to do.

        It's not just fanboys doing that, but "official" educational channels. Both an internal training (more like an all-hands knowledge share) and some dumb training session by a guy from Anthropic explicitly, repeatedly recommended users anthropomorphize their chatbots on no uncertain terms. "Treat it as your super-smart coworkers", "I named mine Becky and I go to her with all my work challenges", etc.

        • customguy 19 hours ago
          ChatGPT just hit me with "I smiled at this:" ...
      • crooked-v 21 hours ago
        What I'd like is the TNG ship computer: an extremely smart system that nonetheless limits its "creativity" to the clearly marked holodeck, where simple prompting is treated like the casual throwaway diversion it is, and actual authors are still putting in substantial effort and personal preferences and viewpoints into their work.
    • alterom 21 hours ago
      >. AI is going to be internet hug box echo chambers at an unbelievable scale.

      More like hotbox of one's own farts, but yeah

  • madhu_ghalame 2 hours ago
    Exactly. Token costs are only going to increase over time. The later we start relying on manual work, the more expensive and unsustainable it will become. Investing in automation now gives us a significant long-term advantage and much greater confidence in our processes.
  • senshan 20 hours ago
    This could serve as a beacon for those who still care:

    Richard Feynman: "I have the advantage of having found out how hard it is to get to really know something, how careful you have to be about checking the experiments, how easy it is to make mistakes and fool yourself..."

    https://www.youtube.com/watch?v=tWr39Q9vBgo

    • beezle 19 hours ago
      just a psa for others: there is at least one, likely multiple youtube accounts that are making AI generated Feynman slop. Don't trust anything posted prior to the current dishearenting era.
      • Loughla 19 hours ago
        Don't you mean after the current era?
  • leecommamichael 20 hours ago
    I read the study. The wager is:

      $0.10 awarded for a correct answer
    
      $0.10 deducted for a wrong answer
    
      $0.00 (no change for a refusal to answer)
    
    
    There was no detail on whether participants were actually going to be paid, or if they were in the negative at the end, be expected to pay their losses.

    I am inclined to believe the effects of the study are real, but not nearly as pronounced as the data. If there were more serious amounts of money on the table, I think common sense would prevail.

  • rho138 22 hours ago
    I appreciate that they’re bringing this study to light… however, none of the source links actually trace back to the quoted study - instead pointing to their own domain or business insider.
  • andy99 22 hours ago
    This was posted earlier but didn’t get traction, and I made the following comment:

    Id want to know if “AI” makes a material difference vs just having access to the wrong answer. Like someone could be given search access that successfully retrieved wrong answers to questions, would that give the same results. How much do uniquely AI characteristics, like sycophancy or the conversational aspect play into this, vs people just being willing to believe what they read?

    • fxwin 22 hours ago
      You can read the study here: https://osf.io/preprints/psyarxiv/5y6m4_v1

      TLDR: They studied both cases (Access to a LLM/Chat interface which gives a wrong answer when asked, and access to pregenerated (wrong) answers). Both experiments yielded similar results.

      • tsimionescu 21 hours ago
        That is not a good summary. In all their versions of the experiment, they presented the tool as AI, and used actual LLM answers. In the first example, participants were directly interacting with a real, though small, LLM, but they had technical issues because of that - 10% of the time the LLM setup failed to present an answer at all. So they redid the experiment with pre-generated answers - the participants saw the same UI, but when they asked the question of the AI, they instead got one of 3 pre-generated answers from that AI, to avoid the technical issues.
        • fxwin 12 hours ago
          im not sure which part of my summary you take issue with then? the reason for study 1b is not relevant to answer the question asked in the comment i replied to
          • tsimionescu 11 hours ago
            It is relevant, because of (1) framing (people thought they were getting answers from AI, not a Google search, and any associated reputation works in that way); and (2) sycophancy and other similar characteristics of the answer's text, which were present in the answers presented in 1b and wouldn't be in a Google search.

            Basically, the difference between 1a and 1b is not at all relevant to the question of whether the observed behavior is caused by AI or simply by faulty tools. The difference between 1a and 1b was designed specifically to be transparent to the actual test takers, and only to eliminate some confounding variable (technical issues in 1a).

  • imoverclocked 21 hours ago
    I saw a post that some teachers are now asking students to ask ChatGPT to do their assignments… and then critique it where it is wrong.

    I feel like this might be an interesting method for places where the default response tends to be someone submitting copy-pasta: just give the ai-response and then have community discussion around what it missed.

    • rsoto2 16 hours ago
      nice try clammy sammy.

      We're gonna continue to push for scientific based approaches to education not some lunatic's idea of an educational experiment that waste's children's education. We tried this with Bill Gates already.

  • TimByte 13 hours ago
    Obviously without a proper RAG pipeline or web search it is just going to hallucinate with maximum confidence. It would be more interesting to see the results with top tier models that have proper alignment to refuse to answer when token probability is low. As it is they just proved that people tend to trust well written text in a chat ui
  • lukeschlather 21 hours ago
    Did they ask any substantive questions? It sounds like they asked movie trivia questions, in a way that the LLM was almost guaranteed to be wrong. So extremely low-stakes where there is no conceivable downside to being wrong.
    • tsimionescu 21 hours ago
      That was exactly the point - there was no downside to saying you just don't know, people didn't know, but they still chose to use (what they thought was) AI and got it wrong that way.
  • SubiculumCode 21 hours ago
    Instead of monetary rewards, they should preface it with, "This is AI model is not very accurate on these types of questions" and then see how they fare. In some areas, AI is quite accurate, so their experience may lead to different expectations.
  • nlawalker 21 hours ago
    And that’s a net win for a lot of people in a lot of situations; that’s the part that needs acknowledgment and study, sociologically.

    You can’t say it didn’t warn us, it’s on the tin: “AI might be wrong.”

  • idopmstuff 21 hours ago
    "The study, authored by Capraro with Chiara Marcoccia of École Normale Supérieure and Walter Quattrociocchi of Sapienza University of Rome, deliberately used questions where AI models typically fail: visual details from films, such as the colour of a team’s uniform in Bend It Like Beckham."

    I get why they used questions where AI models fail, but it also really reduces the value of this study. Nobody is really asking AI the color of a team's uniform in a movie, and if they do and confidently get it wrong, it just doesn't matter at all.

    Asking trivial questions also feels like it would affect the rate at which people are willing to confidently say things that are wrong. If you ask me some question of pointless trivia and I ask ChatGPT, I'll probably just repeat the answer because who cares. If you ask me something even mildly important and I ask ChatGPT, I'll either verify the information before I repeat it to you, or I'll qualify that I looked it up with ChatGPT and didn't verify. But some things are just so unimportant that they don't even warrant the disclaimer.

  • luciana1u 19 hours ago
    accuracy you can fix with a second source. the 2x confidence is the part without an undo button.
  • mcmisieck 13 hours ago
    So basically AI made people to be Americans ;)
  • onlyrealcuzzo 21 hours ago
    The TL;DR delivers:

    > Researchers found AI advice suppressed judgment suspension from 44% to 3%, accuracy from 27% to 9%, while confidence rose from 30% to 76%. People trusted wrong AI answers.

    I didn't see the article mentioned what they were actually asked, but I'm surprised confidence was only ~30%.

  • zahlman 15 hours ago
    > made people less accurate

    > sudy

    Intentional?

  • abernard1 21 hours ago
    Wouldn't it be wild if there were real-world penalties for being wrong?

    And people who follow bad advice will get bad results, and people who follow good advice will get good results?

    I wonder if this will impact the quality and prevalence of certain AI models in the future.

  • Joel_Mckay 19 hours ago
    Also may be irrelevant for people at some jobs, and cause cognitive skill losses over time.

    https://www.youtube.com/watch?v=axOcn--n_lM

    https://www.anthropic.com/research/AI-assistance-coding-skil...

    We should remember LLM do have legitimate use-cases like search, as we enter the "Trough of disillusionment" in the hype cycle. =3

    https://en.wikipedia.org/wiki/Gartner_hype_cycle

  • formvoltron 21 hours ago
    is that better or worse than 3 beers?
  • 1659447091 22 hours ago
    AI assisted/induced Dunning–Kruger effect
  • m0llusk 21 hours ago
    Right now we don't have AI. We have LLMs. LLMs compose responses and have now ability to say they don't know any answer or are unsure. This combines with the chipper manner of their communication to make for a dangerous mix.
  • diamondtea 20 hours ago
    An A for an I makes the whole world dumb
  • greatsage_sh 19 hours ago
    [flagged]
  • fbastosbkk 18 hours ago
    [flagged]
  • ojinai 16 hours ago
    [dead]
  • scoriiu 21 hours ago
    [flagged]
  • al_borland 23 hours ago
    [dead]
  • rsanek 21 hours ago
    Another case of study design not really supporting the headline. According to the paper, the comparison here was "no advice" or "AI advice": there was no "placebo" where you have advice / references from non-AI materials. Think e.g. Google's instant answer box, pre-AI. It's hard to say without the comparison, but I have to imagine displaying that would have similar effects.

    There's plenty of other issues with the study, one of the primary ones being that they chose a relatively 'dumb' AI (Step 3.5 Flash), but also, specifically hand-selected wrong answers. People that are used to using competent AIs that are mostly correct would be operating off of experience that suggests they should trust AI. In this case, by design, they shouldn't have, but it's hard to fault the user for that.

  • fwlr 22 hours ago
    Very concerning - I’ve tried n+m, n*m, n^m, and even m^n, but I can’t seem to reproduce the popular “10x” claim.
    • warkdarrior 21 hours ago
      Did you try n*m where n=10 and m=x?