@[email protected] to

[email protected]English • 9 months ago

ChatGPT Answers Programming Questions Incorrectly 52% of the Time: Study

cross-posted to:
[email protected]

788

ChatGPT Answers Programming Questions Incorrectly 52% of the Time: Study

@[email protected] to

[email protected]English • 9 months ago

cross-posted to:
[email protected]

To make matters worse, programmers in the study would often overlook the misinformation.

The research from Purdue University, first spotted by news outlet Futurism, was presented earlier this month at the Computer-Human Interaction Conference in Hawaii and looked at 517 programming questions on Stack Overflow that were then fed to ChatGPT.

“Our analysis shows that 52% of ChatGPT answers contain incorrect information and 77% are verbose,” the new study explained. “Nonetheless, our user study participants still preferred ChatGPT answers 35% of the time due to their comprehensiveness and well-articulated language style.”

Disturbingly, programmers in the study didn’t always catch the mistakes being produced by the AI chatbot.

“However, they also overlooked the misinformation in the ChatGPT answers 39% of the time,” according to the study. “This implies the need to counter misinformation in ChatGPT answers to programming questions and raise awareness of the risks associated with seemingly correct answers.”

Chat

@[email protected]
link
fedilink
English
56•9 months ago
GPT-2 came out a little more than 5 years ago, it answered 0% of questions accurately and couldn’t string a sentence together.

GPT-3 came out a little less than 4 years ago and was kind of a neat party trick, but I’m pretty sure answered ~0% of programming questions correctly.

GPT-4 came out a little less than 2 years ago and can answer 48% of programming questions accurately.

I’m not talking about mortality, or creativity, or good/bad for humanity, but if you don’t see a trajectory here, I don’t know what to tell you.
- @[email protected]
  link
  fedilink
  English
  104•9 months ago
  Removed by mod
  - @[email protected]
    link
    fedilink
    English
    25•9 months ago
    Perhaps there is some line between assuming infinite growth and declaring that this technology that is not quite good enough right now will therefore never be good enough?
    
    Blindly assuming no further technological advancements seems equally as foolish to me as assuming perpetual exponential growth. Ironically, our ability to extrapolate from limited information is a huge part of human intelligence that AI hasn’t solved yet.
    - @[email protected]
      link
      fedilink
      English
      2•9 months ago
      Removed by mod
  - @[email protected]
    link
    fedilink
    English
    16•9 months ago
    I appreciate the XKCD comic, but I think you’re exaggerating that other commenter’s intent.
    
    The tech has been improving, and there’s no obvious reason to assume that we’ve reached the peak already. Nor is the other commenter saying we went from 0 to 1 and so now we’re going to see something 400x as good.
    - @[email protected]
      link
      fedilink
      English
      6•9 months ago
      I think the one argument for the assumption that we’re near peak already is the entire issue of AI learning from AI input. I think numberphile discussed a maths paper that said that to achieve the accuracy that we want, there is simply not enough data to train it on.
      
      That’s of course not to say that we can’t find alternative approaches
    - @[email protected]
      link
      fedilink
      English
      3•9 months ago
      We’re close to peak using current NN architectures and methods. All this started with the discovery of transformer architecture in 2017. Advances in architecture and methods have been fairly small and incremental since then. The advancements in performance has mostly just been throwing more data and compute at the models, and diminishing returns have been observed. GPT-3 costed something like $15 million to train. GPT-4 is a little better and costed something like $100 million to train. If the next model costs $1 billion to train, it will likely be a little better.
    - @[email protected]
      link
      fedilink
      English
      1•
      edit-2
      9 months ago
      Removed by mod
      - @[email protected]
        link
        fedilink
        English
        1•9 months ago
        In general, “The technology is young and will get better with time” is not just a reasonable argument, but almost a consistent pattern. Note that XKCD’s example is about events, not technology. The comic would be relevant if someone were talking about events happening, or something like sales, but not about technology.
        
        Here, I’m not saying that you’re necessarily right or they’re necessarily wrong, just that the comic you shared is not a good fit.
        
        @[email protected]
        link
        fedilink
        English
        1•9 months ago
        Removed by mod
        
        @[email protected]
        link
        fedilink
        English
        1•9 months ago
        I don’t think continuing further would be fruitful. I imagine your stance is heavily influenced by your opposition to, or dislike of, AI/LLMs
        
        @[email protected]
        link
        fedilink
        English
        1•9 months ago
        Removed by mod
  - @[email protected]
    link
    fedilink
    English
    10•9 months ago
    That comes off as disingenuous in this instance.
- @[email protected]
  link
  fedilink
  English
  29•9 months ago
  The study is using 3.5, not version 4.
  - @[email protected]
    link
    fedilink
    English
    3•9 months ago
    4 produces inaccurate programming answers too
    - @[email protected]
      link
      fedilink
      English
      6•9 months ago
      Obviously. But it is FAR better yet again.
      - @[email protected]
        link
        fedilink
        English
        1•9 months ago
        Not really. I ask it questions all the time and it makes shit up.
        
        @[email protected]
        link
        fedilink
        English
        2•9 months ago
        Yes. But it is better than 3.5 without any doubt.
- Snot Flickerman
  link
  fedilink
  English
  23•
  edit-2
  9 months ago
  Removed by mod
  - @[email protected]
    link
    fedilink
    English
    9•9 months ago
    We are running these things on computers not designed for this. Right now, there are ASICs being built that are specifically designed for it, and traditionally, ASICs give about 5 orders of magnitude of efficiency gains.
- @[email protected]
  link
  fedilink
  English
  4•9 months ago
  Lemmy seems to be very near-sighted when it comes to the exponential curve of AI progress, I think this is an effect because the community is very anti-corp
- @[email protected]
  link
  fedilink
  English
  4•9 months ago
  Given the data points you made up, I feel it’s safe to assume that this plateau will now be a 10 year stretch
- @[email protected]
  link
  fedilink
  English
  3•9 months ago
  In what year do you estimating AI will have 90% accuracy?
  - @[email protected]
    link
    fedilink
    English
    5•9 months ago
    No clue? Somewhere between a few years (assuming some unexpected breakthrough) or many decades? The consensus from experts (of which I am not) seems to be somewhere in the 2030s/40s for AGI. I’m guessing accuracy probably will be more on a topic by topic basis, LLMs might never even get there, or only related to things they’ve been heavily trained on. If predictive text doesn’t do it then I would be betting on whatever Yann LeCun is working on.
- @[email protected]
  link
  fedilink
  English
  1•9 months ago
  We only need to keep doing incremental improvements in the technology and avoid destroying ourselves in the meantime. That’s all it takes for us to find ourselves in the presence of superintelligent AI one day.

[email protected]

[email protected]

You are not logged in. However you can subscribe from another Fediverse account, for example Lemmy or Mastodon. To do this, paste the following into the search field of your instance: [email protected]

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

4.54K users / day
9.56K users / week
17.3K users / month
31.6K users / 6 months
63K subscribers
13.5K Posts
566K Comments
Modlog