“Our primary conclusion across all scenarios is that without enough fresh real data in each generation of an autophagous loop, future generative models are doomed to have their quality (precision) or diversity (recall) progressively decrease,” they added. “We term this condition Model Autophagy Disorder (MAD).”

Interestingly, this might be a more challenging problem as we increase the use of generative AI models online.

  • argv_minus_one@beehaw.orgBanned
    link
    fedilink
    English
    arrow-up
    19
    ·
    3 years ago

    Note that humans do not exhibit this property when trained on other humans, so this would seem to prove that “AI” isn’t actually intelligent.

    • PenguinTD@lemmy.ca
      link
      fedilink
      English
      arrow-up
      7
      ·
      3 years ago

      do we even need to prove this? Like anyone study a bit how generative AI works know it’s not intelligent.

  • frog 🐸@beehaw.org
    link
    fedilink
    English
    arrow-up
    5
    ·
    3 years ago

    Good!

    Was that petty?

    But, you know, good luck completely replacing human artists, musicians, writers, programmers, and everyone else who actually creates new content, if all generative AI models essentially give themselves prion diseases when they feed on each other.

      • frog 🐸@beehaw.org
        link
        fedilink
        English
        arrow-up
        4
        ·
        3 years ago

        I absolutely agree! I’ve seen so many proponents of AI argue that AI learning from artworks scraped from the internet is no different to a human learning by looking at other artists, and while anyone who is actually an artist (or involved in any creative industry at all, including things like coding that require a creative mind) can see the difference, I’ve always struggled to coherently express why. And I think this it. Human artists benefit from other human art to look at, as it helps them improve faster, but they don’t need it in the same way, and they’re more than capable of coming up with new ideas without it. Even a brief look at art history shows plenty of examples of human artists coming up with completely new ideas, artworks that had absolutely no precedent. I really can’t imagine AI ever being able to invent, say, Cubism without having seen a human do it first.

        I feel like the only people that are in favour of AI artworks are those who don’t see the value of art outside of its commercial use. They’re the same people who are, presumably, quite happy playing the same same-y games and watching same-y TV and films over and over again. AI just can’t replicate the human spark of creativity, and I really can’t see it being good for society either economically or culturally to replace artists with algorithms that can only produce derivations of what they’ve already seen.

  • feeltheglee@beehaw.org
    link
    fedilink
    English
    arrow-up
    5
    ·
    3 years ago

    You know how when you’re on a voice/video call and the audio keeps bouncing between two people and gets all feedback-y and screechy?

    That, but with LLMs.

  • Amax@lemmy.ca
    link
    fedilink
    English
    arrow-up
    5
    ·
    3 years ago

    MadAI’s disease.

    I guess we didn’t learn when we did it with cows.

  • Exaggeration207@beehaw.org
    link
    fedilink
    English
    arrow-up
    4
    ·
    3 years ago

    I only have a small amount of experience with generating images using AI models, but I have found this to be true. It’s like making a photocopy of a photocopy. The results can be unintentionally hilarious though.

  • Cybrpwca@beehaw.org
    link
    fedilink
    English
    arrow-up
    3
    ·
    3 years ago

    So we have generation loss instead of AI making better AI. At least for now. That’s strangely comforting.

  • coolin@beehaw.org
    link
    fedilink
    English
    arrow-up
    3
    ·
    3 years ago

    For the love of God please stop posting the same story about AI model collapse. This paper has been out since May, been discussed multiple times, and the scenario it presents is highly unrealistic.

    Training on the whole internet is known to produce shit model output, requiring humans to produce their own high quality datasets to feed to these models to yield high quality results. That is why we have techniques like fine-tuning, LoRAs and RLHF as well as countless datasets to feed to models.

    Yes, if a model for some reason was trained on the internet for several iterations, it would collapse and produce garbage. But the current frontier approach for datasets is for LLMs (e.g. GPT4) to produce high quality datasets and for new LLMs to train on that. This has been shown to work with Phi-1 (really good at writing Python code, trained on high quality textbook level content and GPT3.5) and Orca/OpenOrca (GPT-3.5 level model trained on millions of examples from GPT4 and GPT-3.5). Additionally, GPT4 has itself likely been trained on synthetic data and future iterations will train on more and more.

    Notably, by selecting a narrow range of outputs, instead of the whole range, we are able to avoid model collapse and in fact produce even better outputs.

    • shanghaibebop@beehaw.orgOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      3 years ago

      We’re all just learning here, but yeah, that’s pretty interesting to learn about effective synthetic data used for training.