Crossposted from https://fedia.io/m/[email protected]/t/4218790

TL;DR: In the last year, the Wikimedia Foundation has fired several union organizers, including those that worked on the Community Tech team - a team dedicated to building features for the volunteer community that edits Wikipedia.

As the Wiki Workers Union tries to get the Wikimedia Foundation to recognize their union, it is worth remembering that this is not the first time that the Foundation has worked against the community.

Wikimedia Enterprise is a betrayal of the volunteer movement community of Wikipedia editors, as the Wikimedia Foundation is providing privileged access to big tech AI companies to the Wikipedia corpus - a body of work that the Foundation does not own.

Movement volunteer communities contributed to Wikipedia under copyleft licenses - licenses that work to ensure that the work remains free (as in speech). The big tech AI companies do not license derivative works under copyleft licenses and often do not even attribute where the works came from.

This means that volunteers are working for big tech for free, and the Wikimedia Foundation is selling privileged access to that free labor.

It is against that backdrop that the current unionization struggle unfolds.

  • Jhex@lemmy.world
    link
    fedilink
    English
    arrow-up
    52
    arrow-down
    3
    ·
    2 days ago

    The minute Wikipedia allowed AI to train on its knowledge base I stopped donating to them… what a sad sad world we are living in

    • antonim@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      arrow-down
      23
      ·
      2 days ago

      As opposed to not allowing them to do it and having every AI company scrape the site 24/7 anyway?

      • yoasif@fedia.ioOP
        link
        fedilink
        arrow-up
        29
        arrow-down
        1
        ·
        2 days ago

        They are sitting on $296M. Why aren’t they suing the AI companies to defend the contributors?

        • antonim@lemmy.world
          link
          fedilink
          English
          arrow-up
          7
          arrow-down
          20
          ·
          2 days ago

          Eh, as an editor I haven’t noticed I’m being attacked by AI companies. WMF would need some better argument if they’d want to sue successfully, especially aginst companies that are sitting on 100x more money than them.

          BTW it’s already established that training AI models on copyrighted materials is legal without approval of the copyright holder, and IMO it’s unikely WP’s material would be an exception.

          • yoasif@fedia.ioOP
            link
            fedilink
            arrow-up
            14
            ·
            2 days ago

            BTW it’s already established that training AI models on copyrighted materials is legal without approval of the copyright holder, and IMO it’s unikely WP’s material would be an exception.

            That isn’t accurate - this is a highly unsettled question, and there are multiple cases in litigation today.

            WMF would need some better argument if they’d want to sue successfully, especially aginst companies that are sitting on 100x more money than them.

            A better argument than what - that they are openly violating the licenses under which the encyclopedia is distributed? What more do they need?

            • Bongles@lemmy.zip
              link
              fedilink
              English
              arrow-up
              7
              ·
              2 days ago

              I think people have generally (incorrectly) seen certain settlements and deals, like with Disney as “this is legal now” when really it hasn’t been established yet. It’s just gone away for some companies because money.

              • Wirlocke@lemmy.blahaj.zone
                link
                fedilink
                English
                arrow-up
                5
                ·
                2 days ago

                I genuinely think Disney licensed to Sora because they saw how unprofitable it was and knew it would go under anyways.

                Because Disney got OpenAI to pay for a license they now have no use for.

            • antonim@lemmy.world
              link
              fedilink
              English
              arrow-up
              3
              ·
              2 days ago

              A better argument than what - that they are openly violating the licenses under which the encyclopedia is distributed?

              That’s a bit more convincing.

              this is a highly unsettled question, and there are multiple cases in litigation today

              I followed Kadrey v. Meta a little bit, and the conclusion was in favour of the training being fair use. What are the other ongoing cases?

              • melfie@lemmy.zip
                link
                fedilink
                English
                arrow-up
                2
                ·
                2 days ago

                When there’s a dispute, courts consider the following four issues in deciding whether a use is fair use:

                1. why the party used the copyrighted material (for instance, for commercial versus educational purposes)
                2. whether the copyrighted work is informational or for entertainment
                3. how much of the copyrighted work the party used, and
                4. whether and how the use affects the market for or value of the copyrighted work.
                1. Commercial
                2. Informational
                3. All of it
                4. Definitely does

                Based on this, not sure fair use holds up that well, but I’m not a lawyer.

                • antonim@lemmy.world
                  link
                  fedilink
                  English
                  arrow-up
                  2
                  ·
                  1 day ago

                  Regarding the first point, AI companies usually claim their work is scientific/educational. #3 can be responded to by claiming they actually use none of the original work, i.e. they don’t reproduce any of it; they liken training AI to learning from the materials, and usimg knowledge from some book isn’t just ‘fair use’, it’s the expected use of the book.

                  It’s mostly bullshit and sophistry, of course, but I’m afraid they will win in most of these cases since they have more money to dump into top-tier lawyers.

      • NOT_RICK@lemmy.world
        link
        fedilink
        English
        arrow-up
        19
        ·
        2 days ago

        I mean, yes. With this type of argument you could rationalize away criminalizing or forbidding anything.

  • melfie@lemmy.zip
    link
    fedilink
    English
    arrow-up
    9
    ·
    2 days ago

    Sure, train models on copyleft data as long as all the training data and pipeline code to train the model is open sourced under the terms of the license.

  • TheOrcWhoWrites@lemmy.world
    link
    fedilink
    English
    arrow-up
    6
    ·
    2 days ago

    <!–StartFragment–>

    Seemingly, like other community sites, they think that people will continue to contribute to the slop machines for free, probably because they are losers.

    <!–EndFragment–>

    that is absurd. i know someone personally who just became a wiki editor volunteer. just like any corporation, they treat their lower level staff and volunteers like shit.