• the16bitgamer@piefed.ca
    link
    fedilink
    English
    arrow-up
    194
    ·
    4 days ago

    Looks for my username. Sees my 8 open repositories in there. Sees my poorly coded Uni projects are also in there.

    Oh lord my code is actively helping making AI worse.

    • iammike@programming.dev
      link
      fedilink
      arrow-up
      72
      ·
      4 days ago

      Mission failed successfully!

      Glad to be part of the crew with shitty code in Github to taint them plagiarism machines!

    • Rentlar@lemmy.ca
      link
      fedilink
      arrow-up
      28
      ·
      4 days ago

      Woo! My crappy code and group projects are in there too!

      May all AI generated code be in one giant main loop thanks to my influence 😈

      • Err(()).unwrap()@lemmy.world
        link
        fedilink
        arrow-up
        13
        ·
        4 days ago

        Thinking back to the sins I’ve committed in C# as a high schooler… maybe I should start publishing my works. Put some poison in the soup.

    • Rose@slrpnk.net
      link
      fedilink
      arrow-up
      3
      ·
      3 days ago

      Oh lord my code is actively helping making AI worse.

      I checked out, it has some of my repositories. They crawled this stuff in 2025 and they probably won’t update it. Whoever uses this dataset will have to deal with some super garbage, I tell ya.

    • Hudell@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      12
      arrow-down
      1
      ·
      4 days ago

      you joke but someone has actually traced to one of my commits a specific behavior that people on the internet have been complaining from LLMs recently.

    • Leon@pawb.social
      link
      fedilink
      arrow-up
      5
      ·
      4 days ago

      Hope some group makes it their mission to start flooding github with garbage code.

    • Axolotl@feddit.it
      link
      fedilink
      arrow-up
      4
      ·
      4 days ago

      For some reason the included only the repo for my github profile that has no code in it and not the other ones, my guess is that they avoided me because of the GPL license on all my code

      Spoiler

      also, some repositories were my first projects so they are way more shit than the avarage AI code and it would be funny if they poisoned themselves and i have some archived project that also have shit code but only on github because i switched to codeberg and rewrote them

  • mindbleach@sh.itjust.works
    link
    fedilink
    arrow-up
    27
    arrow-down
    1
    ·
    4 days ago

    Genuinely surprised it’s even opt-out.

    These companies train on Disney DVDs. Permission is not a factor. Training is transformative use, as much for counting letter frequency as for building a chatbot that can sort of code.

    • WhyJiffie@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      30
      ·
      4 days ago

      no, officer, you misunderstand! I’m not pirating this movie, I’m just training my intelligence on it! it is transformative use, see, I can now write you this summary!

        • Wiz@midwest.social
          link
          fedilink
          arrow-up
          1
          ·
          2 days ago

          I would say this is a gray area of law. It hasn’t been tested yet. There are a few factors in determining “fair use”. One of those factors is commercialization, which could nullify fair use. Another is the amount you’re using.

          • mindbleach@sh.itjust.works
            link
            fedilink
            arrow-up
            1
            ·
            2 days ago

            Commercial parody is commonplace fair use. If it surpasses the original… oh well.

            The amount being used becomes negligible as the corpus grows. Filtering a billion webpages into one gigabyte means their average contribution is one byte. Play with the numbers all you like, it’s gonna come out to less than a paragraph.

            If we went full copyright maximalist, and criminally banned using anything but public-domain / BSD / CC0 works, the commercial impact would not be much different. The technology itself shifts the economics of text, images, video, and code.

  • Bieren@lemmy.today
    link
    fedilink
    arrow-up
    9
    arrow-down
    1
    ·
    3 days ago

    If someone wants to scrap the shitty ass code I have on GitHub, have at it. Talk about poisoning AI

  • brucethemoose@lemmy.world
    link
    fedilink
    arrow-up
    66
    arrow-down
    1
    ·
    4 days ago

    Also, all this reminds me of drama in the Skyrim and Minecraft modding scenes, when devs publish stuff under Apache or MIT or whatever.

    Then the devs find out they don’t like what others are doing with their code. Drama ensues.


    …That’s kinda the deal with permissive licenses. Or posting publicly, like here on Lemmy. People will do things you don’t like with your code or content.

      • AlteredEgo@lemmy.ml
        link
        fedilink
        arrow-up
        2
        ·
        4 days ago

        The most important thing is to protect the people’s right to use open source / weight AI models without licensing. Because the AI companies permanently monopolizing AI models and collecting rent forever on the commons, that would be a tragedy. It would forever shift the balance of power to use advanced technology.

          • AlteredEgo@lemmy.ml
            link
            fedilink
            arrow-up
            1
            arrow-down
            2
            ·
            3 days ago

            Well truly profitable applications for LLMs are pretty limited so far, but they will come. Lets say you can replace a tax advisor / accountant with a specialized LLM agent (one that doesn’t hallucinate tax code lol). Imagine oyu have an open source project that uses a coding AI agent to frequently update and retrain the tax LLM with the new regulations and laws and unit tests all this. So it’s ultimately easy to use, you just talk to the tax AI agent. Or alternatively, because they managed to make open weight LLMs illegal because they are piracy because they use books and posts without permission to train and can’t afford to pay the license fees. Now everybody has to use a commercial AI tax service. At the same price at what it cost previously to hire a real human.

            As AI models get better they replace human labor. This is a problem under capitalism, otherwise we could just all just be working less and chill. But it gets even worse if “we” (like the 99%) can’t make use of open source models. Then the plutocrats get to own this newly replaced labor as free labor and collect rent.

            This is why I believe the IP argument of anti-AI is so dangerous. They can afford to pay license fees, we / open source can’t. See patents, or see healthcare insurance purposefully making healthcare costs go up.

      • FishFace@piefed.social
        link
        fedilink
        English
        arrow-up
        3
        arrow-down
        2
        ·
        4 days ago

        LLMs training on data don’t prevent anyone from themselves learning on that data, so I don’t see how the commons is going to be killed here.

        • Natanael@infosec.pub
          link
          fedilink
          arrow-up
          2
          ·
          3 days ago

          It heavily dampens the feedback loops which lead to new contributors joining. If more people feel encouraged to just let an LLM hack something for them then fewer things are learned by fewer people, less knowledge is developed and propagated, etc…

          And more of that which was shared is worse because there’s no quality control in those sharing pipelines, and ironically more ego behind it when the sharer has done less (see: that person in the group project who did nothing useful and took all credit)

    • Toga77@lemmy.world
      link
      fedilink
      arrow-up
      36
      arrow-down
      4
      ·
      4 days ago

      Nah if you’re a massive AI company and you scrape without contribution, you’re a huge piece of shit.

      It’s just stealing plain and simple like anything else.

      They’re not a small user getting open source software, they’re scraping what is already done to try and make you obsolete.

    • unknownuserunknownlocation@kbin.earth
      link
      fedilink
      arrow-up
      16
      ·
      4 days ago

      There are permissive licenses, and then there are copyleft licenses. Permissive licenses go in the direction of “do whatever the fuck you want”. Copyleft licenses are more like “use it for whatever the fuck you want but if you change it give it back to everyone else with the same conditions”. The people who have projects with copyleft licenses are the ones who are (rightfully) pissed about their projects being used to train AI.

      • brucethemoose@lemmy.world
        link
        fedilink
        arrow-up
        6
        arrow-down
        1
        ·
        4 days ago

        Huggingface isn’t violating copyleft licenses here, I don’t think. And research projects that use it, with citations and documentation, wouldn’t either.

        Now, if some business comes along and tries to make proprietary code derived from the dataset, that’s where things get hairy. But the people doing that are responsible for the potential violation, not Huggingface.

    • amio@lemmy.world
      link
      fedilink
      arrow-up
      16
      ·
      4 days ago

      On the other hand, totally did not see this particular shit coming. I wonder if we’ll see a rash of “permissive, except LLMs can fuck right off” licenses.

      • Lev@europe.pub
        link
        fedilink
        arrow-up
        16
        arrow-down
        7
        ·
        4 days ago

        Permissive licences should be the ones to fuck off entirely. GPL or death

        • GreyEyedGhost@piefed.ca
          link
          fedilink
          English
          arrow-up
          13
          ·
          4 days ago

          I’ve said it before and I’ll say it again: there is a place for permissive licenses. A great example is reference code for open standards, eg. TCP/IP.

          • Lev@europe.pub
            link
            fedilink
            arrow-up
            8
            ·
            4 days ago

            Permissive licences are useful compared to copyleft ones only for unfree actors, which we should fight integrally at every step of the way

        • boonhet@sopuli.xyz
          link
          fedilink
          arrow-up
          6
          arrow-down
          1
          ·
          4 days ago

          No, libraries should remain permissive IMO. Applications can be restrictive if they want to. I just don’t see any point in a copyleft library or framework.

            • boonhet@sopuli.xyz
              link
              fedilink
              arrow-up
              5
              arrow-down
              1
              ·
              4 days ago

              World runs on them though. I don’t mean business to consumer shit, all that can rot in hell. I mean business to business. There’s shit out there that’s incredibly niche, takes a ton of effort to develop, and there’s no way it would ever be achieved without a huge financial incentive (because it’s just so niche and it turns out paying tens or hundreds of people takes money). And someone’s gotta pay the people writing all your open source code so most of them need day jobs anyway, which will be difficult to have without any commercial software existing. Sometimes the “all our code is GPL, but you can pay us to host it for you” model works, but a lot of the time it doesn’t.

    • MonkeMischief@lemmy.today
      link
      fedilink
      arrow-up
      3
      ·
      4 days ago

      “Look, I wrote this neat highly advanced machine vision thing to help people with accessibility needs communicate with loved ones! ❤️. MIT licensed I guess! Let’s make the world better!”

      Raytheon, Northropp, Boeing, Microsoft, three-letter-agencies suddenly fork it as a base for their own “projects.”

      😐

  • onlinepersona@programming.dev
    link
    fedilink
    English
    arrow-up
    14
    arrow-down
    3
    ·
    3 days ago

    Get off of Github if you think this is a problem 🤷 There are alternatives like Forgejo (Codeberg), Gitlab, and Radicle (decentralised).

  • SirDimples@programming.dev
    link
    fedilink
    arrow-up
    2
    ·
    2 days ago

    Wow, got about 9 repos of mine there and a few from my startup’s, only public ones are scraped so I say fair enough. Happy to have switched to running my own git infra a year ago

  • ProbablyUnwise@anarchist.nexus
    link
    fedilink
    English
    arrow-up
    21
    ·
    4 days ago

    meanwhile I’m just here scraping GitHub repos for API keys and credentials 🤷‍♂️

    for legal reasons I must insist this is a joke, and in Minecraft.

  • Melllvar@startrek.website
    link
    fedilink
    English
    arrow-up
    18
    arrow-down
    1
    ·
    4 days ago

    A number of my repos are listed.

    But the weird part is that it also lists a repo I don’t recognize. The repo does actually exist on my github account, but it’s marked as “ignored”, and the description says it was automatically exported from Google Code. The code seems to be a MacOS shareware file encryption tool called “BitClamp”, published circa 2008.

    No idea how it got on my account.

      • Diurnambule@jlai.lu
        link
        fedilink
        arrow-up
        5
        ·
        3 days ago

        They got some broken Linux configuration from me, some project with many securities fails in it and a big amount of virus codes I got from the time I was hypefocusing on worms…

  • heliotrope@retrofed.com
    link
    fedilink
    English
    arrow-up
    34
    ·
    4 days ago

    Bad News: My old GitHub repos are there.

    Good News: I wrote that shit when I was 12. The code runs, but it’s not good and not inventive.

    • lyralycan@sh.itjust.works
      link
      fedilink
      arrow-up
      4
      ·
      4 days ago

      For me, they only have one of my repos, and I deleted all my repos months ago, so whether they actually have the code or just the title idk

      Good news: If anyone wants the one they got, I host it myself here instead. Fuck M$, AI thieves etc.

      I mean they can probably steal it from my site too but I’ve taken precautions, and the second biggest reason to migrate off corpo accounts entirely – my site is a much smaller hacker target than Github.

  • dextro@feddit.org
    link
    fedilink
    arrow-up
    23
    arrow-down
    1
    ·
    edit-2
    4 days ago

    I don’t see a problem as long as they stick to AGPL when building a product out of it

    Edit: Oh they also scraped my proprietary code 🧐

  • brucethemoose@lemmy.world
    link
    fedilink
    arrow-up
    27
    arrow-down
    2
    ·
    4 days ago

    Well… I’d rather the dataset be public and there, with an ostensible centralized opt-out, instead of every AI startup frantically rescraping the same things their predecessors did.

    • qaz@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      8
      ·
      4 days ago

      True, I understand why this is better than all those companies scraping it individually (for both ability to opt out and site load), but the way they handled opt out is still quite silly.

    • undefinedTruth@lemmy.zip
      link
      fedilink
      arrow-up
      3
      ·
      4 days ago

      If they only include repositories that come with a proper open source license technically they don’t even need to provide an opt-out option. So, good thing that at least it exists.

    • undefinedTruth@lemmy.zip
      link
      fedilink
      arrow-up
      4
      arrow-down
      2
      ·
      4 days ago

      Second that. After all all my open repositories are all either licenced under GPL or MIT, so complaining would be a kind of hypocritical.

      • gnutrino@programming.dev
        link
        fedilink
        English
        arrow-up
        11
        ·
        edit-2
        4 days ago

        GPL is dodgy to be included in training data for AIs that are then used to generate non-openaource code tbh. Even MIT loses the attribution it’s supposed to have once laundered through AI…

        Speaking as someone that’s been into FOSS for a long time, it does piss me off how quickly copyright got thrown under the bus the moment it became inconvenient for people with money.

        • undefinedTruth@lemmy.zip
          link
          fedilink
          arrow-up
          5
          ·
          4 days ago

          Yes, but what HuggingFace is doing here is the distribution of a data set. And so long the data set itself is open that doesn’t conflict with GPL. HuggingFace is not responsible about how others use that data set.

          Whether LLMs themselves violate copyright for being trained on MIT or GPL code is an entirely different discussion.