Lemmy account of [email protected]

  • 0 Posts
  • 8 Comments
Joined 2 years ago
cake
Cake day: October 7th, 2024

help-circle




  • There are parental controls, but it’s rare. FOSS projects usually aren’t very keen on building systems of authority, even parental ones.

    For smartphones check out /e/OS (Murena Smartphones). They added parental controls recently, not sure how good those are though.

    For computers your best bet’s probably a Linux with Gnome. They also added something recently.

    You might be able to combine both with a custom DNS service, surely there’s something out there for that usecase (NextDNS?). But that’s already getting pretty complex.



  • I mean, they “plan” to publish that info, right?

    I’d bet they used at least CommonCrawl (with it being the majority of data), arXiv, Wikipedia and the set containing all of Github. Probably not a lot of distillation.

    CommonCrawl is one of the reasons small websites and social instances get DDoS’ed by rules-ignoring AI crawlers. Wikipedia data basically always gets used without paying them. Github… well, it’s a prime example of how they broke millions of licenses.

    There’s also other stuff commonly used. Don’t get me started on the training sets for image generation. It’s beyond disgusting (and I’m not even talking just about theft and cultural destruction at this point).

    This stuff is a bottomless pit, and if those (F)OSS models really aim to be up at the top they’ll have to break every possible moral, ethical and legal rule just like everyone else. Even more so if they omit closed-source training sets.