AI Data Laundering: How Academic and Nonprofit Researchers Shield Tech Companies from Accountability

Posted September 30, 2022September 30, 2022 by Andy Baio

Yesterday, Meta’s AI Research Team announced Make-A-Video, a “state-of-the-art AI system that generates videos from text.”

We’re pleased to introduce Make-A-Video, our latest in #GenerativeAI research! With just a few words, this state-of-the-art AI system generates high-quality videos from text prompts.

Have an idea you want to see? Reply w/ your prompt using #MetaAI and we’ll share more results. pic.twitter.com/q8zjiwLBjb
— Meta AI (@MetaAI) September 29, 2022

Like he did for the Stable Diffusion data, Simon Willison created a Datasette browser to explore WebVid-10M, one of the two datasets used to train the video generation model, and quickly learned that all 10.7 million video clips were scraped from Shutterstock, watermarks and all.

In addition to the Shutterstock clips, Meta also used 10 million video clips from this 100M video dataset from Microsoft Research Asia. It’s not mentioned on their GitHub, but if you dig into the paper, you learn that every clip came from over 3 million YouTube videos.

So, in addition to a massive chunk of Shutterstock’s video collection, Meta is also using millions of YouTube videos collected by Microsoft to make its text-to-video AI.

Non-Commercial Use

The academic researchers who compiled the Shutterstock dataset acknowledged the copyright implications in their paper, writing, “The use of data collected for this study is authorised via the Intellectual Property Office’s Exceptions to Copyright for Non-Commercial Research and Private Study.”

But then Meta is using those academic non-commercial datasets to train a model, presumably for future commercial use in their products. Weird, right?

Not really. It’s become standard practice for technology companies working with AI to commercially use datasets and models collected and trained by non-commercial research entities like universities or non-profits.

In some cases, they’re directly funding that research.

For example, many people believe that Stability AI created the popular text-to-image AI generator Stable Diffusion, but they funded its development by the Machine Vision & Learning research group at the Ludwig Maximilian University of Munich. In their repo for the project, the LMU researchers thank Stability AI for the “generous compute donation” that made it possible.

The massive image-text caption datasets used to train Stable Diffusion, Google’s Imagen, and the text-to-image component of Make-A-Video weren’t made by Stability AI either. They all came from LAION, a small nonprofit organization registered in Germany. Stability AI directly funds LAION’s compute resources, as well.

Shifting Accountability

Why does this matter? Outsourcing the heavy lifting of data collection and model training to non-commercial entities allows corporations to avoid accountability and potential legal liability.

It’s currently unclear if training deep learning models on copyrighted material is a form of infringement, but it’s a harder case to make if the data was collected and trained in a non-commercial setting. One of the four factors of the “fair use” exception in U.S. copyright law is the purpose or character of the use. In their Fair Use Index, the U.S. Copyright Office writes:

“Courts look at how the party claiming fair use is using the copyrighted work, and are more likely to find that nonprofit educational and noncommercial uses are fair.”

A federal court could find that the data collection and model training was infringing copyright, but because it was conducted by a university and a nonprofit, falls under fair use.

Meanwhile, a company like Stability AI would be free to commercialize that research in their own DreamStudio product, or however else they choose, taking credit for its success to raise a rumored $100M funding round at a valuation upwards of $1 billion, while shifting any questions around privacy or copyright onto the academic/nonprofit entities they funded.

Not sure if I mentione but stable diffusion is a model created and released by CompVis at University of Heidelberg.
The LAION dataset is created by the German charity of the same name. We support both.
This is outlined in the announcement post.
Has implications for stuff above
— Emad (@EMostaque) August 16, 2022

This academic-to-commercial pipeline abstracts away ownership of data models from their practical applications, a kind of data laundering where vast amounts of information are ingested, manipulated, and frequently relicensed under an open-source license for commercial use.

Unforeseen Consequences

Years ago, like many people, I used to upload my photos to Flickr with a Creative Commons license that required attribution and allowed non-commercial use. Yahoo released a database of 100 million of those Creative Commons-licensed images for academic research, to help the burgeoning field of AI. Researchers at the University of Washington took 3.5 million of the Flickr photos with faces in them, over 670,000 people (including me), and released the MegaFace dataset, part of a research competition sponsored by Google and Intel.

I was happy to let people remix and reuse my photos for non-commercial use with attribution, but that’s not how they were used. Instead, academic researchers took the work of millions of people, stripped it of attribution against its license terms, and redistributed it to thousands of groups, including corporations, military agencies, and law enforcement.

In their analysis for Exposing.ai, Adam Harvey and Jules LaPlace summarized the impact of the project:

[The] MegaFace face recognition dataset exploited the good intentions of Flickr users and the Creative Commons license system to advance facial recognition technologies around the world by companies including Alibaba, Amazon, Google, CyberLink, IntelliVision, N-TechLab (FindFace.pro), Mitsubishi, Orion Star Technology, Philips, Samsung1, SenseTime, Sogou, Tencent, and Vision Semantics to name only a few. According to the press release from the University of Washington, “more than 300 research groups [were] working with MegaFace” as of 2016, including multiple law enforcement agencies.

That dataset was used to build the facial recognition AI models that now power surveillance tech companies like Clearview AI, in use by law enforcement agencies around the world, as well as the U.S. Army. The Chinese government has used it to train their surveillance systems. As the New York Times reported last year:

MegaFace has been downloaded more than 6,000 times by companies and government agencies around the world, according to a New York Times public records request. They included the U.S. defense contractor Northrop Grumman; In-Q-Tel, the investment arm of the Central Intelligence Agency; ByteDance, the parent company of the Chinese social media app TikTok; and the Chinese surveillance company Megvii.

The University of Washington eventually decommissioned the dataset and no longer distributes it. I don’t think any of those researchers, or even the people at Yahoo who decided to release the photos in the first place, ever foresaw how it would later be used.

408 of about 4,753,520 face images from the MegaFace face recognition training and benchmarking dataset. Visualization by Adam Harvey of Exposing.ai.

They were motivated to push AI forward and didn’t consider the possible repercussions. They could have made inclusion into the dataset opt-in, but they didn’t, probably because it would’ve been complicated and the data wouldn’t have been nearly as useful. They could have enforced the license and restricted commercial use of the dataset, but they didn’t, probably because it would have been a lot of work and probably because it would have impacted their funding.

Asking for permission slows technological progress, but it’s hard to take back something you’ve unconditionally released into the world.

As I wrote about last month, I’m incredibly excited about these new AI art systems. The rate of progress is staggering, with three stunning announcements yesterday alone: aside from Meta’s Make-A-Video, there was also DreamFusion for text-to-3D synthesis and Phenaki, another text-to-video model capable of making long videos with prompts that change over time.

But I grapple with the ethics of how they were made and the lack of consent, attribution, or even an opt-out for their training data. Some are working on this, but I’m skeptical: once a model is trained on something, it’s nearly impossible for it to forget. (At least for now.)

Like with the artists, photographers, and other creators found in the 2.3 billion images that trained Stable Diffusion, I can’t help but wonder how the creators of those 3 million YouTube videos feel about Meta using their work to train their new model.

A mysterious voice is haunting American Airlines’ in-flight announcements and nobody knows how

Posted September 23, 2022September 26, 2022 by Andy Baio

Here’s a little mystery for you: there are multiple reports of a mysterious voice grunting, moaning, and groaning on American Airlines’ in-flight announcement systems, sometimes lasting the duration of the flight — and nobody knows who’s responsible or how they did it.

Actor/producer Emerson Collins was the first to post video, from his Denver flight on September 6:

The weirdest flight ever.
These sounds started over the intercom before takeoff and continued throughout the flight.
They couldn’t stop it, and after landing still had no idea what it was. pic.twitter.com/F8lJlZHJ63
— Emerson Collins (@ActuallyEmerson) September 23, 2022

Here’s an MP3 of the audio with just the groans, moans, and grunts, with some of the background noise filtered out.

This is the only video evidence so far, but Emerson is one of several people who have experienced this on multiple different American Airlines flights. This thread from JonNYC collected several different reports from airline employees and insiders, on both Airbus A321 and Boeing 737-800 planes.

Other people have reported similar experiences, always on American Airlines, going as far back as July. Every known incident has gone through the greater Los Angeles area (including Santa Ana) or Dallas-Fort Worth. Here are all the incidents I’ve seen so far, in chronological order:

July – American Airlines, JFK to LAX. Bradley P. Allen wrote, “My wife and I experienced this during an AA flight in July. To be clear, it was just sounds like the moans and groans of someone in extreme pain. The crew said that it had happened before, and had no explanation. Occurred briefly 3 or 4 times early in the flight, then stopped.” (Additional flight details via the LA Times.)
August 5 – American Airlines 117. JFK to LAX. Wendy Wanderman wrote, “It happened on my flight August 5 from JFK to LAX and it was an older A321 that I was on. It was Flight 117. There was flight crew that was on the same plane a couple days earlier and the same thing happened. It was funny and unsettling.”
September 6 – American Airlines. Santa Ana, CA to Dallas-Fort Worth. Emerson Collins’ flight. “These sounds started over the intercom before takeoff and continued throughout the flight. They couldn’t stop it, and after landing still had no idea what it was… I filmed about fifteen minutes, then again during service. It was calmer for a while mid flight.”
Mid-September – American Airlines, Airbus A320. Orlando, FL to Dallas-Fort Worth. Doug Boehner wrote, “This happened to me last week. It wasn’t the whole flight, but periodically weird phrases and sounds. Then a huge ‘oh yeah’ when we landed. We thought the pilot left his mic open.”
September 18 – American Airlines 1631, Santa Ana, CA to Dallas-Fort Worth. Boeing 737-800. An anonymous report passed on by JonNYC, “Currently on AA1631 and someone keeps hacking into the PA and making moaning and screaming sounds 😨 the flight attendants are standing by their phones because it isn’t them and the captain just came on and told us they don’t think the flight systems are compromised so we will finish the flight to DFW. Sounded like a male voice and wouldn’t last more than 5-10 seconds before stopping. And has [intermittently] happened on and off all flight long.” (And here’s a second person on the same flight.)

Interestingly, JonNYC followed up with the person who reported the incident on September 18 and asked if it sounded like the same voice in the video. “Very very similar. Same voice! But ours was less aggressive. Although their volume might have been turned up more making it sound more aggressive. 100% positive same voice.“

Official Response

View from the Wing’s Gary Leff asked American Airlines about the issue, and their official response is that it’s a mechanical issue with the PA amplifier. The LA Times followed up on Saturday, with slightly more information:

“Our maintenance team thoroughly inspected the aircraft and the PA system and determined the sounds were caused by a mechanical issue with the PA amplifier, which raises the volume of the PA system when the engines are running,” said Sarah Jantz, a spokesperson for American.
Jantz said the P.A. systems are hardwired with no external access and no Wi-Fi component. The airline’s maintenance team is reviewing the additional reports. Jantz did not respond to questions about how many reports it has received and whether the reports are from different aircrafts.

This explanation feels incomplete to me. How can an amplifier malfunction broadcast what sounds like a human voice without external access? On multiple flights and aircraft? They seem to be saying the source is artificial, but has anyone heard artificial noise that sounds this human?

Why This Is So Bizarre

By nature, passenger announcement systems on planes are hardwired, closed systems, making them incredibly difficult to hack. Professional reverse engineer/hardware hacker/security analyst Andrew Tierney (aka Cybergibbons) dug up the Airbus 321 documents in this thread.

So… We've had a good dig into this.

The A321 passenger announcement system looks to be physically discrete to the interphone and other systems.

We're struggling to see a path. https://t.co/qVdJR6cUm0
— Cybergibbons 🚲🚲🚲 (@cybergibbons) September 23, 2022

They don't tend to put things in planes that are not needed.

You can speak on the interphone (the cabin telephone system) from lots of places including the belly panel and engines. But it's wired.
— Cybergibbons 🚲🚲🚲 (@cybergibbons) September 23, 2022

“And on the A321 documents we have, the passenger announcement system and interphone even have their own handsets. Can’t see how IFE or WiFi would bridge,” Tierney wrote. “Also struggling to see how anyone could pull a prank like this.”

This report found by aviation watchdog JonNYC, posted by a flight attendant on an internal American Airlines message board, points to some sort of as-yet-undiscovered remote exploit.

pic.twitter.com/5Ol4zuSb9u
— 🇺🇦 JonNYC 🇺🇦 (@xJonNYC) September 20, 2022

We also know that, at least on Emerson Collins’ flight, there was no in-seat entertainment, eliminating that as a possible exploit vector.

They did not! This was a watch inflight entertainment on your phone flight
— Emerson Collins (@ActuallyEmerson) September 23, 2022

Theories

So, how could this happen? There are a handful of theories, but they’re very speculative.

Medical Intercom

The first to emerge was this now-debunked theory came from “a former avionics guy” posting in r/aviation on Reddit:

The most likely culprit IMHO is the medical intercom. There are jacks mounted in the overhead bins at intervals down the full length of the airplane that have both receive, transmit and key controls. All somebody would need to do is plug a homemade dongle with a Bluetooth receiver into one of those, take a trip to the lav and start making noises into a paired mic.
The fact that the captain’s announcements are overriding (ducking) it but the flight attendants aren’t is also an indication it’s coming from that system.

If this was how it was done, there’s no reason the prankster would need to hide in the bathrooms: they could trigger a soundboard or prerecorded audio track from their seat.

However, this theory is likely a dead end. JonNYC reports that an anonymous insider confirmed they no longer exist on American Airlines flights. And even if they existed, the medical intercoms didn’t patch into the announcement system. They only allow flight crew to talk to medical staff on the ground.

someone says:
“These don’t exist on AA. We use an app on our iPhone to contact medical personnel on the ground. No such port exists, not since the super80 and they were inop’d.”
— 🇺🇦 JonNYC 🇺🇦 (@xJonNYC) September 24, 2022

Pre-Recorded Audio Message Bug

Another theory, also courtesy of JonNYC, is that there’s an issue with the pre-recorded audio messages (“PRAM”), which were replaced in the last 60 days, within the timeframe of all these incidents. Perhaps some test audio was added to the end of a message, maybe by an engineer who worked on it, and it’s accidentally playing that extra audio?

.. as an intermittent inflight announcement? Maybe the IT version of blowing a slide with a beer in your hand. 😂"

⬆️THE "CRAZY THEORY" PART PLEASE ⬆️
— 🇺🇦 JonNYC 🇺🇦 (@xJonNYC) September 25, 2022

It's probably the PRAM… Pre-Recorded Announcement Machine.

These have solid state storage, techs just load files they get from *somewhere*, test procedure for audio is less than 20 minutes to check it out, and it can be interrupted by inflight announcements.
— Mɪᴄʜᴀᴇʟ Tᴏᴇᴄᴋᴇʀ (@mtoecker) September 23, 2022

Artificial Noise

Finally, some firmly believe that it’s not a human voice at all, but artificial noise or audio feedback filtered through the announcement system.

Nick Anderegg, an engineer with a background in linguistics and phonology, says it’s the results of “random signal passed through a system that extracts human voices.”

An amp malfunction that inputs the signal through algorithms meant to isolate the human voice. All the non-human aspects of the random signal will be stripped out, and the result will appear human. https://t.co/2bXYVCFs2l
— Nick Anderegg loudly supports human rights (@NickAnderegg) September 26, 2022

Anderegg points to a sound heard at the 1:20 mark in Emerson’s video, a “sweep across every frequency,” as evidence that American Airlines’ explanation is accurate.

The tone sweep is just a sign that it’s artificial. Random signals (i.e. interference), when passed through systems designed to isolate the human voice, will make them sound human. It’s attempting to extract a coherent signal where there is none, so it’s approximating one
— Nick Anderegg loudly supports human rights (@NickAnderegg) September 26, 2022

Personally, I struggle with this explanation. The wide variation of utterances heard during Emerson’s three-hour flight are so wildly different, from groans and grunts to moans and shouts, that it’s difficult to imagine it as anything else but human. It’s far from impossible, but I’d love to see anyone try to recreate these sounds with random noise or feedback.

Any Other Ideas?

Any other theories how this might be possible? I’d love to hear them, and I’ll keep this post updated. My favorite theory so far:

Flying hurts the clouds and their screams are picked by the PA system. Seems pretty obvious 🙄
— Horus First (@HorusFirst) September 23, 2022

Photoillustration of an American Airlines jet headed towards a scared cartoon cloud

Online Art Communities Begin Banning AI-Generated Images

Posted September 9, 2022September 13, 2022 by Andy Baio

As AI-generated art platforms like DALL-E 2, Midjourney, and Stable Diffusion explode in popularity, online communities devoted to sharing human-generated art are forced to make a decision: should AI art be allowed?

Collage of dozens of images made with Stable Diffusion, indexed by Lexica

On Sunday, popular furry art community Fur Affinity announced that AI-generated art was not allowed because it “lacked artistic merit.” (In July, one AI furry porn generator was uploading one image every 40 seconds before it was banned.) Their new guidelines are very clear:

Content created by artificial intelligence is not allowed on Fur Affinity.
AI and machine learning applications (DALL-E, Craiyon) sample other artists’ work to create content. That content generated can reference hundreds, even thousands of pieces of work from other artists to create derivative images.
Our goal is to support artists and their content. We don’t believe it’s in our community’s best interests to allow AI generated content on the site.

Last year, the 27-year-old art/animation portal Newgrounds banned images made with Artbreeder, a tool for “breeding” GAN-generated art. Late last month, Newgrounds rewrote their guidelines to explicitly disallow images generated by new generation of AI art platforms:

AI-generated art is not allowed in the Art Portal. This includes using tools such as Midjourney, Dall-E, and Craiyon, in addition fractal generators and websites like ArtBreeder, where the user selects two images and they are combined into a new image via machine learning.
There are cases where some use of AI is ok, for example if you are primarily showcasing your character art but use an AI-generated background. In these cases, please note any elements where AI was used so that it is clear to users and moderators.
Tracing and coloring over AI-generated art is something best shared on your blog, as it is much like tracing over someone else’s art.
Bottom line: We want to keep the focus on art made by people and not have the Art Portal flooded with computer-generated art.

It’s not just long-running online communities: InkBlot is a budding art platform funded on Kickstarter in 2021 that went into open beta just this week. They’ve already taken a “no tolerance” policy against AI art, and updating their terms of service to exclude it.

Hi, we mentioned a few days ago that we have a no tolerance for AI art & working on updating our ToS in coming day for this which you can see in tweet here: https://t.co/5NCCKDYVWv
— 🦋InkBlot @ MEMBERSHIP DRIVE (@inkblot_art) September 9, 2022

Platforms that haven’t taken a stand are now facing public pressure to clarify their policies.

DeviantArt is one of the most popular online art communities, and increasingly, members are complaining that their feeds are getting flooded with AI-generated art. One of the most popular threads in their forums right now asks the staff to “combat AI art” by limiting daily uploads, either by segregating it under a special category or to ban it entirely.

@DeviantArt You were kind of the last art site dedicated to art, but everyday I check the site now more and more its Ai. 10 out of 25 on your front page is Ai gen images. I guess this actually might be the end of a lot of art sites? I hope someone steps in and makes a new site pic.twitter.com/1Kez5FFQQF
— Zakuga Art (@ZakugaMignon) September 6, 2022

ArtStation has also been quiet as AI-generated images grow in popularity there. “Trending on ArtStation” is one of the most popular prompts for AI art because of the particular aesthetic and quality of work found there, which nudges the AI to generate work scraped from it, leading to a future ouroboros where AI models will be trained on AI-generated art found there.

Every time I go to DA or Artstation these days the front pages are flooded with unmodified AI generated slop. Its ugly and makes the sites feel lesser. I go to these places to be inspired, not demoralized.
— RJ Palmer (@arvalis) September 9, 2022

However you feel about the ethics of AI art, online art communities are facing a very real problem of scale: AI art can be created orders of magnitude faster than traditional human-made art. A powerful GPU can generate thousands of images an hour, even while you sleep.

Lexica, a search engine that solely indexed images from Stable Diffusion’s beta tests in Discord, has over 10 million images in it. It would take a lifetime to explore everything in it, a corpus made by a relatively small group of beta testers in a few weeks.

Left unchecked, it’s not hard to imagine AI art crowding out illustrations that took days or weeks for someone to make.

To keep their communities active, community admins and moderators will have to decide what to do with AI art: allow it, segregate it, or ban it entirely.

Perfect Tides, a Coming-of-Age Point-and-Click Adventure, Kickstarts a Sequel

Posted August 31, 2022September 3, 2022 by Andy Baio

There’s no shortage of amazing games so far this year, but my personal favorite is an underdog: Perfect Tides, a ’90s-esque point-and-click adventure about growing up as a teen on a sleepy island resort town in the early 2000s, finding an escape from real-life feelings of loneliness and loss in discussion forums and late-night AIM chats.

Mara and her friend Lily on the beach… definitely not on drugs

The first game from Meredith Gran, creator of the decade-long comic series Octopus Pie, it approaches challenging subjects with the confidence of someone who created narrative comics every week for ten years. I can’t think of another comics artist who has dived into game design like this, but it pays off with uniquely charming pixel art and animation, colorful writing, and a story that genuinely moved me by the end. It navigates complex feelings about family, old friends, and new loves, while also being genuinely funny.

Let’s put it this way: I’ve been playing videogames for the last 35 years, but Perfect Tides is the first time I felt compelled to write a walkthrough (spoilers!) and actively participate in forums to help people finish it.

This is a long way of saying that you should play Perfect Tides on Steam or Itch, and then go back the Kickstarter for its sequel, Perfect Tides: Station to Station, which has only six days to go and still needs another $20,000 to cross the finish line. (Update: It hit the goal!)

you should probably go back this project

But you don’t have to take my word for it! Kotaku said the original game was “one of the year’s best,” the “kind of game you don’t even see coming, yet turns out to be incredible” and “perfectly captures the intensity and struggle of adolescence.” AV Club called it a “harrowing, funny, beautiful, horrifying, and ultimately reassuring work of art.” Polygon summed it up as “devastatingly honest.” My favorite review was from Buried Treasure’s John Walker, who wrote, “It is the most extraordinary exploration of what it is to be a teenager, told with such heart, such truth.”

Spoilers Ahoy

If you’ve already played Perfect Tides, I want to mention two key moments that are so wonderful, and yet so easy to miss in your first playthrough, they’re worth replaying it for. THESE ARE SPOILERS!

First, if you didn’t manage to patch things up with Lily, you missed a long sequence with her in the final season of the game. (To get the full experience of that sequence, you’ll need to find a specific MP3 and put into the game directory when prompted: a remarkable breaking-the-fourth-wall sidestep around copyright licensing that I’ve never seen in a game before.)

Second, there are two major endings. If it feels anticlimactic, you likely didn’t resolve your conflicts with Lily, Simon, and your family. There are 95 possible points, but you don’t need them all to get the best ending. Feel free to use my 100% completion guide for help getting there.

Perfect Tides isn’t perfect. Like any classic point-and-click adventure, there are some clunky bits here and there, and you’ll likely need the occasional hint or glance at a playthrough to finish. But it’s so worth it.

Exploring 12 Million of the 2.3 Billion Images Used to Train Stable Diffusion’s Image Generator

Posted August 30, 2022December 20, 2023 by Andy Baio

One of the biggest frustrations of text-to-image generation AI models is that they feel like a black box. We know they were trained on images pulled from the web, but which ones? As an artist or photographer, an obvious question is whether your work was used to train the AI model, but this is surprisingly hard to answer.

Sometimes, the data isn’t available at all: OpenAI has said it’s trained DALL-E 2 on hundreds of millions of captioned images, but hasn’t released the proprietary data. By contrast, the team behind Stable Diffusion have been very transparent about how their model is trained. Since it was released publicly last week, Stable Diffusion has exploded in popularity, in large part because of its free and permissive licensing, already incorporated into the new Midjourney beta, NightCafe, and Stability AI’s own DreamStudio app, as well as for use on your own computer.

But Stable Diffusion’s training datasets are impossible for most people to download, let alone search, with metadata for millions (or billions!) of images stored in obscure file formats in large multipart archives.

So, with the help of my friend Simon Willison, we grabbed the data for over 12 million images used to train Stable Diffusion, and used his Datasette project to make a data browser for you to explore and search it yourself. Note that this is only a small subset of the total training data: about 2% of the 600 million images used to train the most recent three checkpoints, and only 0.5% of the 2.3 billion images that it was first trained on.

Screenshot of the LAION-Aesthetic data browser, showing results from a search for Swedish artist Simon Stålenhag with thumbnail images — Screenshot of the LAION-Aesthetic data browser in Datasette

Go try it right now at laion-aesthetic.datasette.io! (Update: It’s now offline. See below for details.)

Read on to learn about how this dataset was collected, the websites it most frequently pulled images from, and the artists, famous faces, and fictional characters most frequently found in the data.

Data Source

Stable Diffusion was trained off three massive datasets collected by LAION, a nonprofit whose compute time was largely funded by Stable Diffusion’s owner, Stability AI.

All of LAION’s image datasets are built off of Common Crawl, a nonprofit that scrapes billions of webpages monthly and releases them as massive datasets. LAION collected all HTML image tags that had alt-text attributes, classified the resulting 5 billion image-pairs based on their language, and then filtered the results into separate datasets using their resolution, a predicted likelihood of having a watermark, and their predicted “aesthetic” score (i.e. subjective visual quality).

Collage of some of the images with the highest “aesthetic” score, largely watercolor landscapes and portraits of women

Stable Diffusion’s initial training was on low-resolution 256×256 images from LAION-2B-EN, a set of 2.3 billion English-captioned images from LAION-5B‘s full collection of 5.85 billion image-text pairs, as well as LAION-High-Resolution, another subset of LAION-5B with 170 million images greater than 1024×1024 resolution (downsampled to 512×512).

Its last three checkpoints were on LAION-Aesthetics v2 5+, a 600 million image subset of LAION-2B-EN with a predicted aesthetics score of 5 or higher, with low-resolution and likely watermarked images filtered out.

For our data explorer, we originally wanted to show the full dataset, but it’s a challenge to host a 600 million record database in an affordable, performant way. So we decided to use the smaller LAION-Aesthetics v2 6+, which includes 12 million image-text pairs with a predicted aesthetic score of 6 or higher, instead of the 600 million rated 5 or higher used in Stable Diffusion’s training.

This should be a representative sample of images used to train Stable Diffusion’s last three checkpoints, but skewing towards more aesthetically-attractive images. Note that LAION provides a useful frontend to search the CLIP embeddings computed from their 400M and 5 billion image datasets, but it doesn’t allow you to search the original captions.

Source Domains

We know the captioned images used for Stable Diffusion were scraped from the web, but from where? We indexed the 12 million images in our sample by domain to find out.

Nearly half of the images, about 47%, were sourced from only 100 domains, with the largest number of images coming from Pinterest. Over a million images, or 8.5% of the total dataset, are scraped from Pinterest’s pinimg.com CDN.

User-generated content platforms were a huge source for the image data. WordPress-hosted blogs on wp.com and wordpress.com represented 819k images together, or 6.8% of all images. Other photo, art, and blogging sites included 232k images from Smugmug, 146k from Blogspot, 121k images were from Flickr, 67k images from DeviantArt, 74k from Wikimedia, 48k from 500px, and 28k from Tumblr.

Shopping sites were well-represented. The second-biggest domain was Fine Art America, which sells art prints and posters, with 698k images (5.8%) in the dataset. 244k images came from Shopify, 189k each from Wix and Squarespace, 90k from Redbubble, and just over 47k from Etsy.

Unsurprisingly, a large number came from stock image sites. 123RF was the biggest with 497k, 171k images came from Adobe Stock’s CDN at ftcdn.net, 117k from PhotoShelter, 35k images from Dreamstime, 23k from iStockPhoto, 22k from Depositphotos, 22k from Unsplash, 15k from Getty Images, 10k from VectorStock, and 10k from Shutterstock, among many others.

It’s worth noting, however, that domains alone may not represent the actual sources of these images. For instance, there are only 6,292 images sourced from Artstation.com’s domain, but another 2,740 images with “artstation” in the caption text hosted by sites like Pinterest.

Artists

We wanted to understand how artists were represented in the dataset, so used the list of over 1,800 artists in MisterRuffian’s Latent Artist & Modifier Encyclopedia to search the dataset and count the number of images that reference each artist’s name. You can browse and search those artist counts here, or try searching for any artist in the images table. (Searching with quoted strings is recommended.)

Of the top 25 artists in the dataset, only three are still living: Phil Koch, Erin Hanson, and Steve Henderson. The most frequent artist in the dataset? The Painter of Light™ himself, Thomas Kinkade, with 9,268 images.

From a list of 1,800 popular artists, the top 10 found most frequently in the captioned images

Using the “type” field in the database, you can see the most frequently-found artists in each category: for example, looking only at comic book artists, Stan Lee’s name is found most often in the image captions. (As one commenter pointed out, Stan Lee was a comic book writer, not an artist, but people are using his name to generate images in the style of comic book art he was associated with.)

Some of the most-cited recommended artists used in AI image prompting aren’t as pervasive in the dataset as you’d expect. There are only 15 images that mention fantasy artist Greg Rutkowski, whose name is frequently used as a prompt modifier, and only 73 from James Gurney.

(It’s worth saying again that these images are just a subset of one of three datasets used to train the AI, so an artist’s work may have been used elsewhere in the data even if they’re not found in these 12M images.)

Famous People

Unlike DALL-E 2, Stable Diffusion doesn’t have any limitations on generating images of people named in the dataset. To get a sense of how well-represented well-known people are in the dataset, we took two lists of celebrities and other famous names and merged it into a list of nearly 2,000 names. You can see the results of those celebrity counts here, or search for any name in the images table. (Obviously, some of the top searches like “Pink” and “Prince” include results that don’t refer to that person.)

Donald Trump is one of the most cited names in the image dataset, with nearly 11,000 photos referencing his name. Charlize Theron is a close runner-up with 9,576 images.

Collage of generated portraits of Donald Trump and Charlize Theron from Stable Diffusion

A full gender breakdown would take more time, but at a glance, it seems like many of the most popular names in the dataset are women.

Strangely, enormously popular internet personalities like David Dobrik, Addison Rae, Charli D’Amelio, Dixie D’Amelio, and MrBeast don’t appear in the captions from the dataset at all. My hunch was that the CommonCrawl data was too old to include these more recent celebrities, but based on the URLs, there are tens of thousands of images from last year in the data. (If you can solve this mystery, get in touch or leave a comment!)

Fictional Characters

Finally, we took a look at how popular fictional characters are represented in the dataset, since this is subject matter that’s enormously popular using Stable Diffusion and Craiyon, but often impossible with DALL-E 2, as you can see in this Mickey Mouse example from my previous post.

“realistic 3d rendering of mickey mouse working on a vintage computer doing his taxes” on DALL·E 2 (left) vs. Stable Diffusion (right)

For this set of searches, we used this list of 600 fictional characters from pop culture to search the image dataset. You can browse the results here, or search for any other character in the images table. (Again, be aware that one-word character names like “Link,” “Data,” and “Mario” are likely to have many more results unrelated to that character.)

Characters from the MCU like Captain Marvel (4,993 images), Black Panther (4,395), and Captain America (3,155) are some of the best represented in the dataset. Batman (2,950) and Superman (2,739) are neck and neck. Luke Skywalker (2,240) has more images than Darth Vader (1.717) and Han Solo (1,013). Mickey Mouse barely breaks the top 100 with 520 images.

NSFW Content

Finally, let’s take a brief look at the representation of adult material, another huge difference between Stable Diffusion and any other model. OpenAI rigorously removed sexual/violent content from its training data and blocked potentially NSFW keywords from prompts.

The Stable Diffusion team built a predictor for adult material and assigned every image a NSFW probability score, which you can see in the “punsafe” field in the images table, ranging from 0 to 1. (Warning: Obviously, sorting by that field will show the most NSFW images in the dataset.)

In their announcements of the full LAION-5B dataset, LAION team member Romain Beaumont estimated that about 2.9% of the English-language images were “unsafe,” but in browsing this dataset, it’s not clear how their predictors defined that.

There’s definitely NSFW material in the image dataset, but surprisingly little of it. Only 222 images got a “1” unsafe probability score, indicating 100% confidence that it’s unsafe, about 0.002% of the total images — and those are definitely porn. But nudity seems to be unusual outside of that confidence level: even images with a 0.9999 punsafe score (99.99% confidence) rarely have nudity in them.

It’s plausible that filtering on aesthetic ratings is removing huge amounts of NSFW content from the image dataset, and the full dataset contains much more. Or maybe their definitions of what is “unsafe” are very broad.

More Info

Again, huge thanks to Simon Willison for working with me on this: he did all the heavy lifting of hosting the data. He wrote a detailed post about making the search engine if you want more technical detail. His Datasette project is open-source, extremely flexible, and worth checking out. If you’re interested in playing with this data yourself, you can use the scripts in his GitHub repo to download and import it into a SQLite database.

If you find anything interesting in the data, or have any questions, feel free to drop them in the comments.

Update

On December 20, 2023, LAION took down its LAION-5B and LAION-400M datasets after a new study published by the Stanford Internet Observatory found that it included links to child sexual abuse material. As reported by 404 Media, “The LAION-5B machine learning dataset used by Stable Diffusion and other major AI products has been removed by the organization that created it after a Stanford study found that it contained 3,226 suspected instances of child sexual abuse material, 1,008 of which were externally validated.”

The subset of “aesthetic” images we analyzed was only 2% of the full 2.3 billion image dataset, and of those 12 million images, only 222 images were classified as NSFW. As a result, it’s unlikely any of those links go to CSAM imagery, but because it’s impossible to know with certainty, Simon took the precaution of permanently shuttering the LAION-Aesthetic browser.