The learned aesthetics of YouTube channels
How ChatGPT interprets what cryptocurrency and thirst trap channels look like
After China and India, the third country with the largest population is MrBeast’s YouTube channel: 509 million and counting. If every subscriber gave him a penny he would make in one instant more than a hundred well-paid computer science PhDs make in a year.
The legend of MrBeast goes something like this:
He was a teenager making videos without his parents knowing.
Then, he and some friends spent 1,000 days on 18 hour Skype calls reverse engineering the YouTube algorithm.
They figured out how each detail (colors, camera angles, text, thumbnail design, voice, audio, music, etc.) impacted metrics like clickthrough-rate (CTR).
His channel became ultra viral.
Opinions of MrBeast vary. Some commend his philantropic Africa-well-building pursuits. Others condemn his evil-villain-esque real-life Squid Games. But what most people agree on is that MrBeast is one of the most emblematic figures of the “quantitative YouTube era.” This era, which I just named 10 seconds ago, is the era responsible for the homogenization of YouTube content. At some point we stopped having legendarily quirky videos. Instead, YouTube became a hellscape of templated thumbnails and titles. Secrets about the algorithm began leaking in online blogs: put your face on the thumbnail, point at things, and if you lack all self-respect make a really exaggerated face.
Unfortunately, these cringey thumbnails are what the Machine™ has learned to promote. Or, is it humans who preferred the thumbnails first and the machines are just a mirror that amplifies things? Potato tomato for the purposes of this article. Regardless of how it came to be, the reality we live in is one of cultural convergence.
This reality is likely the fault of a group of PhDs. A group of someones built a crude algorithm that ingests an image––like a thumbnail––and converts it into numbers. These numbers encode several things, like the colors in your image (think RGB values) and can even encode shapes like curves and sharp edges. This process is called “feature extraction” in ML parlance and when the <<algorithm>> chews all these numbers what is trying to do is to “learn a representation.” Simply, the machine learns what numbers corresponds to what shapes and what shapes make a pointy finger or an agape mouth. This process is what first allowed entities like banks and post offices to create software that read the numbers that people handwrote on checks and letters. A few demos available here (explainer) and here.
But those are not the PhDs that are to blame for how the aforementioned technology was abused. They created gunpowder. Then came the people who engineered the gun and the bullets. The truly devilish idea came like this: the above system converts a YouTube thumbnail into numbers. Millions and billions of people click on videos every day, often by browsing thumbnails. What if, for every click, we gave more “weight” to certain thumbnails (a.k.a. groups of numbers). If people click on videos that have red thumbnails a lot more, then why not serve those videos more often? Now, replace “red” for “pointy finger, open mouth, color correction, etc.” People who were not making videos with those types of thumbnails were getting less clicks and less views and so they must comply, lest they get stuck forever with 71 subscribers, like yours truly.
It seems easy (and entirely reasonable) to blame MrBeast for all of this. But in reality, this was bound to happen as people with marketing degrees who did not fail Algebra I entered the online arena. Quantitative marketing was already taking off and Google was one of the biggest preachers of this terrible gospel. They began A/B testing everything with religious zeal circa 2006. Their “41 shades of blue” became infamous and marked the replacement of designers by data chez Google. “Last quarter using 45 106 237 on Gmail led us to a 0.04% increase in app time spent.” Their lead visual designer resigned over this philosophy.
All of the above really is just background knowledge for the unsettling realization I had 2 days ago. I was making slides for an upcoming presentation. I needed YouTube screenshots of channels with certain characteristics. They needn’t be real (in fact, fake would be preferred) so I asked ChatGPT to generate them. Only after the deed was done I stopped to ponder the implications.
People have long feared the homogenization of culture that will take place when more and more people use AI tools. This is because AI generates text and image in the same way that generic brands copy distinct products. A bottle of Heinz ketchup has distinct elements: shape, color, font, composition, etc. By identifying the most important elements, you distill the essence of a “bottle of ketchup”, and thus you create a convincing copy.

So, what is the essence of a cryptocurrency channel, or a channel that shares geopolitical misinformation? Here is one through ChatGPT’s eyes.
AI models learn representations by observing millions of images of things. After looking at millions of images of ketchup bottles or YouTube channel screenshots, it can “learn the representation” of these things: it can distill their essence. Reward algorithms created an environment where aesthetics converged into homogenous styles. Generative AI models have now given us the opportunity to visualize what the mode/common denominator of these styles is. Here are some other fun ones:


Each of these screenshots simultaneously reveals the essence of the depicted channel as well as the interests of the watcher (see the sidebar)1. For example, the misinformation channel is particularly interesting as it reveals common conspiratorial themes and characters. The thirst traps channel indicates that the viewer is a gen z boy based on the subscriptions on the sidebar.
The corollary of this article is that this AI experiment lets you experiment the compression of culture in real time. You can check for yourself how mainstream certain tastes, styles, and aesthetics really are.
I will let you guess the prompts that generated the screenshots above and let you ponder the implications of these results ( ͡° ͜ʖ ͡°).
Appendix: Other (Artificial) Screenshots


An interesting thing to notice is that the main content is all artificially generated. The depicted channels don’t exist. However, all the channels from the sidebar I checked (under “subscriptions”) are real. The technical reason this happens is interesting but slightly off-topic––a convo for the comment section, if needed. Suffices to say that the channels shown on the sidebar are what the LLM believes best aligns with the target audience of the depicted channel.









