Neanderthal
A thought experiment in embedding visuals directly in streaming Markdown prose.
View on githubModern writing chose speed over visible meaning
26 reusable letters can write down anything anyone says, and nobody has to draw a new picture for every object in the world. Writing got fast, cheap to copy, and able to carry any subject at all. What it gave up was visibility. A word can name a thing perfectly well without ever showing it to you.
So the pictures moved out of the sentence. They became citations, thumbnails, and links. If you want to connect a word to the thing it names, you stop reading, scroll, click, and open a tab for exploration. What's left inside the line is thin. Arrows, currency signs, math notation, interface icons, emoji. Emoji carry tone. A citation tells you evidence exists somewhere else. Neither one lets you see the thing at the moment its name shows up.
Modern text didn't forget images. It pushed them out of the reading flow. It wasn't always like this.
The evolution of pictographic script

Before we wrote words, we drew. Neanderthals pressed deliberate finger marks into a French cave wall more than 57,000 years ago. Artists in Sulawesi painted a hunting scene 51,200 years ago, the earliest known picture that tells a story. None of it is writing and none of it can be read as sentences, but it does prove one thing: humans held memory in pictures long before scripts held language.
Then pictures learned to speak. For most of writing's history, images weren't placed next to the text. They were the text.
Egyptian hieroglyphs were a full writing system, not a gallery of drawings. A sign could carry a sound, a whole word, or a category of meaning. The signs called determinatives, or classifiers, sat at the end of a word and were never spoken aloud. They told the reader what kind of thing the word was: a person, a place, an action, an idea. Letters carried the sound and the picture carried the sense. Scribes also packed signs into balanced square groups and turned birds and figures to face the start of the line, so a page of text held together as an image.

Maya glyphs got to the same place by another road. Scribes combined word signs with syllable signs inside a roughly square block, one large main sign with smaller ones clipped to its edges. The same word could be written as a single picture, spelled out in syllables, or both at once. B'alam, jaguar, could be a jaguar's head or b'a-la-ma. Each block was compact, visual, and linguistic in one stroke. You could look at it and read it.

They weren't alone.
- Cuneiform started around 3300 BCE in Uruk as simplified drawings of real things, grain and cattle and tools, before speed and the stylus flattened them into abstract wedges.
- Aztec and Mixtec codices were painted, not spelled. The images carried the story and the glyphs supplied names, places, and dates.
- Naxi Dongba in southwestern China is the last picture writing still in use, held by a few dozen elderly priests. You can still see pictographic script roadsigns in that part of China.
Across continents that never met, scribes landed on the same instinct. A picture can do work inside a sentence that a word can't do as well.
Put the picture where the meaning happens
Take a sentence:
Neanderthals shaped Mousterian stone tools to produce sharp cutting flakes.
If you've never seen a Mousterian tool, that name gives you nothing to hold on to. You either take the sentence on trust or leave the page to go find out.
Neanderthals shaped Mousterian stone tools to produce sharp cutting flakes.
It isn't there just as a decoration. It's a piece of visual evidence, always there to explore when you are curious.
This is easier to try than to describe, so I made a playground, try below.
How do deep-sea organisms create cold light to survive in the midnight zone?
Why this thought experiment
Research on multimedia learning is direct about this. People understand words and pictures better when the two sit together, which is called the spatial contiguity principle. It states that people learn better when corresponding words and pictures are placed close together rather than far apart on a page or screen.
The same research warns you the other way too. Interesting but irrelevant images, the ones researchers call seductive details, reliably make comprehension worse. Readers keep the decoration and lose the argument.
So the goal was never a picture on every line. It's zero distance between an unfamiliar idea and the one image that resolves it: an unfamiliar object, a mechanism, a spatial structure, a comparison, a piece of primary evidence. If the picture doesn't make the nearby idea easier to understand, it doesn't belong there.
None of this is a rollback to hieroglyphs. Words should keep doing what only words do, which is grammar, sequence, nuance, abstraction, argument. Images should do what only images do.
What's new is that AI can join the two at the moment of explanation. It can spot the visual concept, find a trustworthy image, place it beside the right phrase, and keep the source attached.
Neanderthal is a thought experiment. Ancient scripts kept the picture inside the sentence and modern writing pushed it out. This puts it back, with AI placing the image next to the word that needs it.
Right now the capsule engine pulls its visuals from Wikimedia and DuckDuckGo, leaning on Wikimedia for scientific and encyclopedic subjects and DuckDuckGo for people and culture. The experiment runs on general knowledge, so feel free to fork it and customise it for whatever you are working on.