Chimpanzees, Computers, and Shakespeare
Why the Infinite Monkey Theorem Is Misleading, AI and What It Really Teaches Us About Innovation
Imagine a room filled with chimpanzees, each hunched over a typewriter. They tap away randomly, day after day, year after year. The famous thought experiment - the infinite monkey theorem. The idea tells us that, given infinite time or infinite monkeys, they would eventually produce the complete works of William Shakespeare by pure chance.
It’s a beautiful, mind-bending, fun idea about probability. It’s also, as recent mathematics has shown, profoundly misleading in any universe we actually inhabit.
The Maths Bites Back
A 2024 study (https://www.sciencedirect.com/science/article/pii/S2773186324001014) by mathematicians Stephen Woodcock and Jay Falletta put hard numbers on the fantasy. Even if every one of Earth’s roughly 200,000 chimpanzees typed one key per second on a 30-key keyboard for the entire remaining lifespan of the universe (the point at which heat death sets in), they would almost certainly never produce Shakespeare’s collected works (nearly 885,000 words of poetry, drama, and insight).
A single chimp, in those settings, has only a ~5% chance of randomly typing even the word “bananas” in its lifetime. The odds of producing even a short coherent sentence collapse into absurdity. The full canon would require impossible timescales.
The theorem survives as elegant mathematics, but as a practical intuition about creativity or discovery, it fails. Pure randomness at biological scales and finite time simply doesn’t cut it.
The Real Mistake Isn’t the Monkeys, It’s the Search
Here’s where the story gets interesting, and where it connects directly to how real progress happens in science, technology, and pharmaceutical innovation.
It’s tempting to say we’ve simply swapped in better monkeys: bigger, faster, digital ones. But that framing quietly repeats the same error the maths just challenged. Computers, and the AI systems running on them, aren’t chimpanzees at all, however fast. A chimp at a keyboard carries no information from one keystroke to the next; every letter is a fresh coin flip, which is exactly why the odds never improve, no matter how long it runs. An AI model does the opposite. Each token, each proposed molecule, each note in a melody is chosen *conditional* on everything that came before it and everything the model has absorbed about how language, chemistry, or music tends to behave. That conditioning is what collapses an astronomical search space down to a tractable one - not superior typing speed, but the elimination of almost every option that a blind search would have wasted time on.
That distinction matters more than it sounds. A model trained on the accumulated structure of English, Elizabethan drama, or protein-ligand binding isn’t sampling uniformly from the space of possible outputs - it’s sampling from a probability distribution shaped by everything humanity already knows about what tends to work. Large language models can now produce a decent, certainly coherent Shakespearean sonnet, a passable pastiche of a scene from *Hamlet*, or a plausible new plot structure within minutes. Not through brute-force randomness, but by having internalised the deep structures of form, psychology, and language, and then navigating that space with purpose. You can’t do it because you can’t internalize all of that - you can write something that’s passable, but you miss some of the things that matter.
Scale this with parallel compute, active learning, and rapid iterative feedback, and you move from “impossible by chance” to “routinely achievable with directed effort.”
The monkeys were never the point. The search strategy is.
The Real Lesson for Discovery
This isn’t just a thought experiment about literature. It’s a profound analogy for how innovation actually works - especially in fields like drug discovery, where the search space is astronomically larger than all of Shakespeare’s works combined.
Chemical space contains an estimated 10⁶⁰ or more drug-like molecules. Random screening of compounds is monkey-typing at planetary scale: slow, expensive, and mostly fruitless for complex targets.
The asymmetric advantage comes from replacing random search with intelligent, data-driven search: machine learning models that predict molecular properties before a single compound is synthesised, generative chemistry platforms that propose new molecular scaffolds optimised for binding, safety, and synthesisability, and closed-loop systems that design, simulate, test with real assay data, and learn from the result - active learning and Bayesian optimisation doing the work of deciding what to try next, rather than reinforcement learning tuned to human preference, which is a different tool built for a different job.
This isn’t hypothetical. Insilico Medicine’s rentosertib - a TNIK inhibitor for idiopathic pulmonary fibrosis, a disease with no treatments that reverse its progression - had its target identified by an AI platform trawling multi-omics data, and its molecular structure designed by a generative chemistry engine. It entered Phase III trials in July 2026. That’s the pipeline this post is actually about: not a thought experiment, a drug in human trials right now.
What once required screening millions of compounds over years can now be guided toward promising regions of chemical space in days or weeks. We’re not waiting for infinite random trials. We’re building better searchers.
This is the core of asymmetric learning: the organisations and teams that develop superior ways of navigating vast possibility spaces - through better data, better models, better feedback loops - pull ahead dramatically. The difference between random exploration and directed intelligence isn’t incremental. It’s transformative.
The Catch: Nobody Knows (Knew) What Shakespeare Looks Like in Advance
There’s a limit to the monkey analogy worth naming honestly, because it’s the difference that actually matters for anyone doing this work.
The monkeys are searching for a known, hard target. We have the complete works of Shakespeare already; success is an exact match against a fixed string. Drug discovery doesn’t work that way, and neither does genuine creative work. It’s a question no-one has the answer to - yet. Nobody has the answer key for the molecule that will safely treat a disease with no existing treatment, or for the novel that hasn’t been written yet. The objective itself is uncertain, contested, and often redefined mid-search, as new safety data or new critical judgement reshapes what “success” even means.
That’s a harder problem than typing towards Shakespeare, in one sense - and an easier one in another. You can’t verify progress against a known answer, but you also aren’t constrained to reproduce something that already exists. The real work of directed search isn’t just navigating towards a target efficiently. It’s deciding, continuously, what the target should be.
Beyond Parroting: Toward Genuine Novelty
Critics will rightly note that current AI systems are sophisticated remixers and predictors rather than true originators *ex nihilo*. They still benefit enormously from human curation, taste, and judgement.
But it’s worth asking how much of a contrast that really is with Shakespeare himself. He didn’t emerge from a probability distribution, but he wasn’t working from nothing either. He was a highly trained pattern-matcher, with Plutarch, Holinshed, Italian novellas in his reading list (unlike mine, he put them on his bookshelves once read), and the conventions of the sonnet and the five-act structure, iterating within tight formal constraints and a specific culture. Directed search, in other words, isn’t unique to machines - it’s arguably a reasonable description of how human creativity has always worked, at a much smaller and slower scale. It’s writing a song when someone tells you the key, the time signature, and the length and style. The interesting frontier isn’t “AI versus human insight” as two opposed categories. It’s how the two forms of directed search compose: AI proposing, humans steering and selecting on taste and judgement that no model yet has, AI iterating again.
In pharma, this already looks like AI proposing candidates that medicinal chemists would never have reached for, followed by expert refinement and experimental validation. Each cycle narrows the gap between what’s proposed and what actually works.
The Cost of Getting It Wrong Quietly
There’s a failure mode worth naming, because it’s asymmetric in a way that’s easy to miss.
A chimpanzee’s garbled text is obviously garbled - nobody mistakes even that banana example, “xjqpz banana zzzqx”, for a sonnet. An incorrect answer from a directed search doesn’t announce itself in the same way. A molecule can look right on paper - clean binding predictions, plausible synthesis route - and still fail in a way that only shows up in a clinical trial. A generated passage can scan perfectly and say nothing. The risk of intelligent search isn’t that it fails more often than random search; it’s that when it fails, it can fail *convincingly*, in a form that passes casual inspection precisely because it was optimised to look right.
That’s a genuinely asymmetric cost, and it’s the reason validation loops - real assay data, real experimental feedback, real human judgement in the loop - aren’t a bolt-on to directed search. They’re what keeps the asymmetry working in your favour instead of against you.
What This Means Practically
If you work in R&D, strategy, or innovation:
- Stop thinking in terms of “more shots on goal” through brute-force randomness.
- Start thinking in terms of better aim through superior models and feedback, and be precise about which technique is actually doing the aiming - active learning and Bayesian optimisation for scientific search, not tools built for aligning conversational tone.
- Invest in the infrastructure that turns data into directed intelligence: high-quality datasets, simulation capabilities, rapid experimental loops, and human-AI collaboration workflows.
- Build in validation that catches confidently-wrong answers, not just obviously-wrong ones - the failure mode that matters most is the plausible one.
- Measure not just activity (compounds screened, models trained) but asymmetric progress - how much faster or more effectively you’re navigating the space that matters.
The infinite monkey theorem was always a story about limits. The real story today is about transcending those limits through intelligence and design - not by building a better monkey, but by leaving the whole metaphor behind.
We don’t need every chimpanzee on Earth typing forever. We need better-directed search, pointed at problems worth solving, with humans still deciding what “worth solving” means.
That’s how the next breakthroughs - in literature, in science, in medicine - will actually arrive.




