Creativity Is a System, Not a Talent
The moat was never the work. It is the machinery that makes the work.
Kees Bakker · Sep 2026 · 12 min read
Ask twenty-five language models for a metaphor about time, fifty times each. You get 1,250 answers and two ideas. Time is a river. Time is a weaver. That is from the paper that won best paper at NeurIPS last year: across seventy-odd models and 26,000 real prompts, different models agree with each other about 80 percent of the time. Wharton ran the human version. People brainstorming toy products with ChatGPT overlapped on 94 percent of their concepts, and nine of them, working alone, named their toy “Build-a-Breeze Castle.”
I run design at an AI company. I use these tools all day and I am not here to tell you they don’t work. They work. That is the problem. The standard take is that AI made execution cheap, so ideas are the scarce input, so be more creative. Every clause is true and the advice is useless, because individual creativity is precisely what AI raises. The cleanest study on this, in Science Advances, gave writers AI ideas and got stories rated more novel and more useful, with the weakest writers gaining most, and the whole set of stories 10.7 percent more similar to each other. Everyone got better. The pool got narrower. The floor rose and the ceiling converged.
So “be more creative” fails on contact. More creativity from the same generator is more samples from the same distribution. The question that matters is who owns a different distribution, and that is not a question about you. It is a question about the organisation you work in.
What actually got cheap
AI did not make ideas cheap. It made generation cheap: plausible candidates, in any medium, on demand, at close to zero marginal cost. The economists had this in 2018, before the tools were any good. When the price of prediction falls, the value of the thing next to it rises, and the thing next to prediction is judgement: knowing which output is worth having. “Having better prediction raises the value of judgment.” Swap generation for prediction and it still reads true.
This has happened before, and the artists did not lose.
In April 1874, thirty painters who had given up on the Salon hung their work in a rented studio at 35 boulevard des Capucines. The studio belonged to Nadar, a photographer. Painting, faced with a machine that reproduced the visible world faster and cheaper than any hand, walked upstairs into the rooms of the machine and changed what painting was for. The Impressionists did not out-render the camera. They stopped competing on rendering.
A century later it happened to type. In 1985 the LaserWriter and PageMaker shipped and setting type stopped being a trade. North America had about 4,000 typesetting firms at the peak; by 1995 a handful were left. The first thing everyone made was ransom notes, eleven fonts on one page. Then the layer above typesetting became a profession. Graphic design with standards and a bar is largely a product of the decade its production step went free.
That is the pattern. When a production step goes to zero, the step disappears and the layer directly above it professionalises. If your work touches generation, the only question worth your time is what the layer above generation is, and whether you own it.
More experiments is not the answer
The reflex is volume. Generate more, test more, ship more, let the market pick. There is a respectable version of this (the equal-odds rule: hit rate stays constant, so hits scale with attempts) and there is the poster version, Ira Glass telling you to close the gap between your taste and your work by doing a lot of work.
Both are true for a person. Neither survives a generator. Equal odds assumes every attempt is an independent draw. When every attempt comes out of a model that agrees with every other model 80 percent of the time, attempts are not independent. You are re-rolling inside a small patch of the space and calling it exploration. A 2019 test of the equal-odds rule found the quantity-to-quality link only holds for people high in openness. The trait that widens the distribution is the trait that makes volume work. Volume is downstream of range. It is not a substitute for it.
Design already has its tell. Someone built a scanner for the signatures of AI-made interfaces (the purple accent, the glass panel, the centred hero in Inter, the component library nobody touched) and ran it over Show HN launches. His original title was the finding: submissions tripled and now mostly have the same vibe-coded look. I can see it from across the room and so can you, and so, increasingly, can the people we are selling to. Merriam-Webster’s word of 2025 was “slop.” Macquarie’s was “AI slop.” The culture named the output of generation without a system above it before most companies noticed they were producing it.
The market has priced it. About half of all new articles published are now AI-generated, a share that has been flat for five quarters. Of the articles that rank on the first page of Google, 14 percent are AI. Of the ones ChatGPT and Perplexity cite, 18 percent. Half the supply, a seventh of the demand. The generator did not fail. It made exactly what it was asked for: competent, legible, interchangeable.
The cleanest experiment in advertising this year came from Ipsos and Syracuse. Ten human-made ads, briefs reverse-engineered, AI versions generated, 3,000 people. Most could not tell which was which. The human ads landed 11 points above benchmark. The AI ads landed 5 below. The title of the study is that the AI ads were good enough, and that’s the problem. Good enough is what gets approved. Good enough is what quietly loses. I have approved good enough. Everyone has. It is the most expensive word in the building.
One honest caveat, because the ad industry’s own data has a confession in it. The IPA found creatively awarded campaigns running twelve times as efficient as non-awarded ones through 2008. By 2018 the multiple was under four and heading to zero. Creativity as the industry measures it, in awards, decoupled from effectiveness a decade before the models arrived. So the moat is not “creativity” in that sense either. It is something more specific and less glamorous.
The moat is process, not people
Hamilton Helmer’s list of durable advantages has seven entries and one of them was built for this. Process Power: routines embedded in an organisation that produce a better product and can only be copied through an extended commitment. The barrier is time, because improvements have to propagate through the whole system and because most of the knowledge is tacit. His example is Toyota, and the HBR study of the Toyota Production System opens on the paradox that matters: almost nobody has copied Toyota successfully, even though Toyota has been extraordinarily open about how it works. Everyone can read the manual. The manual is not the capability.
Buffett has a line that looks like it kills creativity as a moat: “A moat that must be continuously rebuilt will eventually be no moat at all.” Read it again. Creative output must be continuously rebuilt. Every campaign expires, every interface gets cloned, every distinctive look becomes a template in the next model’s training set. The output is, in his phrase, a Roman candle. The machine that produces the output on schedule, at a bar, with a recognisable point of view, does not expire. It compounds. The machine is the moat. The work is the receipt.
Every institution people cite for creativity was designed as a machine. Mervin Kelly built Bell Labs around a 700-foot corridor so a physicist walking to lunch passed every door in the wing. He put theorists next to experimentalists and metallurgists next to engineers on the same project, gave people years instead of quarters, and made it a rule that a senior researcher could not turn a junior away. Ten Nobel Prizes and five Turing Awards came out of that floor plan. Pixar’s Braintrust is a meeting with two rules: everyone in the room has made a film, and the room has no authority. The director owns the decision. Catmull is explicit that removing the power to mandate solutions is what makes people honest. Apple’s industrial design group under Ive was about fifteen people, most of whom had worked together for fifteen to twenty years. None of these is a talent story. They are architecture, cadence and tenure.
Watch where the market is moving creative work. In 2008, 42 percent of ANA member brands ran an in-house agency. In 2023, 82 percent. The stated reasons are cost, agility and brand knowledge, and the plain reading of those reasons is cadence and context: work that has to ship weekly, made by people who hold the whole picture. I have sat on both sides of this. An agency is never your only client. That is not a character flaw; it is the business model, and it means the agency structurally cannot hold enough context to know which questions matter. What stays external is lengthening in tenure and narrowing in scope. The strategic relationship survives. The production relationship comes home. Companies are building the machine whether or not they call it that.
The Authorship Stack
Name the machine so it can be built. Three layers, each resting on the one below, all of it sitting on a commodity.
authorship_stack: taste: the bar # what is good, held by someone who can tell authorship: the asset # a point of view the market recognises without the logo infrastructure: the loop # routines that make both repeatable at cadence generation: commodity # any model, any vendor, any day
Taste is the bar. Not aesthetic preference. The ability to tell, reliably and before the market does, which of a hundred plausible candidates is the one. Rick Rubin was asked on camera what he is paid for, given he barely plays an instrument and cannot work a soundboard. “The confidence that I have in my taste and my ability to express what I feel.” That is the whole job description. Karpathy’s read on automation is the mirror image: the more verifiable a task, the faster it automates, and creative work lags because nobody has written the verifier. Taste is the verifier. It is tacit, which is why nobody has written it down, which is why it is not in the training set. The Stanford data on young workers losing ground in AI-exposed jobs is precise about what got exposed: codified knowledge, the kind in textbooks and procedures. Tacit knowledge is still safe because nobody managed to codify it.
Taste has a known failure mode: it does not scale through one person’s hands. Someone whose job is holding the bar without executing the work used to be called a creative director, or a design manager, and that role is coming back for a specific reason. A generation of builders came up without pre-AI reps. They can generate anything. Someone still has to be able to tell.
Authorship is the asset. Taste selects. Authorship is what it selects toward: a point of view the market recognises without the logo. The ad researchers call the visual version distinctive assets. The IPA and System1 measured the compound version across 4,000 ads and found consistent campaigns produce 27 percent more very large brand effects than inconsistent ones. Authorship is the thing a model cannot originate, because it is defined against the mean rather than sampled from it. It is also what a model can extend once it exists, which is the correct reading of Coca-Cola’s AI Christmas ads. They tested at the top of System1’s scale two years running, not because the model was creative, but because it was handed thirty years of a distinctive asset to re-render. The authorship was already there. Generation extended it.
Infrastructure is the loop. The routines that turn a bar and a point of view into output on a schedule. This is where Helmer’s opacity lives and where almost every company I have seen is thin. The parts are known. A candour mechanism with no authority, so work can be called ugly early; Catmull’s line is that all of Pixar’s films suck at first and the system is built to say so out loud. Psychological safety, which in the original research explained most of the variance in whether teams learned from their own work, and which Google later found underneath every other team dynamic it measured. Attempting before instruction: 53 studies say struggling with a problem before being shown the answer produces better learning. Proximity and tenure, Kelly’s corridor and Ive’s fifteen. And a bar-holder with standing to reject, sitting inside the loop, not at the end of it as sign-off. Sign-off at the end is where good enough gets approved.
Nothing in that list is new. What is new is that the generation step underneath it is free, which means the loop can run at a cadence that used to be impossible, and the three layers above it are now the entire cost structure.
AI-native creative work is its own discipline
This is the claim the “use AI to iterate faster” crowd cannot make, because it means admitting the discipline changed shape.
The evidence on how people actually use these tools is consistent and unflattering. In the Science study on ChatGPT and professional writing, most participants pasted the prompt in and submitted the output lightly edited or untouched. In the BCG study, consultants with GPT-4 produced work rated over 40 percent better on tasks inside the model’s competence, and were 19 points less likely to be right on tasks outside it, because they trusted the tool the same way in both cases. METR’s trial found experienced developers 19 percent slower with early-2025 tools while believing, afterwards, that they had been 20 percent faster. The tool did not fail. The operator’s model of the tool did. I have done all three of these things. So have you.
The people who won in the BCG study had a shape. Some divided tasks and delegated cleanly; some integrated the model into every step and verified continuously. Both are descriptions of a loop with a human verifier inside it, and that loop is the discipline. Generate, verify, at cadence, with taste as the verifier and authorship as the target. A survey of 900 designers this year has 91 percent using AI weekly and 62 percent naming inconsistent output as their biggest problem. Inconsistency is not a tooling defect. It is what a commodity generator does when there is no stack above it.
So the discipline is not prompting and it is not iteration speed. It is running the Authorship Stack against a generator: holding the bar tighter because candidates are cheaper, defining the point of view harder because the default is the mean, and building the loop on purpose because the generator will run without one and deliver slop on time.
The objection, and why it is right
The strongest case against all of this is that taste is not a moat, it is alpha: a decaying edge, valuable only relative to a baseline, and the baseline is rising faster than the taste-is-the-moat crowd wants to admit. This year’s exceptional is next year’s default. The model learns the cues.
That is right about individuals. Individual taste is alpha. It decays and the model absorbs it. That is the argument for the stack, not against it. A person’s taste is a snapshot. An organisation’s loop regenerates taste faster than the baseline rises, because the loop is what produced the taste in the first place. Toyota’s practices have been public for decades and the advantage is older than that. The output decays. The machine that makes the next output is the durable thing. Helmer calls this hysteresis. Buffett calls it enduring.
Which brings it back to the river. Every model, asked for a metaphor about time, reaches for the same one, because the river is the mean of everything ever written about time. The mean is now free, in every medium, on demand. The moat was never the work. It is the organisation that can be relied on, week after week, to arrive somewhere the mean does not.