I started with some simple product screens and the name Really Great written in italics. My first eight attempts to build an identity system with AI were worse than that. I quickly learned the models' default taste is generic: prompt them with adjectives and you get the blandest plausible version of bold, warm, and confident.
What fixed it was changing the process.
When that didn't work, I found a blog post by Lucas Crespo: vague words make vague work. I went and found 32 images that evoked the feeling I wanted people to have in the product. The AI's job flipped from inventing a brand out of my adjectives to naming what my references had in common. It sorted the pile into three named directions, I picked two, and had it write the feeling down as a short brief. Every asset after that was generated from references plus that brief. This was the biggest unlock of the project. My taste lives in images, and until that point I had fed it zero.
There was no old brand to defend. The incumbent was that typed-italics placeholder on our existing screens. It still entered every round as a live candidate, and for the first week it kept winning.
“I would go with what we have over either of these.”Me, round one, verbatim from the log
Keeping "no change" on the table as a real outcome meant I couldn't talk myself into shipping something new that lost to a typed-out placeholder.
Every dead direction got named for the neighborhood it drifted into. Royal blue and gold looked like Laker colors. Slate on cream failed a side-by-side gut check against Greenhouse. A meadow palette went full John Deere. A lime accent read as Mountain Dew. Seven color systems died this way in four days. The names are what made the kills useful: each one became a constraint the next round could be checked against, instead of a vague feeling that something was off.
Pages of options side by side told me nothing. What I could actually react to was one concept executed at full conviction across believable surfaces, mark and type and color and copy all committed, with the reasoning printed on the page.
“This page was the most helpful thing we've done.”Me, on the first full-conviction system page
Hedged middles produce zero signal per round. When exploring, execute the bet at full strength and name its risks right on the page.
Logo candidates got composited into our actual product nav at actual size, and most died there. The winning color system got built onto a live branch of our product overnight, and I judged it the next morning on real data at working density. One direction that looked gorgeous on big clean poster cards turned to mud on a dense screen. My eye also caught drift that the automated checks missed. I'm not a designer and the process never needed me to be one. It needed my reaction to real artifacts at real size.
Every reaction I gave got written into a running process log word for word, more than 900 lines by the end, dictation typos included. The system that finally won was assembled from fragments my eye had approved weeks apart, across rounds that otherwise died: one model's sky treatment, the other model's dawn light, a script wordmark from the first week. The two best fragments came out of a blind head-to-head, the same brief given to two models with neither allowed to see the other's page until I'd judged both. Recombination is only possible if the fragments got written down when they happened.
“This feels connected but different… feels like we've finally broken out of that loop.”Me, the moment it finally worked
The brand book itself went through nine versions. Before the second big pass, I wrote a success contract: what was locked, what was allowed to improve, and the rule that a more complete version is not automatically a better one. Ten scoped improvements, worked one at a time, each prototyped and compared head to head against the incumbent version before it could land. Without a written finish line, AI iteration never ends, because there is always one more plausible improvement.
The most useful thing I brought was specific feedback and a refusal to settle. Halfway through the color search I had versions I would have shipped, and I said so in the log: they looked like cleaner versions of what we already had. Shippable and right turned out to be different events. The first version of the brand book got graded barely a C minus, with everything it needed already sitting in the reference folder. Every hard grade came with specifics about what was wrong, and those specifics are what the next round was built from. Being easy to please gets you the generic thing the model handed you first.
Seven color systems died the same two deaths, a claimed neighborhood or a color with no job in the system, while I kept asking for the next variant. When the same kill reason shows up twice, the generator is the bug. Change how you're producing candidates, and stop asking for one more.
When the AI dressed candidates in made-up product screens, I caught myself reacting to fake features and fake copy instead of the identity. Real captures labeled with their build date, or honestly abstract demonstrations. Nothing in between produces signal.
My first visual direction won a five-way bake-off on beautiful standalone pages, then collapsed the first time it touched real product mocks. Working density on day one would have killed it on day one.
Hand-coding wordmark ideas in CSS from written descriptions capped the quality at competent CSS, and lost to the typed-italics placeholder eight times. Discovery happens in image generation. Code comes back in for production.
They redraw, they can't copy. Every generative cleanup of the winning wordmark came back subtly wrong. A mechanical trace of the original image fixed it in one pass and became the standard path from picked image to production vector.
The calendar says eight weeks, but most of that was the project sitting untouched, and even the intense stretch was a side thing. I was doing tons of other work the whole time, coding and content, and I'd turn to the brand when I wanted a creative break: review a round, react, send the next one off to generate. Claude Code and Codex did the building, Midjourney generated the world imagery, and the final mark was traced mechanically because image models redraw things, they can't copy them. Days before the color breakthrough, this went into the log:
“I am really frustrating and disappointed in where this is all ending up. I am not sure if I should just settle for where we started.”Me, dictated, typo preserved
The breakthrough came days later, and everything went from there: the winning system locked, merged into the product, and the brand book built on top of it.
The AI produced every artifact in that book and made none of the decisions. Killing round after round, naming why, and recognizing the real thing when it finally showed up stayed my job the entire way through. I still don't fully understand why round twenty-two worked when the fifteen before it didn't. I knew it when I saw it. The whole process existed to put enough real things in front of my eyes for that to happen.