Skip to content
meirlabs
Back to all articles
Meir RosenscheinJuly 21, 20265 min read

How Do You Make Brand Assets and Mascots With AI Consistent Across a Campaign?

TL;DR

Stop describing the character in words; attach it as a picture. Every generation carries the actual anchor image plus one instruction, same character, just re-posed, so identity lives in the pixels, not the prose. From that anchor you build a kit of approved poses, pick one per piece, and composite headlines in HTML on top. On soclever's mascot, Lottie, this held one axolotl steady across 48 upload-ready ad files.Launch offer: Early clients get 50% off their first build, so your real cost is about half these figures. Book a free AI plan to lock it in.

By the numbers
$0.067Per generated image, so a full six-ad art refresh runs about $0.40 plus a retry or two
48Upload-ready ad files from one reference kit: 6 courses, 7 native canvases, one render doing double duty
0Words baked into the art; every headline is HTML on top of the image

Most founders picture the mascot problem as a prompt-writing problem. Get the description precise enough, they think, and the model will draw the same character every time. Wrong frame. "A peach-gold axolotl with feathery gill-plumes" has a thousand different faces, and the model picks a new one on every call. The fix is not a better sentence. It is refusing to describe the character in words at all.

The canonical anchor for Lottie, soclever's mascot: a peach-gold axolotl drawn front-on, the one image every later generation is conditioned against
This one file is the character. Every ad below is generated by attaching this picture, never by re-typing what Lottie looks like.

Why does the same prompt give you a different character every time?

Because a text description asks the model to re-invent the character from scratch on every generation, and every re-invention is a fresh chance for the model's own defaults to leak in. Proportions shift, the palette warms or cools, the styling drifts. You are not requesting the same axolotl twice; you are requesting the model's best guess at your words twice, and its guesses are not identical. Words can pin a genre ("friendly cartoon axolotl"). They cannot pin an individual. For that you need the individual itself in the room.

How do you lock the character in place?

You attach the actual anchor image as part of the request, before the text, and let it carry the identity. The prompt then only says what is different this time. On soclever, the internal rulebook that governs every Lottie generation states it as numbered law:

  • The reference image is ALWAYS attached. Identity lives in the picture, not the prose.
  • Appearance is NEVER re-typed in text. "Text competing with the reference is the #1 cause of drift."
  • One state change per call. Never combine a pose, a prop, and a mood change at once.

The single load-bearing instruction in the prompt is this: "Keep the mascot's identity (colors, face, proportions) EXACTLY consistent with the attached reference image. It is the same character, just re-posed." Notice what it refuses to do. It never re-states Lottie's colors or shape. It asserts only that whatever is in the attached picture must stay, then hands the model the one thing allowed to change: the pose. The image below was generated exactly this way, from the anchor above.

Lottie re-posed as the patient coach: sitting low, arms open, unmistakably the same axolotl as the anchor
Same face, same gill-fronds, same proportions as the anchor, new pose. The model copied identity from the attached pixels and only invented the sitting posture.

How do you get variety without it drifting?

You build a kit. From the one anchor you generate a set of approved moods and personas, a coach, a professor, a scout, and each of those becomes a reference in its own right, but every one traces back to the same locked identity. Then you pick the reference per piece to pre-load the right tone before the model reads a word of the brief. On soclever's six-course ad campaign the mapping was editorial, not mechanical:

"Teach a kid who's given up"   ->  coach pose    (patient, sitting low)
"Grow your tutoring business"  ->  focused pose
"Understand medical results"   ->  encouraging pose
"Spot fake news"               ->  scout pose
"Master Excel"                 ->  professor pose (glasses, book)
"Build a portfolio"            ->  plain 3/4 anchor

That is chosen variation, not drift. Drift is what happens when an unreferenced image is generated from text alone and the model's defaults creep in. Picking the coach pose over the professor pose is a deliberate creative choice among pre-approved options; the character underneath never changes. Holding it there is a set of guardrails, and every guardrail is a scar, each one written after a real failure:

  • Gill-fronds drifting crimson. Warm scenes pulled Lottie's gold gill-fronds toward red to match the palette. Fix: a standing rule that the fronds "must stay the mascot's natural warm gold/peach/orange exactly as in the reference image, NEVER crimson, red, or recolored to match the scene palette."
  • The mascot getting distorted when it needed to shrink. The model simplified proportions instead of scaling. Fix: "If the mascot needs to appear smaller in the frame, scale it down; never recolor it."
  • A stray human wandering into the scene. Fix: "NO real human figures, faces, or photographic people anywhere," so the axolotl stays the only face in the whole brand.

Why keep every word out of the art?

Because image models cannot render legible type, and copy baked into the image can never be re-laid-out or re-edited without paying for a whole new generation. So the art is generated with a blanket ban: "ABSOLUTELY NO text, words, letters, numbers, labels, captions, signage, or speech bubbles anywhere in the image." Every headline, chip, and wordmark is composited on top afterward as real HTML text, rendered crisp at any size. That split buys three things: pixel-perfect type in any language, one source illustration re-laid-out natively for every format, and the freedom to reword a headline without regenerating a single pixel of art.

The finished ad: the same coach-pose Lottie, now with the headline and call-to-action composited in crisp HTML over the untouched illustration
One AI illustration, zero baked-in text. The headline, the button, and the wordmark are all live HTML rendered on top, so the same art fills seven platform sizes.

And each platform size is rendered natively, not cropped. LinkedIn wants 1200 by 627, X wants 1200 by 675, Meta wants 1200 by 628, all "landscape with a text column" but off by a handful of pixels. Crop one master image to hit all three and you clip the mascot or the headline differently on each; instead one square illustration is re-laid-out for seven exact canvases. A gotcha worth stealing: when you drop square art into a non-square frame, pin the container to a fixed aspect ratio, or you get white letterbox bands above and below the art.

What does this not solve?

It is honest to say what it does not do. Reference conditioning holds a still character steady; it does not give you animation or a rigged 3D model. It does not guarantee pixel-for-pixel identity either, expressions and angles still vary slightly, which is exactly why a human approves each pose into the kit before it becomes a reference. And it only works on a model that accepts an image as input (soclever runs Gemini's image model at about $0.067 a picture). What you get is not a clone on every call. It is the same recognizable character, every time, for roughly the price of a rounding error. Locking a character to a reference kit is the same kind of one-workflow build as any other AI adoption project: scope it tight, ship it, and let the pipeline pay for itself on every campaign after.

Sources