AI creative generation: what separates the teams who publish

RESEARCH LEAD | TRIPLE WHALE | DECEMBER 2025 TO FEBRUARY 2026

Triple Whale was building AI creative generation into Moby. Ecommerce brands run performance ads on Meta, and the ceiling on how many they can run is how fast creative gets made. The team's leading assumption was that bulk generation was the highest-value thing to build. Generating an image was already possible. The question was why brands would not put spend behind the output.

Moby is Triple Whale's AI experience. Creative generation was being built into it as a studio where you describe what you want, get images back, turn them into ads, and publish to Meta.

My job was to work out what the studio had to do before anyone would publish what it made. I led the research from discovery through early concept testing and wrote the product requirements the designs were built against.

14Interviews 3Publishing segments 6Journeys mapped 4Conditions to publish

Approach

Recruiting on publishing behavior

Fourteen interviews across two phases with performance marketers, creative leads, founder-operators, and agency leads. The mix that mattered was what people actually did with AI creative: publishing it, generating and never publishing, or not using it at all. Role, brand maturity, and creative approach were secondary cuts. The comparison across those three groups produced the central finding.

Reading usage before asking about it

Generation dashboards and Moby transcripts, to see what people were doing in the product and to pre-identify who sat in which segment before fieldwork started.

Competitive analysis

What other creative tools did well, where they broke down, and what customers had already tried and abandoned.

A decision point between phases

Discovery closed with a review with product, design, and leadership to make a go, pivot, or adjust call before any investment in phase two. The risk being managed was real: discovery could have shown the direction did not fit how creative gets made.

Phase two, testing the direction

Concept testing on the designs, covering brand controls, the editor, entry points from Creative Analysis and Moby chat, and copy and image together.

How the team followed it

The designer and PM observed most interviews. We debriefed after sessions and synthesized together, so findings were shared as they formed rather than written up and handed over. Weekly syncs and a running project channel carried the work between sessions.


What we found

Five themes, each mapping to one of four conditions the product had to meet before anything would ship.

Publishability is the bar

CLEARED

What people adopt is a way to get ads live. The threshold is whether the output clears approval. Participants who published succeeded by scoping AI to bounded changes: color swaps on handbags, pose variations, short clips. Those became one participant's best performing ads of the month. Another team found AI winners in testing and killed them anyway, because leadership said they looked fake. The pattern held across the sample. People who pushed AI toward net-new hero creative published the least.

What people check for is credibility. Beauty brands need shade and packaging accuracy. Food brands cannot get meat to look real. A Hungarian CMO described AI copy that was grammatically correct and that nobody would ever say, and faces that read as unfamiliar in her market while passing for US audiences. One participant used an AI truck image he thought cleared the bar, then watched two of ten user testers flag it. The question people are asking is whether a human designer would have made that choice.

“AI can look a bit chilly. The texture will be too perfect.”

Founder and CEO, personal care brand

Creation starts from a reference

CONNECTED

Once teams know what is working, they show it rather than describe it. A past winner, a competitor's ad, a saved example. Teams keep collections in Notion, Foreplay, and Pinterest specifically as starting points. One participant took a competitor's ad and swapped in her product, headline, and call to action. Another briefs with a screenshot and the instruction to use that format. “Adapt this” was closer to the real request than “generate something new.”

The concept is the unit of work

CONTROLLED

Every ad starts as an angle. The headline, the hook, the product treatment, and the call to action get conceived together, and people brief them that way. Tools that return an image and leave the copy elsewhere force people to reassemble something they already thought of as whole. One participant said downloading the image to layer text on herself would be slower than her current process.

Small edits have to stay small

CONTROLLED

The failure that ended trials was a small edit triggering full regeneration. Ask for a different background, get a different product. Fix one thing, create another. Participants described wanting to circle an element and change only that. Chat prompting alone could not get them there.

“We should be able to circle it somehow and change the specific detail without touching other things.”

Senior paid media specialist

Cadence is the end state

CONTINUOUS

One-off generation is not the goal. Teams want a target volume every week without quality falling off, and they cannot hit it. One participant runs four parallel creative sources and still cannot get what he needs when he needs it. Another waits two weeks on his creative team, by which point the month is over. The constraint is publishable volume. Generation speed was never the bottleneck.


Six journeys, one studio

Discovery produced six ways creative production gets triggered. Three are reactive and three are planned. Each carried its own capability requirements.

Select a journey to see how it runs and what the product has to provide at each stage.

The journeys became the instrument the designs were tested against. I outlined them and wrote the requirements. The designer built the capabilities. The PM held the feasibility line on what could actually be built. Then the designer and I walked each capability back through the journeys to check whether someone could get from a performance signal to a published ad without falling out of the flow. That pass is where the designs came together.


What we recommended

The first recommendation was to not build bulk generation yet.

Bulk was the team's leading hypothesis, and the evidence pointed elsewhere. Teams could not use volume they could not publish, and the thing stopping them from publishing was control. Ten mediocre variants create ten review problems. The recommendation was to sequence: build targeted, controlled flows for static Meta ads first, get the publish rate up, then expand format and volume.

Output quality depends on the model, and the model was not something the team controlled. What the team controlled was what went into it and what happened after. The rest of the requirements concentrated there.

  • Give the model more to work with. Brand kit, voice, naming conventions, product imagery, reference images, and dimensions feed generation up front rather than getting corrected afterward.

  • Give people precise control over the output. Select an area, change that element, leave the rest intact. This carried the most weight, because full regeneration was the thing that ended trials.

  • Generate the concept, not the asset. Copy comes back with the image, tailored per variant.

  • Put the entry point where the analysis happens. Generation triggered from Creative Analysis carries performance context into the brief, which is the handoff people were doing by hand.

  • Make the proven format repeatable. A template holds what stays constant, and an automation runs it on a cadence, so teams stop rebuilding context every session.

  • Measure success on the share of generated creative that gets published. Volume was already easy.

The approval gate went in as a flagged gap. Every brand has someone who signs off, so direct-to-publish was going to fail. Permission structures limited how far approver handoffs could be built into the platform, so the recommendation named the problem and left it open rather than proposing a partial fix.


What it became

Five screens from the designs at engineering handoff, in the order the recommendations run.

Final designs at engineering handoff.

Creative Analysis, with generation attached. Sort the winners, select, and create ads from the same view. Hook score sits in the table, one of the metrics participants named as missing from the platform.
Brand Vault. Logos, palette, guidelines, and naming conventions, stored once and applied to every generation.
Area selection in the editor. Mark the element, say what changes, queue it, apply when ready.
Variants with copy attached. Three concepts, each carrying its own primary text and headline.
Published. Live ads, with a link back into Creative Analysis.

Where it landed

The phase one review came back a go, and most of the recommendations were built. Some flows did not make it.

Moby Studio launched and is live today.

Precision editing became a capability beyond creative, applied to other generated outputs including dashboards and slides.

The sequencing held. Static shipped first, then video, then bulk generation. Product detail pages moved to a separate stream.

The findings gave the team building the creative automation specialist their starting material, a baseline on what the experience needs before their co-design pilot begins.

The approval gate is still open, and so is whether publish rate holds as the measure the team runs on. Getting more of a team into the platform, including the people who sign off on creative, is a live question.


Reflection

The finding that mattered most came from a recruiting decision made before any interview happened. Splitting the sample by publishing behavior rather than role meant the comparison was available in the data. Had the sample been cut by job title, the interviews would have produced a list of feature requests, and the pattern underneath, that the people pushing AI hardest were the ones shipping least, would not have surfaced.

The bulk generation recommendation is the one I think about. Saying no to the thing a team is excited about only works if you can say what would make it a yes. Bulk shipped later, once the flows around it could carry it.