Back to blog

The Visual Intelligence Gap Is Closing

Why Google's Pomelli experiment signals a shift from generating assets to understanding design systems

Sergio Giannone

Sergio Giannone

Share this post

I’ve been watching AI systems struggle with visual content for years. Ask a language model to write you an essay, and it’ll produce something coherent, occasionally even brilliant. Ask it to design a poster that actually looks professional, and you’d get something that screamed “computer-generated” from across the room.

That gap between linguistic and visual intelligence has defined AI’s practical limitations. Text generation became routine. Visual design? Not so much.

Thanks for reading Cowritten.ai! Subscribe for free to receive new posts and support my work.

Google’s new Pomelli experiment (currently only available in the US - unless you use a VPN, that is) suggests that the gap is finally closing (well, sort of - Spoiler: it doesn’t quite work as you’d expect… yet). And they’re not alone in this shift. Over the past year, we’ve seen Canva launch its own foundational design model, Figma introduce AI-powered prototyping tools, and Gamma transform presentation creation with AI agents. What’s happening across these platforms tells us something important: maintaining coherence across decisions has become the breakthrough, not just generating prettier pictures.

What agents couldn’t do

The frustrating thing about AI and design has always been consistency. A model could generate a single image that looked decent. It could write copy in a particular tone. It could even suggest colour palettes. But keeping fonts, colours, tone, and visual style aligned across a campaign? The system would collapse into contradictions.

I’ve seen teams try. They’d generate assets piecemeal, then spend hours trying to force everything into alignment. The AI would produce three versions of a social post, each in a subtly different style, requiring a designer to harmonise them manually. The promise was automation; the reality was extra work.

Pomelli’s approach feels different. The tool extracts what it calls “Business DNA” from an existing website (tone, fonts, colours, image style) and uses that profile as a constraint system for everything it generates.

Cowritten'.ai’s “Business DNA” according to Pomelli.The insight here isn’t technological wizardry. Design is a system of relationships, not a collection of assets. Most AI design tools treat each output as independent. Pomelli, it seems, treats the brand as the unit of work.

Canva has taken a similar path. The company launched its own foundational model in late October this year, trained on design elements to generate designs with editable layers and objects rather than flat images, and to understand layout hierarchy, branding rules, and visual coherence. Their model now powers all AI-driven tools across their platform and works across different formats (social media posts, presentations, whiteboards, websites).

From assets to systems

What makes this wave interesting isn’t that companies have built better image generators. They’ve built (or at least made the first steps towards building) something closer to visual reasoning systems that can hold multiple constraints in working memory and apply them consistently across outputs.

This matters more than it sounds.

Good design requires maintaining coherence whilst adapting to context. A social media post, a website banner, and an email header all need to feel like they come from the same brand, even though they serve different purposes and live in different formats.

That kind of multi-constraint problem has been beyond most AI systems. They could optimise for one thing at a time (make this image dramatic or make that copy friendly), but they couldn’t juggle competing requirements simultaneously. The models were powerful, but the orchestration was weak.

According to Canva’s head of product marketing, design hasn’t had a fit-for-purpose model yet because existing AI models generate flat, JPEG-style outputs where text, backgrounds, and other elements cannot be separated. The technical advance is that these systems are now beginning to understand structure, the relationships between elements that make design work.

Figma has approached this differently. Their new Figma Make feature uses Anthropic’s Claude 3.7 Sonnet model and can generate interactive prototypes from text prompts whilst inheriting existing design systems stored in Figma. For teams already working in Figma, this means AI that understands their established design language rather than starting from scratch.

Why small businesses matter

Google is positioning Pomelli for small and medium-sized businesses. That’s the right call. Not because large brands don’t need design systems, but because SMBs expose the real bottleneck.

Big organisations have design teams. They have brand guidelines, asset libraries, and approval processes (well, most of the time). Their problem with AI is that integration with existing workflows creates friction.

Small businesses often have none of that infrastructure. For small to medium-sized businesses, creating impactful, on-brand content can require significant investment in time, budget, and design expertise, making it a major obstacle. They need to generate professional-looking content without hiring specialists. They’re the clearest test case for whether AI design tools actually work, because there’s no safety net of human expertise to catch the failures.

If Pomelli can deliver coherent, on-brand campaigns for businesses that don’t have dedicated designers, that signals the technology has crossed a meaningful threshold. Not because anyone’s being “replaced” (spoiler alert: they aren’t!), but because design thinking becomes accessible to people who couldn’t afford it before.

Gamma has found traction in a similar space. The AI presentation tool lets users describe what they want and generates complete slide decks with a consistent visual design. For solo founders, educators, and marketers who previously struggled with presentation software, having an AI partner that maintains visual coherence across slides changes what’s possible.

The uncanny valley of visual AI

There’s still something slightly off about AI-generated design. You can usually tell. The layouts feel algorithmically safe. The colour choices lack the small imperfections that make human work distinctive. Everything is a bit too centred, a bit too balanced, a bit too obviously optimised.

I suspect that uncanny quality will persist for a while. These systems are pattern-matchers, not taste-makers. They can reproduce the statistical average of “good design,” but they struggle with the intentional asymmetries that make work memorable.

The question is whether that matters. For brand campaigns aimed at growing a small bakery’s Instagram following, statistical competence might be enough. The goal isn’t winning design awards. In their case, looking professional and consistent might suffice.

But if AI design converges on the same safe aesthetic (clean, minimal, balanced), then every brand starts looking like every other brand. The cost of accessibility might be homogeneity.

I don’t have an answer to that yet. Worth watching, though.

What this means for designers

Every time I talk about AI and creative work, someone asks whether this means designers are obsolete.

The short answer: no.

The longer answer: their role is shifting. Designers won’t spend as much time executing routine assets. They’ll spend more time defining systems like brand identities, constraint sets, and quality standards that AI systems can operationalise. It’s the same thing happening with written content.

Pomelli’s “Business DNA” concept is essentially a design system encoded for machine consumption. Someone still has to create that system in the first place. Someone still has to judge whether the outputs actually work. Someone still has to know when to break the rules.

Design is becoming more architectural and less artisanal. Less “make me a poster,” more “define the principles that make our posters work.” That’s a different skill set, and not everyone will adapt comfortably.

The experiment matters more than the product

Pomelli is explicitly labelled an experiment, which means it might not survive beyond the beta phase. Google Labs describes Pomelli as an early experiment that might take some time to get things right.

But the experiment itself carries weight. It signals what’s now technically possible. If Google can build a system that maintains brand coherence across generated assets, so can everyone else. This is a demonstration of where the capabilities are heading.

The technology remains flawed. Still biased toward safe choices. Still missing the edge that makes exceptional design exceptional.

But the threshold’s been crossed. Visual content is no longer AI’s blind spot.

If you’re exploring how AI can support your content and design workflows, reach out through cowritten.ai.

Thanks for reading Cowritten.ai! Subscribe for free to receive new posts and support my work.

Share this post

Sound like

your team?

Most teams I talk to have built something that works some of the time. Let’s find where yours stalls.

Sound like

your team?

Most teams I talk to have built something that works some of the time. Let’s find where yours stalls.