Back to blog

How I cope with a new AI model every week

Reflections on speed, systems, and selective attention

Sergio Giannone

Sergio Giannone

Share this post

In the last few weeks alone, multiple major AI model announcements landed in my inbox. Claude had a new version. Google released Gemini 3, quickly climbing to the top of performance benchmarks. Someone on LinkedIn was breathlessly explaining why that Chinese model would change everything.

More recently, GPT 5.2 claimed back the number one spot. My work chat lit up with the inevitable question: “Should we switch?”

Thanks for reading Cowritten.ai! Subscribe for free to receive new posts and support my work.

This scene repeats itself weekly now. Sometimes daily. And if you’re a senior marketer, you know exactly how exhausting this feels. The promise was that AI would make our work easier. Instead, we’re drowning in decisions about which models to use, when to migrate, and whether we’re falling behind by not adopting the latest release.

Here’s what I’ve learned after a few years of navigating this chaos: you don’t need to keep up with every model. You need a system that turns weekly releases from sources of anxiety into manageable signals for improvement.

Why weekly releases feel different

Acceleration and the expectation treadmill

The pace of AI model releases has compressed what used to be yearly technology cycles into weekly sprints. OpenAI and other leading labs have been shipping significant updates and new models at a pace that feels closer to weeks than years. Announcements arrive with such regularity that Fortune describes it as creating “announcement fatigue” amongst technology leaders. This rate of change is so rapid that traditional evaluation processes are obsolete.

Each release arrives wrapped in language suggesting immediate competitive advantage for early adopters and obsolescence for those who wait. Marketing materials promise transformative capabilities. Social media amplifies the urgency. Your competitors announce they’re using the latest model. The pressure compounds: are you being prudent by waiting, or are you already falling behind?

The real cost: decision fatigue and information toxicity

Survey data using Quantum Workplace’s burnout metrics found that frequent AI users report burnout levels around 45% higher than non‑users, although this is correlational and doesn’t prove AI is the sole cause. This isn’t because the technology is difficult to use, but because the continuous stream of decisions about whether to adopt, migrate, or wait creates what researchers call “decision fatigue”, the degradation of decision quality that occurs when we face too many choices in compressed timeframes.

The psychological burden extends beyond decisions. We’re experiencing what occupational health researchers term “information toxicity”, meaning the erosion of wellbeing from continuous exposure to information suggesting our skills might be obsolete, our choices might be wrong, and our current tools might already be outdated. LinkedIn posts announce which jobs AI will replace. News coverage breathlessly reports each capability advance. The ambient anxiety exhausts us before we even begin our actual work.

The mistake most teams make

Tool-first thinking versus problem-first thinking

When a new model drops, the natural instinct is to ask: “What can this do?” We start with the tool’s capabilities and work backwards to find applications. This approach feels logical but creates a fundamental misalignment. You end up with solutions looking for problems, features nobody requested, and workflows designed around what’s technically possible rather than what’s actually needed.

The teams that successfully navigate model proliferation flip this equation. They start with clear problems: “Our product descriptions lack consistency” or “Campaign ideation takes too long.” Only then do they ask whether new capabilities address these specific challenges. McKinsey’s work on digital and AI transformation consistently finds that organisations realise the most value when they redesign workflows around specific business problems rather than leading with tools.

How tool sprawl erodes brand and capacity

Without clear governance, teams naturally experiment with whatever models seem promising. Marketing tries one tool for social media copy. Product uses another for documentation. Customer service adopts something different for response drafts. Consultancies such as Deloitte note that fragmented digital environments and overlapping tools can increase stress and make work feel more chaotic, especially when governance is weak.

The damage extends beyond stress. When every team uses different models with different prompting styles, brand voice fragments. Output quality varies wildly. Knowledge gets siloed in model-specific formats. The supposed efficiency gains evaporate as teams spend more time coordinating across incompatible systems than they save through automation.

A simple mental model for staying sane

Signal versus noise (three criteria to decide attention)

Not every model release deserves your attention. I use three criteria to filter the signal from the noise. First, capability delta: does this model offer a meaningful improvement over what we’re using? Not marginally better benchmarks, but capabilities that would change how we work. Second, alignment with current use cases: does it address an active problem we’re trying to solve? Third, migration cost: can we switch without redesigning our entire workflow?

If a release doesn’t meet all three criteria, it goes into a “watch” category. I note it exists, but don’t investigate further. This simple filter eliminates roughly 90 per cent of announcements from immediate consideration. The mental relief is immediate — you’re no longer obligated to evaluate everything, just the subset that could genuinely matter for your specific context.

A practical playbook

Choose an opinionated stack (two to three models) and why that matters

Depth beats breadth when it comes to AI tools. Master two to three models consistently rather than maintaining shallow familiarity with many options. Pick a primary model for content generation, perhaps another for specialised tasks like code or analysis, and standardise. Trust me, this is liberating, not limiting. Your team builds genuine expertise. Prompts get refined and reused. Output becomes predictable.

The “opinionated” part matters. Don’t just pick popular models. Choose based on your specific needs. If brand voice consistency is paramount, you might prioritise models with better instruction-following. If you handle sensitive data, privacy policies might trump raw capability. Document why you chose what you chose. This makes future migration decisions clearer — you’re not comparing features in abstract but asking whether new options better serve your documented priorities.

Prompt and context engineering rules for consistent brand output

Models change, but good prompt engineering principles remain stable. Build a prompt library that encodes your brand voice, content standards, and common patterns. I’m not talking about templates, but rather about operational documentation of how your brand speaks. When you do migrate models, these prompts become your compatibility test suite.

Context engineering prevents drift. Don’t just tell the model to “write professionally.” Provide examples of your best content. Include your brand guidelines. Specify what you don’t want as clearly as what you do. Empirical studies in academic and industry benchmarks show that better instructions and richer context can significantly improve LLM output quality, sometimes by dozens of percentage points on specific tasks. That’s the difference between usable first drafts and content that requires complete rewriting.

Simple evals: datasets, assertions and LLM-as-judge

Evaluation doesn’t require sophisticated infrastructure. Start simple. Collect 20-30 examples of good output — product descriptions or blog posts you’re proud of, email campaigns that performed well, support responses customers loved. This becomes your golden dataset. When evaluating new models or prompts, run them against this set. How often does the output match your quality bar?

Code assertions catch structural problems. Every product description needs a clear benefit statement? Write a check for it. Email subjects should be under 50 characters? Automate that validation. These are simple if-then statements that catch obvious failures before humans review.

LLM-as-judge provides scalable quality assessment. Use your current model to evaluate outputs from candidates: “Does this product description match our brand voice guide? Score 1-5 with explanation.” It’s not perfect, but it’s consistent and fast. You can evaluate hundreds of outputs in minutes rather than hours.

Psychological safety, consent and trace handling

Teams need to feel safe admitting when they don’t understand new tools or expressing scepticism about promised benefits. Create space for questions without judgment. Acknowledge that even experts feel overwhelmed by the pace of change. When leaders model uncertainty (”I’m not sure if this new model is worth switching to, let’s test it”) teams feel permission to be human rather than perpetually confident.

Privacy isn’t optional. If you’re reviewing AI outputs to improve prompts, you need explicit consent for trace inspection. Be clear about what you’re reviewing and why. Default to synthetic test data rather than real customer information. Yes, this makes debugging harder. It also maintains trust and likely keeps you compliant with emerging regulations. Build privacy into your architecture from day one rather than retrofitting it after problems emerge.

Common counterarguments and how to answer them

“But what if the next model is transformative?”

Truly transformative advances are rare and obvious. GPT-3 to GPT-4 was transformative. Most week-to-week releases are incremental. If something genuinely revolutionary emerges, you’ll know — not from breathless LinkedIn posts but from concrete examples of previously impossible tasks becoming possible. Until then, a quarterly review window is sufficient.

“Won’t this slow us down?”

Teams that chase every release move fast in circles. Teams with stable, well-tuned systems move steadily forward. McKinsey’s research shows that organisations with disciplined AI adoption achieve 2.5x better outcomes than those attempting to use every available tool. Sustained delivery beats sporadic sprints.

What winning looks like when you stop chasing every drop

Success with AI doesn’t mean using the newest model or the most models. It means building systems that absorb useful improvements whilst maintaining operational stability. It means your team spends time solving customer problems rather than evaluating tools. It means your brand voice stays consistent whether you’re using this week’s model or last quarter’s.

The teams thriving in this environment share common traits. They’ve replaced anxiety with process. They’ve chosen depth over breadth. They’ve built systems that treat model releases as signals for periodic evaluation rather than commands for immediate action. Most importantly, they’ve remembered that AI is meant to augment human capability, not exhaust it.

The next time three model announcements land before your morning coffee, you won’t need to panic. You’ll note them in your watch list, continue with your planned work, and evaluate them properly when your next research window opens.

That’s building something that lasts, not falling behind.

Thanks for reading Cowritten.ai! Subscribe for free to receive new posts and support my work.

Share this post

Sound like

your team?

Most teams I talk to have built something that works some of the time. Let’s find where yours stalls.

Sound like

your team?

Most teams I talk to have built something that works some of the time. Let’s find where yours stalls.