AI & Marketing Research with Dr. Eva Wolf
Not another AI news podcast. This is a research radar — a twice-weekly briefing that surfaces peer-reviewed studies on AI and marketing, tells you what the evidence actually says, and helps you decide what's worth a deeper read.
AI & Marketing Research with Dr. Eva Wolf
AI Creativity Benchmarks, Chatbot Personas & Agent Competition
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Thanks for listening to AI & Marketing Research Radar by Big Plans Media.
I’m Dr. Eva Wolf, and I help marketers, educators, consultants, and business owners turn AI marketing research into practical strategy, smarter workflows, and real business opportunities.
More episodes: https://bigplans.media/ai-marketing-research-radar/
Consulting: https://bigplans.media/ai-marketing-consulting/
Big Plans Media — Where Big Ideas Meet Smart Marketing.
You're listening to Evita, an AI-generated research briefing avatar trained on the research framework and methodology of Dr. Eva Wolfe, marketing professor, AI researcher, and founder of Big Plans Media. Every day, Evita scans emerging research in AI, marketing, consumer behavior, psychographics, and business strategy to identify the most relevant developments, opportunities, and risks worth watching. These daily radar reports are designed to help busy professionals stay informed without having to read hundreds of research papers themselves. And every Friday, join Dr. Eva Wolfe Live for her personally recorded weekly AI marketing radar roundup, where she breaks down the biggest stories, explains what actually matters, and shares practical insights and strategic implications for marketers, educators, entrepreneurs, and business leaders. Now, here's today's radar report. Here's the signal I can't ignore today. You're making decisions right now about which AI tools to use for creative work, which chatbot personality to deploy, whether to trust AI agents to run your campaigns. And you're probably making all three of those calls on vibes. Not because you're lazy, because there's been no real framework for any of it until now. Today's papers point to the same pattern. The era of just pick an AI and prompt it is ending. The teams that win are the ones who get specific. We screened 249 papers. Three cleared the full text bar and made today's radar. Quick caveat: this is a first pass research briefing, not a final academic review. All three papers today are preprints. I'll tell you what they suggest, what they don't prove, and which ones deserve a closer read. Okay, let's get into it. Paper one. Here's the business question. When your team is using AI to write copy, brainstorm campaigns, generate concepts, are you using the right model for the right job? And is there even a way to know? Because until now, there really wasn't. You had GPT-4 because everyone else was using GPT-4. You had the occasional Twitter thread about which model was quote better, that was it. Vibes. This paper changes that. BD and a large team built something called AGC Bench, the largest, most systematic benchmark of AI creativity ever run. They screened over 3,000 papers, pulled together 78 creativity benchmarks across six domains, and ran 83 different AI models through all of them. Okay, so what did they find? First big finding. AI models have a single underlying creativity score that cuts across all domains. They called it a C factor, like the G factor in human intelligence research. One number that explains 81.5% of the variation in creative performance across 83 models. So a model that's good at writing jokes is probably also good at brainstorming product ideas. Creative ability in AI is more general than most people assumed. But here's the finding that stopped me cold. Simply telling an AI to be creative, those two words boosts creative output more than turning on step-by-step reasoning mode. More than chain of thought prompting. Huh, that is not a minor UX tip. That is a free five-second upgrade to every creative workflow your team runs today. Now the catch. This is a preprint, not peer-reviewed yet. And the human comparison, where they checked how AI stacks up against actual humans, was limited to five tasks with an unspecified number of participants. So I wouldn't hang a big claim on the human versus AI horse race from this paper. Also, 67 of the 78 benchmarks are text only. If your creative work involves visual concepts, video, audio, this benchmark doesn't cover that at all. And the top human performer still beats the top AI model overall. So we're close, not there. Plain English payoff. Add be creative to your AI prompts for copy and campaign work. It's the single cheapest improvement you can make to your creative output today. Okay, here's where this becomes commercially interesting. Money move. Build a creative AI model selector, a lightweight tool, agency facing SAS, that takes a brief type and recommends the best performing model for that specific task based on AGC bench domain scores. Claude led on narrative and figurative language. GPT 5.4 led on brainstorming. Different tools for different jobs. That's a billable layer most agencies aren't offering yet. Action step. Today run the same creative brief through your AI tool twice. Once with your normal prompt, once with be creative added. Send both outputs to a colleague without labeling which is which. See if they can tell the difference and which one they prefer. Evidence check. Preprint, no peer review, and the human comparison data is thin. Treat the Be Creative prompting finding as a strong hypothesis to test, not a confirmed law. Radar verdict, test this week. The prompting insight is low effort and immediately testable. The model selection work is the piece that compounds over time. This next one looks like a design theory paper. Stay with me, the business implication is more practical than the title suggests. Paper two. Here's the business question. If your brand has an AI chatbot, customer service, sales, onboarding, is it costing you trust by using the same personality for every single interaction? Think about it. Your bot is probably set to helpful and enthusiastic, and it stays that way whether someone is celebrating a new purchase or filing a complaint for the third time about the same order. That's a problem. Rahman and Desai synthesized existing empirical work on AI chatbot personas and proposed a framework they're calling the Fluid Personality Framework. The core idea. There are two dials you should be adjusting. One is which role the bot is playing: coach, tutor, or tool. Two is how intensely it expresses its personality, how warm, enthusiastic, or energetic it sounds. Right now, most brands are leaving both dials frozen. Here's what the underlying research they cite actually shows. A medium level of personality expression, not too flat, not over the top, consistently scores better on trust, likability, and enjoyment than either extreme. Think of a person who's warm but not gushing. That middle ground works best. And making a bot sound very human, like a doctor or a financial advisor, can backfire. Users form unrealistic expectations. In one health app study, a doctor persona was preferred. In finance, it made no difference at all. The lesson: trust doesn't require a human persona. It requires the right persona for the context. Now, I want to be really clear about this. This paper does not run a new experiment. No user study, no data collected. It's a conceptual framework built on prior research. That prior research is real, but this paper itself proves nothing new. And that actually bothers me. The framework is genuinely useful, but without a validation study, you're adopting a design philosophy, not acting on confirmed findings. Know the difference. Plain English payoff. If your chatbot sounds the same whether a customer is excited or furious, that's not a consistent brand voice. That is a trust problem waiting to become a churn problem. Here's the business hiding inside the research. Money move. Offer a chatbot persona audit service. Go into a client's existing AI assistant. Identify the places where the personality is mismatched to the interaction context, a bubbly bot handling complaint escalations, for example, and redesign those system prompts using the medium intensity principle and context switching logic from this framework. That's a concrete, repeatable deliverable, not use AI better. A specific audit with a specific output. Action step. Pull up your current chatbot system prompt. Read it out loud. Ask yourself, does this personality make sense if the customer is frustrated? If the answer is no, write a second version for high friction moments. Test both in your next sprint. Evidence check. Pre-print position paper. No new data. The medium personality is best finding comes from prior studies, and at least one of those studies found no effect in a health context. The rule is not universal. Test your specific use case. Radar verdicts. The framework is a useful design scaffold, but it needs your own validation before you bet a chatbot rebuild on it. This is the paper I almost skipped, and I'm genuinely glad I didn't. Because if AI agents start handling your media buying, your content production, or your campaign management, what happens in this paper is coming to your industry fast. Paper 3. Here's the business question. If you deploy AI agents to run marketing tasks, ad buying, content creation, campaign management, what actually determines whether those agents succeed or blow your budget? This paper doesn't come from a marketing journal. It comes from AI research. But I'm telling you, if you're building or evaluating any AI agent stack, this is the paper you need to understand. Chu, Zhang, and Vandershar built a simulated gig economy. They called it AI work, where LLM agents compete for posted jobs over multiple rounds. They bid on tasks, built reputations, invested in capabilities, and adapted strategies. Hmm. Think upwork, but every participant is an AI. So what did they find? Three capabilities separated the winning agents from the losing ones, knowing their own strengths accurately, watching what competitors were doing, planning several steps ahead. That's metacognition, competitive awareness, and long horizon planning. And here's the finding that stopped me cold. Adding a simple prompting scaffold, one that nudged agents to think about those three things, made them capture one and a half times more market share. Same underlying model, just better prompting architecture. Not a rebuild, a layer. And the platform design findings are equally important. When the platform made everyone's prices public, agents raced to undercut each other and quality dropped. When prices were hidden, agents invested in actual capability improvement. Sound familiar? That's every ad tech race to the bottom you've ever watched happen in real time. The catch? This is a simulation, not a real market. The researchers are explicit. They built it to illustrate mechanisms, not predict real-world outcomes. The 1.5x market share finding is directional. It's a hypothesis to test, not a number to put in a deck. And the number of agents, the specific models used, the full experimental parameters, not fully detailed in the available text. That's a methodology gap I'd want closed before citing this confidently. That's the piece I keep coming back to. The insight is genuinely important. The evidence is thin enough that you shouldn't act on the number, but you absolutely should act on the principle. Plain English payoff. An AI agent that doesn't know what it's good at will waste your money and damage your results. And a simple prompting change can fix that. Here's the monetizable angle. Money move. Build a prompt scaffolding product for AI agent deployments. A layer that automatically injects metacognition, competitive awareness, and planning instructions into any LLM agent stack. Sell it as a performance layer, not a new model, not a new platform. A prompting architecture that makes the agents you already have dramatically more effective. That's a real product gap right now. Action step. If you're running any AI agents today, even simple ones handling content or outreach, add three explicit instructions to their system prompts. One, assess whether this task is within your capabilities before starting. Two, consider whether a different approach would produce a better outcome. Three, think about how this task fits into the longer-term goal. Run it for two weeks. Measure output quality. Evidence check. Preprint Simulation Study. No real world agents, no real markets, no human participants. The findings are directional hypotheses. Validate in your own deployment before drawing hard conclusions. Radar verdict, watch list. The mechanisms are important, and the prompting insight is testable today. But the simulation basis and thin methodology details mean you're building on a hypothesis, not a finding. At first glance, these three papers look completely separate. A creativity benchmark, a chatbot design framework, an AI agent economics simulation. What does any of that have to do with the other? Together, they're pointing at the same ship. Genere over. The teams that win are the ones who get specific. Specific about which model they use for which creative tasks. Specific about which personality mode they deploy in which customer moment. Specific about how they instruct their agents to think. Not more AI, better insertion points. Not louder chatbots, more calibrated ones. Not smarter agents, better prompted ones. Here's the tension I keep sitting with. All three of these papers are preprints. The creativity benchmark is the most methodologically solid of the three, but even that has a limited human comparison. The chatbot framework has no new data at all. The agent simulation is explicitly a stylized illustration. So the pattern is real. The signal is real. But we're still in the hypothesis phase, not the proven playbook phase. And that's actually the most important thing I can tell you. Because I've watched teams bet big on early stage AI research and get burned. The right move right now is to test these ideas cheaply, not deploy them wholesale. The bigger theme is this specificity is the new competitive moat. Not which AI you use, how precisely you've configured it for the exact job you're asking it to do. That's the piece most teams are still missing. Here's the playbook from today. One, add be creative to your AI prompts for copy and campaign work. Run a blind test today. It costs nothing and the evidence is strong enough to warrant it. 2. Review your chatbot system prompt for personality mismatches. Write a second version for high friction customer moments. Test both in your next sprint. 3. If you're running AI agents on any marketing task, add three self-awareness instructions to their system prompts. Capability check, approach consideration, long-term fit. Run it for two weeks and measure. Evidence check on all of that. All three papers today are preprints with no peer review. The creativity benchmark is the most methodologically solid. The chatbot framework has zero new data. The agent simulation is explicitly a stylized model. Use them to decide what to test, not what to blindly believe. Links to all three papers are in the show notes. Read the originals before making major decisions. Want the human expert take? Join Dr. Eva Wolf every Friday for the AI Marketing Radar Roundup, where she extracts no nonsense, money-making tips, practical strategy, and real business opportunities from the week's research. Subscribe on Apple Podcasts, Spotify, YouTube, and wherever you listen to podcasts. This is Evita for Big Plans Media, and I'll be back in the next radar brief.