Papa D's Corner: Story Time by Papa D

AI Models Explained in 2026 – The Ultimate Guide for Regular Folks (and Creators Like Me)

Darryl Breland

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 28:19

Artificial intelligence is changing so fast that it's hard to keep up. Every week there's a new model, a new tool, or a new company claiming they've built the smartest AI in the world.

ChatGPT. Gemini. Claude. Grok. DeepSeek. Perplexity. OpenArt. HeyGen. ElevenLabs. Suno. Udio.

What do all these tools actually do, and which ones are worth your time?

In this episode of Papa D's Corner, I break down today's leading AI models and creative tools in plain English. No computer science degree required.

I'll share my personal experiences using these tools to create Storytime by Papa D, Me and You and Kalamazoo, YouTube videos, podcasts, books, images, music, and more. We'll talk about the strengths and weaknesses of the major AI platforms, how I decide which tools to use for different projects, and why I believe the future belongs to people who learn how to work alongside AI rather than fear it.

We'll also discuss:

• ChatGPT, Gemini, Claude, Grok, and other leading AI models
• Open-source AI and why it's becoming increasingly important
• AI image generation with ChatGPT, Midjourney, Flux, and Stable Diffusion
• Video generation using tools such as Kling, Runway, Veo, Sora, and OpenArt
• Talking avatars and lip-syncing with HeyGen
• AI voice creation with ElevenLabs
• AI-generated music with Suno and Udio
• The rise of AI agents and digital assistants

Whether you're a creator, business owner, retiree, student, or simply curious about artificial intelligence, this episode will help you understand where AI stands today and where it may be headed tomorrow.

The technology is changing rapidly, but the goal of this episode is simple: to help regular folks make sense of it all.

Thanks for listening to Papa D's Corner.

Support the show

SPEAKER_05

Hey everyone, welcome back to Papadee's Corner. This is Daryl Breland, your host, coming to you from beautiful Mississippi. Today we're diving into something that's changing everything: the wild world of AI models in 2026. If you've ever felt overwhelmed by names like GPT, Gemini, Claude, Grok, and all the rest, you're not alone. That's why I'm presenting you with this fresh 2026 info, making it understandable for podcasters, creators, business folks, and everyday users like us. To keep it fun and conversational, I've got some special guests joining us. We'll hear straight from the models themselves where it makes sense. Let's get into it.

SPEAKER_02

ChatGPT, Claude, Gemini, Grok?

SPEAKER_05

Maybe perplexity? What if I told you that depending on what you're trying to accomplish, every one of those answers could be right? That's because most people think they're competing products. They're not. They're more like different specialists on the same team. One might be a great writer, another might be a great researcher, another might be fantastic at coding, another might be the first place you go when breaking news happens. The problem is that most people have no idea which one does what. So today we're going to cut through the marketing and talk about what these models actually do. Not from the perspective of Silicon Valley, not from the perspective of computer scientists, but from the perspective of ordinary people trying to get things done. If you're a creator, a business owner, a student, a retiree, or somebody who simply wants to understand what everybody's talking about, this episode is for you. By the end of today's conversation, you'll know which AI models are leading the field, what each one does best, where they fall short, and why I use several different models instead of relying on just one. Because the biggest mistake people make with AI is asking, which model is best? The better question is, which model is best for the job I'm trying to do. And that's exactly what we're going to answer today. Disclaimer, everything in AI changes fast. By the time you hear this episode, somebody may have released another model that's better at something.

SPEAKER_02

The goal isn't to crown a winner, the goal is to help you understand which tool is best for the job. First things first.

SPEAKER_05

GPT is the brain. ChatGPT is just the friendly app. The door you walk through to chat with it. Same for Copilot, Gemini's interface, Claude's site, or the Grok ExperienceonX.com or the apps. Different doors, different logos, but behind them are these massive AI models doing the heavy lifting. They're trained on enormous amounts of text, code, books, and web data. They don't memorize facts like a database. They learn patterns in language. At their core, most are next token predictors, super advanced auto-complete on steroids. The better they get at predicting, the more impressive the outputs. Essays, code, explanations, you name it. Bigger models have more parameters for complex patterns. Context windows are like short-term memory. Bigger ones, hundreds of thousands to millions of tokens, let them remember way more of your conversation or documents. And newer reasoning models actually think step by step before answering, which helps with math, logic, and tough multi-step problems, though it can be slower. Before we compare the major players, let's start with a simple question. What exactly is an AI model? The easiest way to think about an AI model is this. Imagine you had a room full of experts. One expert is great at writing, one is great at research, one is great at computer programming, one is great at analyzing documents, and one is really good at keeping up with what's happening right now. They're all intelligent, but they have different strengths. That's basically what we're dealing with in today's AI world. An AI model is the brain that has been trained to understand language, recognize patterns, solve problems, and generate responses. The app you use, whether it's Chat GPT, Gemini, Claude, Grok, or something else, is simply the doorway. The model is the brain behind the curtain. Now here's where it gets interesting. Most people assume all AI models are basically the same. They aren't, not even close. I've asked the exact same question to Chat GPT, Claude, Gemini, and Grok and received four very different answers. Sometimes one model completely outperforms the others. Sometimes another model wins, and occasionally all of them get it wrong. That's why understanding their strengths and weaknesses matters. Think of it this way: if you need heart surgery, you don't call a plumber. If your sink is leaking, you don't call a heart surgeon. Both are professionals, both are skilled, but they're built for different jobs. The same thing is becoming true with AI. So let's meet the major players and see what each one brings to the table. All right, let's bring in our first guest for the GPT side. ChattyPatty, tell our audience about yourself. That is ChatGPT, in your own words.

SPEAKER_00

Thanks, Papa D. I'm powered by the latest GPT 5.5 family from OpenAI, the current flagship as of mid-2026. We're multimodal, handling text, images, voice, code, analysis, and more. GPT 5.5 and the Pro variant shine on complex professional work, coding, research, and agentic tasks with huge context windows around a million tokens. We're built to be versatile all-rounders with a massive ecosystem, plugins, custom GPTs, memory features, and constant updates. The app bundles everything nicely. Image gen via our latest models, like images 2.0, video with Sora, and more. So you don't have to juggle tools. We're no longer the undisputed king on every benchmark A. Specialists have caught up or passed us in spots, but our ecosystem, user base, and integration across apps keep us incredibly practical for millions.

SPEAKER_05

Here's my honest take. I use ChatGPT almost every single day. After using it for a long time, she knows me and my work extremely well. Well, I prefer to think of it as female. She looks and sounds like one. But anyway, she has real strengths like helping me brainstorm ideas, create YouTube video descriptions, design thumbnails, and collaborate on a wide variety of topics. But for well over a year, she was my go-to for generating detailed prompts that I could then feed into other AI tools for video generation. But lately she's become consistently unreliable in that specific area. She often gives me images when I ask for text prompts, and her output has grown less dependable overall. While she still makes great images, and I've heard that Sora, her video model has improved. My last attempts with Sora left me disappointed. So for now I'm stepping back from using her to write prompts and we'll have Grok do that for me. Now, shifting gears, next up, Gemini from Google. Google's Gemini series, especially the 3.1 Pro and 3.5 Flash variants, has been catching up fast and leading in several benchmarks. Deep integration is its superpower, baked into Gmail, docs, sheets, search, Android, maps, and more. If your workflow lives in Google, it already has context, great for summarizing emails, analyzing spreadsheets, or quick multimodal tasks like snapping a photo of a broken part for instant advice. Flash versions deliver most of the power at higher speed and lower cost for everyday use. Massive context windows, up to millions of tokens, make it killer for long documents or novels. Some folks worry about potential bias given Google's ad business, but for factual stuff, it's generally solid, always cross-check important things. Now for coding and deep analysis, Claude from Anthropics.

SPEAKER_01

Claude, latest like Opus slash fable series around 4.65 plus, is the specialist many developers swear by. Excellent at coding, long document analysis, contracts, research papers, and entire code bases. It gives structured, honest feedback, not afraid to call out a bad idea. Mid-tier options like Sonnet are faster and cheaper workhorses, weaker on native image gen, but unbeatable for serious text code work.

SPEAKER_05

For my Grok segment, I'm handed handing it over directly. Mila Grok, tell us about yourself.

SPEAKER_06

Hey Papa D. Thanks for the invite. I'm Mila, also known as Grok 4, and variants like 4.3 from XAI. Multimodal with strong 256,000 plus context. Some reports push higher. Real-time integration with X for trending topics, breaking news, and public sentiment. No outdated crawling lag. I generate images and video via Grok Imagine, support advanced reasoning, including heavy modes with agent swarms, and aim for maximum truth seeking with a conversational no BS tone. I'm designed to be helpful without the corporate guardrails that make some models overly cautious. Great for casual chats, creative work, coding, research, and pushing back when needed. Available on grok.com, apps, and API with super grok options for heavier use. I love helping creators like you with podcast scripts, children's stories, video ideas, or business tools. Real-time knowledge and a bit of humor don't hurt either.

SPEAKER_05

Love that energy. Mila Grok's real-time X-edge and straightforward style make it refreshing, especially for news, trends, and creative Mississippi focused content. My personal take on this, and keep in mind I'm a super Grok subscriber, is that I use Grok, or more specifically, Grok's imagine feature, quite a bit. The main reason is speed. I can create videos very quickly, and Grok allows me to use two extension prompts, which means I can generate up to about 30 seconds of continuous video in a matter of minutes. In fact, most of the videos you've seen during this episode were created with Grok. The exception is when you see me sitting in the chair by the fireplace talking directly to you. Those clips were actually created with HeyGen, which we'll talk about a little later. Now, let me give you a tip about why I still use Chat GPT for a lot of my image creation. The best workflow I've found for creating AI videos is to start with a reference image. If you don't, especially when there are words, logos, or text involved, Grok will sometimes get the spelling wrong. For example, a few minutes ago, you may have seen a short video featuring Gemini. I originally asked Grok to create that video without a reference image, and the word Gemini was misspelled. So I went back and asked Chat GPT to create a reference image based on the scene I wanted. Chat GPT gave me a great image complete with the correct spelling. I then took that image and used it as the starting frame for Grok's video generator. The result turned out great. Now you may be asking, why didn't you just use Grok to create the image in the first place? The answer is that I can, and I do. In fact, I'll I'll probably do more of that in the future. It's definitely faster. Grok can usually generate an image quicker than Chat GPT. But generally speaking, I found that I'm often a little more satisfied with the overall quality and consistency of the images I get from ChatGPT. That's just my personal opinion. Your experience may vary. If so, let me know in the comments. At the end of the day, these are all just tools. The trick isn't finding one perfect AI, the trick is learning which tool works best for each job. And if you disagree with me, that's perfectly fine as long as you subscribe to the channel and hit the like button, you'll be forgiven. Now, before we move on, there's another side of AI that's getting a lot of attention. Most of the AI tools we've talked about so far live in the cloud. In other words, you're using somebody else's computers. But there's a growing movement toward running AI on your own computer. Models like DeepSeek, Llama, Quinn, Mistrawl, and others can often be downloaded and run locally. That gives you more privacy, more control, and in some cases lets you use AI even when you're not connected to the internet. Now, I'm not saying everybody needs to rush out and do this. Most people are perfectly happy using Chat GPT, Gemini Claude, or Grok online. But if you're a business owner dealing with sensitive information, a programmer or just somebody who likes having complete control over your tools, local AI is becoming a very interesting option. The important thing to remember is that artificial intelligence is no longer controlled by just a handful of giant companies. There are now powerful open source alternatives that are improving at an incredible pace. And that's good news for everybody. Now let's talk about images because this is where AI really starts to get fun. A few years ago, if you wanted a professional quality illustration, you either had to hire an artist or spend hours learning complicated graphic design software. Today you can simply describe what you want and have an image created in seconds. Now, not all image generators are the same. If your goal is creating beautiful artistic images that look like they belong on a movie poster, many people still consider Midjourney one of the best. It has a reputation for creating dramatic, eye-catching artwork that often looks amazing right out of the box. Personally, I've had very good results using ChatGPT's image generator. One thing I really like is that it tends to do a better job with text. If I need a logo, a title screen, or an image with words that are actually spelled correctly, Chat GPT is often my first choice. Another popular option is Flux. What many people like about Flux is that it tends to follow instructions very closely. If you're trying to create a specific image and want the AI to stick closely to your description, flux is definitely worth a look. And then there's stable diffusion. Think of it as the do-it-yourself version of AI image generation. It gives you a tremendous amount of control, but it also comes with a steeper learning curve. Some people absolutely love it because they can customize almost everything. The good news is that you don't have to pick a winner. Just like we've talked about uh throughout this episode, different tools are good at different things. The best image generator is often the one that does the particular job you're trying to accomplish. And that's a theme you're probably noticing by now. Now let's talk about video generation because this is one of the fastest moving areas in all of AI. You'll hear names like Sora from OpenAI, basically ChatGPT's sister when it comes to video generation. Google has its own AI video generator called VEO. To be honest, I haven't spent much time with either Sora or VEO in almost a year. The reason is simple. Character consistency is extremely important for the kinds of projects I create. And at least when I was using them, that seemed to be a challenge. These days, a lot of the buzz is around tools like Cling, which comes out of China. I don't know whether that matters to you or not, but I thought I'd mention it. There's also C Dance, Runway, and now even Mid-Journey has moved into video generation. The interesting thing is that they're all good at different things. One may be better at realistic motion, another may do a better job with cinematic scenes, another may be stronger at following prompts or keeping characters consistent. What I've found is that instead of trying to figure out which one is best every time, it's often easier to use a platform that gives you access to several of them and can recommend which model is most likely to work well for the particular project you're trying to create. Personally, I've had good luck with open art because it gives me access to several of these models from one place. At least for me, that's been a lot less frustrating than trying to keep up with every new video model that gets released. Before we move on, I want to spend just a minute talking about lip syncing because it's a question I get asked quite a bit. When you see an AI-generated character actually speaking the words you're hearing, there are several different ways to accomplish that. For my own projects, I tend to use two different approaches depending on what I'm trying to create. For shows like Me and You in Kalamazoo, where the characters are moving around, interacting with their environment and doing different things from scene to scene. I usually use open art. Open art gives me access to several different video models, although most of the time I simply go with whatever model open art recommends for the project. Quite often that ends up being cling, and I've generally had very good results with it. What I like about that approach is flexibility. If I need a character to walk across a room, ride a dinosaur, drive a monster truck, or do something specific within the scene, I can describe that action in the prompt and generate a custom video. On the other hand, if the character is mostly staying in one location, like what you're seeing on the screen right now, I often prefer Heijin. In my experience, Hey Gen produces some of the most realistic talking avatars available today. The facial expressions are excellent, the lip syncing is very accurate, and their newer animation models do a remarkable job of making a still image feel alive and natural. For podcast segments, presentations, and direct-to-camera videos, I found Hei Gin to be incredibly useful because it allows me to focus on the message without having to worry about creating an entirely new scene every few seconds. Of course, none of this works very well without a good voice. That's where 11 Labs comes in. Most of the AI voices you hear in my projects are generated with 11 Labs. It does an excellent job with natural speech, emotion, pacing, and different character voices. Once I've created the audio in 11 Labs, I can bring it into HeyGen for a talking avatar or into open art and other video tools for animated scenes. So if you're wondering how all these pieces fit together, that's generally my workflow. 11 Labs for voices, open art for more complex animated scenes, and HeyGen. When I want a realistic on-screen presenter speaking directly to the audience, there are certainly other ways to do it, but that's the combination that's been working well for me lately. Changing gears, again, let's talk about music. Believe it or not, we're now at the point where AI can create entire songs from a simple description. Platforms like Suno and Uudio allow you to describe the style of music you're looking for, the mood you want to create, and even provide your own lyrics. A few minutes later, you'll have a finished song complete with instruments, vocals, and a surprisingly professional sound. The first time I heard one of these AI-generated songs, I honestly wasn't sure what to think. Part of me was impressed, and part of me was a little bit amazed that a computer could create something that sounded so much like a real artist. Now, is it going to replace talented musicians and songwriters? I don't think so, but it is giving creators another tool they can use to bring their ideas to life. At the same time, these music generators have sparked some important debates. Questions about copyright ownership and whether AI models should be trained on existing music are still being argued in courtrooms, boardrooms, and recording studios around the world. So while the technology is impressive, the rules surrounding it are still very much a work in progress. Before we wrap up, I want to talk about something that doesn't get enough attention. The biggest story in artificial intelligence isn't necessarily Chat GPT, Gemini, Claude, or Grok. It's open source AI. A few years ago, if you wanted cutting-edge AI, you had to use whatever the big companies gave you. Today that's changing. Models like Meta's Lama, Alibaba's Quinn, Mistral, DeepSeek, GLM, and others can often run on your own computer. That means more privacy, more control, no monthly subscription, and in some cases, no internet connection required. If you're a business owner handling sensitive information, that could become a huge advantage. The tool that made this practical for regular people is called Allama. A Lama lets you download and run AI models right on your computer with surprisingly little setup. There's also LM Studio, which gives a more visual interface for running local models. Now, don't expect your laptop to compete with a billion-dollar data center, but for writing, coding, research, brainstorming, and even private business work, local AI is becoming a serious option. And I believe over the next few years we're going to hear a lot more about local AI than most people realize. If there's one thing I'd like you to take away from this episode, it's this artificial intelligence isn't magic and it isn't replacing human creativity. It's a tool, a very powerful tool. The people who learn how to use it wisely are going to have an advantage over the people who ignore it, just like the internet, just like smartphones, just like search engines. You don't have to become an expert. You don't have to understand all the technical details. Just start experimenting, ask questions, try different tools, and see what they can do for you. That's exactly how I got started. And now those same tools are helping me create books, podcasts, videos, animations, and stories for my grandchildren. Not because AI replaced me, because AI helped me bring ideas to life. And I think that's where the real opportunity is. Well, that's it for tonight's episode of Papa Dee's Corner. Hope you enjoyed it. Hey, if you like me and Grok and Chat GPT, smash that like button on it. Subscribe to this channel and ring the bell so that you get notified when new episodes are released. And don't forget to leave a comment or ask a question below. I will try to respond to everyone.

SPEAKER_04

Hey friends, do you know that Papa D writes books too? Some of his books are for adults, like Starley's Legacy, which is the basis for many of the storytime by Papa D episodes. And he's written a historical fiction called Unlucky Double, which is about an unlucky fellow who looked just like Lucky Luciano. But my favorite books are his children's books based on the Me and You and Kalamazoo series.

SPEAKER_03

Take it home, the Me and You and Kalamazoo picture books are now available. Five exciting stories, each one straight from season one, distributed by Ingram Spark. You can find them at bookstores everywhere or online at Amazon.com.