Macpreneur: Helping Solopreneurs Streamline Their Businesses on a Mac
The show for solopreneurs who can't imagine running their businesses on anything else than a Mac.
Discover jargon-free tips, tools and strategies to streamline your business, save time ⏱ and money 💸 while enjoying your solopreneur lifestyle without tech-related stress.
Mix of solo shows and interviews with fellow Macpreneurs who share their own tips, tools and strategies allowing them to be more efficient and productive running their businesses on their Mac.
Macpreneur: Helping Solopreneurs Streamline Their Businesses on a Mac
Mac Solopreneurs, Finally Understand What AI Agents Can and Shouldn't Do
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
🆓 Quiz: Discover Your AI Stage (https://macpreneur.com/aistage)
Most people think AI agents are just "smarter automations." They're not. And that distinction could save you thousands in wasted tokens and broken workflows.
In this week's episode, I break down exactly what separates a true AI agent from an AI-powered automation and walk through how I evolved my Virtual Assistant Liaison (VAL) from a simple Claude project into a semi-autonomous system.
Show notes and video version: https://macpreneur.com/episode175
Links to past episodes:
- MP172: Why Mac Solopreneurs Should NOT Start With AI Agents (And What to Do Instead)
- MP173: Mac Solopreneurs Get Better AI Output Using the PTCF Framework
- MP174: How Mac Solopreneurs Build AI Assistants That Already Know Their Business
Highlights
- 00:00 Teaser
- 00:37 Welcome
- 01:51 VAL Evolution Phases
- 03:53 Memory Cascade System
- 06:27 Morning Brief Workflow
- 08:12 What Makes an Agent
- 11:05 Automation vs Agent
- 12:36 News Brief Case Study
- 17:02 Deep Research Starter Agent
- 19:11 Lethal Trifecta Safety
- 23:25 Recap
- 26:42 Outro
Text me topics for future episodes
🆓 7-day FREE trial of Claude Cowork and other Claude Pro features
Here's the link: https://macpreneur.com/claude2606
⚠️ Limited to 3 guest passes and only for people who never had a paid Claude subscription before.
🎤 Want to be a guest on the show? Fill the application form available at https://macpreneur.com/apply or visit the show profile on Podmatch.
👥 Join the Macpreneur Community!
Simplify your digital life and streamline your business with fellow solopreneurs.
✨ Join the Waiting List
âś… Want to be more efficient on your Mac?
Answer a few questions about how you're currently dealing with unnecessary clicks, repetitive typing and file clutter. It's FREE and takes less than 2 minutes!
🆓 Get personalized time-saving tips today!
Follow me:
Introducing VAL: From Assistant to Agent (00:00)
Damien Schreurs In the last episode, I introduced you to VAL, my Virtual Assistant Liaison, an assistant that graduated into an agent. VAL now prepares my morning briefings, runs my weekly reviews, and logs its own work, sometimes (laughs) while I’m not even at my Mac. Today, my Macpreneur friend, I’m opening up the hood, and by the end of this episode, you will know exactly what separates an AI agent from an AI-powered automation and the safest way to taste agent power today.
Welcome to Macpreneur (00:37)
Nova AI Welcome to Macpreneur, the show for seasoned solopreneurs looking to streamline their business on a Mac. Unlock the secrets to saving time and money with your host and technology mentor, Damien Schreurs.
The Story of VAL’s Evolution (00:50)
Damien Schreurs Hello, hello, and welcome to Macpreneur. I truly appreciate you tuning in today. This is the final solo episode of Your Mac, Your AI Stack series, part of Season 7, The AI-Enhanced Macpreneur. If you missed episode 172 on the AI Capability Matrix, or episode 173 on PTCF, or episode 174 on AI Assistants, the links are in the show notes. I promised that in the final episode of the series, I would show you how you can get agent-level power at a fraction of the cost. How? Yes, with AI-powered automations, and today, we draw that line precisely. And I promised that we will talk about what happens when an assistant graduates into an agent, and that is VAL’s story. As I said at the top of the show, VAL stands for Virtual Assistant Liaison. The first iteration was a simple Claude project with a custom instruction that was following the PTCF framework, so persona, task, context, and format. However, it doesn’t have autonomy, and it had very limited capabilities. Phase two was to attach some MCP connectors to the Claude desktop app. So for instance, I’m using NotePlan. I’m using Timing. I’m also using Apple Notes for some of the stuff that I use for Macpreneur and Easy Tech. That gave the Claude project the ability to work on data, even though I still had to babysit it, tell it exactly what to do, and confirm some of the actions it wanted to make. The third iteration was to convert that into CoWork as a project as well, having direct access to a specific folder on my Mac. The big advantage of switching to CoWork was that I could instruct it to do multi-tasks. So basically, I created standard operating procedures that VAL can follow, and it reads the SOP, builds its different step-by-step, so, um, step one, step two, step three, with different milestones, and then executes the work independently, and that meant that it reduced a lot of me babysitting Claude and start delegating outcomes.
How VAL’s Memory Cascade System Works (04:04)
Damien Schreurs So the way VAL works is that I’ve built the knowledge base in NotePlan. In NotePlan, I have created a VAL folder, and in that folder, there are subfolders for context, for archive notes, for standard operating procedures, but also for a memory system. The latest iteration of that memory system is that there is a note for two days’ work, so there is an M1 Today, an M2 This Week, an M3 This Month, an M4 This Quarter, and then an M5 This Year, and an M6 Previous Years. And what I set up as a system is a memory cascade system. Whenever I interact with Claude and with VAL, in particular, it is instructed to log the work that we do together in M1 Today. And then, at the end of the day, there is a scheduled task that runs automatically, and that scheduled task is looking at what was written in M1 Today, then taking what is the most important, summarizing it, and moving what’s required in M2 This Week versus throwing away what is not super useful. But on top of that, scheduled task, at the end of the day, is looking at the whole system. So has the folder structure changed? Did we use a new tool or a new MCP server that needs to be in a note somewhere? So it’s basically self-learning and self-improving the system every evening, and at the end of the week, on Sunday evening, 10:30 PM, there is a weekly…… task that actually does the same thing, look at what happened during the week, and then pushes everything from M2 this week to M3 this month. And on top of that, I have now tasks that are not scheduled actually that I fire up manually, one in the morning for my morning brief. As soon as I start it, as I launch it manually, it will look at a Google Doc that contains all the unread emails from the past 24 hours. It also looks at my weekly note with my weekly tasks and goals, and for the day’s note, a note that is created every day, there is a section in that note where Val can put its own briefing. At the end of that process, it will suggest three tasks that, based on my goals or the context that it has about me, the three tasks that it suggests would be the most valuable for me to tackle for the day. It seems like a great system. The reality is, it’s a work in progress. Where I am today is not where I was a week ago or a month ago. It is evolving all the time. As soon as I add a new MCP connector, it changes things, SOPs, standard operating procedures. As I said, a large language model is a human dialogue simulator that does not really understand what it’s saying. From time to time, it misses things, or the context that it is given is either not, uh, large enough or it’s not clear. I’m constantly refining that system, but the key is this, uh, autonomy, the fact that I can give it a goal or I can tell it, “I want you to help me with these kinds of tasks,” and it will use the tools by itself and it will create the plan to reach the goal.
The Anatomy of an AI Agent: LLM + Harness (08:34)
Damien Schreurs If we look at the anatomy of an agent, it’s a collection of two things. It’s an LLM, but with what is called a harness, and the harness is the deterministic part of the whole system. So, a large language model is probabilistic, it’s just nondeterministic, so it can make stuff up and it doesn’t understand what we ask it to do and what it’s doing. And the reason why an agent works is actually thanks due to that harness. The harness is more like a program, a computer program, that is overseeing what’s happening with the LLM. So, you may have seen that LLMs have included, quote-unquote, reasoning or thinking. This is internal to the LLM, and it is fallible. The harness sits on top of the LLM, and it double-checks what it does, and it guides the hell- LLM into choosing the right tool, using the tools properly. It’s computer code that keeps the model honest and on task. Sometimes I hear people say that large language models are computer programs. No, that’s not correct. A large language model is a simulator. It’s a predictor of tokens. It’s matrix multiplication and using weights, and the output of that model is probabilistic, the complete opposite of a computer program. But the harness is the computer program thing around the large language model. On top of that, harnesses can initiate work, so that’s how schedule task work. A schedule task is a computer program that, in a deterministic fashion, at a set time, kicks in and triggers the LLM. The LLM itself cannot do that. An LLM needs to be prompted. It needs to be triggered. It cannot trigger itself.
AI Agents vs. AI-Powered Automations (11:22)
Damien Schreurs And I use agents, so an LLM that is conducted by a harness, when I have a nondeterministic process, when the way to reach a goal could change from time to time, when you want something that could be probabilistic, generating a draft of a blog post or generating an image. It could be analyzing documents or analyzing a bunch of emails. So, it’s a large language model which is probabilistic, but the sequence of steps is actually known and is fixed. If it’s a deterministic process that includes nondeterministic steps, that’s when you use AI-powered automation. Knowing that difference is what will help you save on tokens, because… If most of the steps are sequential, deterministic steps, with some logic, if, then, else, but at one point, you have one or two AI steps, then you should use an AI powered automation. So tools like n8n, Zapier, make.com are better suited for that. An example of an AI powered automation where I thought I needed an agent and the agent built (laughs) the AI power automation for me is the news brief gathered every day for me, twice a day, 7:25 AM, 7:25 PM. And so, there are two scripts that aren’t fully deterministic and that run at very set times. So, one is using a tool called Block Watcher, and every couple of hours, it fetches new articles from a set of blogs. And then, at 7:25 AM and 7:25 PM, there is another script that goes through all the unread articles, but then uses Haiku, so the Claude AI model, via the API to summarize and curate news articles. And it’s not the AI who is doing that, it is a shared script that is pushing the output of the LLM into a Google Doc. I don’t rely on an LLM to not mess up the edition of a Google Doc, I have a fixed and unbreakable process to make sure that whatever the LLM has outputted as the curated list of articles that I should pay attention to, that is then going into the Google Doc. And the only issues that there could be, my computer not having internet connectivity, or being down, or because I decided to (laughs) upgrade macOS. I wanted to experiment with Open Claw to do that process for me. I thought I could set up Open Claw as an agent that would actually go through the said blogs that I chose. And I thought that I could tell Open Claw, “Okay, do this, fetch the… Go to the website, fetch the thing, curate, and then create the Google Doc or update the Google Doc,” and so on. I realized it was not stable enough. It would not properly do the thing from time to time. And so, then we iterated, and then I ask Open Claw, “Can you create the scripts that would help you do that?” Until I realized that I didn’t need Open Claw anymore (laughs) because I had inadvertently created an AI powered automation, and I didn’t need an agent for that. When the way to go from point A to point B is not known, then an agent is perfect. Another analogy, if you need to go to a city center and that city center has a train station, it’s better to take the train, because you know it will go at different stations until the last one. And that would be an AI powered automation if, during the train journey, you’re using AI along the way, but you continue going through the same path. And if you don’t know whether you need to go to shop A and shop B or shop C before arriving at the city center, then you would take the car, and that would be the equivalent of using an agent.
Navigating Agent Risks and a Safe Starting Point (16:46)
Damien Schreurs Now, there are risks in using agents, is you don’t really control what the agent will do. It could misuse a tool, and it could be instructed, it’s called prompt injection attacks, it could be instructed to exfiltrate information. Before going more deep into Claude Cowork or Codex or even OpenClaw and, and Nous Hermes, there is today an agent that is quite safe to use and does not require specific skills, and it’s called Deep Research. It’s available in ChatGPT, in Claude, in Gemini, in CoPilot. Deep Research is a specialized agent instructed to perform internet research, web research, and it has very limited capabilities or tool use. It can basically do search queries. You could ask a normal LLM in a normal chat to research a few things for you, but it will do a few queries, and that’s it. But with Deep Research, it will prepare a research plan, show you the research plan, allow you to edit the research plan, then anonymously…… go through a five, six, seven step research plan, and create a very detailed report. We’re talking about 10 to 30 pages of report. And usually, when the research agent starts, it takes 10, 15, 20 minutes to complete the research. But it’s 20 minutes where you can do something else, and if you had to do it by yourself, it would have taken you hours. It’s great how much time we can save by using a deep research agent.
The “Lethal Trifecta” of Agent Security (19:12)
Damien Schreurs Now, if we go back to the typical agents, Claude Cowork, Claude Code, Codex, but also, we have the OpenClaw and Nous Hermes, and so on. Going back to security, there is a security researcher called Simon Willison who has introduced the concept of the lethal trifecta, basically three ingredients that increases the risk of getting hacked through agent. Those three ingredients are, number one, access to your private data, number two, exposure to untrusted content, and number three, the ability to communicate with the outside world. The idea is to reduce the risk or to stay safe, you should only allow two of those three things, not the three together. So if you allow access to private data, and you give it the ability to communicate with the outside world, do not expose your agent to untrusted content. And if you need to have your agent read your emails, it will have access to your private data, because it will have access to your emails. Then you have to remove the ability to communicate with the outside world. If you gave access to your Gmail to your agent, the agent should have read-only access, not write. It c- it cannot be allowed to send emails on your behalf, and it cannot have another means of communication. So, no Telegram, (laughs) no WhatsApp, no Discord. Yes, if you want your agent to process private data and have exposure to untrusted content, it should not be allowed to communicate with the outside world. So, you pick two, and you have to put in place guardrails so that the third one is not available. But as soon as you have the three, the trifecta, then you can be exposed to great risks by using AI agent. So examples, an attacker literally email your agent, and then tell it to exfiltrate information on it behalf. But it could also do some web browsing, then visit a webpage, and in that webpage there are some instructions for the agent to exfiltrate data or actually compromise your computer. And also, back in episode 172, I introduce an incident that happened with, with VAL, where it overwrote a note in NotePlan. So, the good news is I was able to revert the situation because of the revision system in NotePlan. As soon as you give an agent the ability to use tools, know exactly what it can do and make sure that you can undo anything that it does. So, it’s either revisions, or you have a time machine backup and you are able to revert some files or folders to an earlier state. My advice, start more with Cowork or Codex, because it’s easier to limit the damage that the agent can do. With OpenClaw and Nous Hermes, they have much more autonomy, and you have to be mindful about how you configure them and what you let those agent do on your behalf.
Recap and Your Next Step (23:27)
Damien Schreurs So, to recap, an AI agent is a non-deterministic large language model combined with a deterministic harness. It is the opposite or the complementary to an AI-powered automation. The easiest way to test agents today is to start using deep research more and more with projects in Claude or ChatGPT, and also with Gemini Gems, it’s possible to have deep research as part of their toolbox. So basically, you could convert an assistant into an AI research agent at the same time. Finally, be mindful of the lethal trifecta, having access to private data, being in contact with untrusted content, and being able to communicate with the outside world. Pick two, but do not let your agent fulfill the three criterias. And now that you’ve seen the full picture, chatbots, assistants, and agents, you might be wondering where you actually stand today and what your smartest next step is. I built a free two-minute quiz. It’s called Discover Your AI Stage at macpreneur.com/aistage, in one word. And I will put the link in the show notes. So once again, it’s macpreneur.com/aistage.
What’s Next for the Macpreneur Podcast (25:17)
Damien Schreurs So that’s a wrap for this series called Your Mac, Your AI Stack. We went from the AI capability matrix to PTCF to AI assistants and today, agents. What’s next? There will be a few more guest interviews, and then season seven will close while I build a buffer for season eight. And I can already reveal the main theme for season eight, and it’s gonna be called The Lean Macpreneur. I’m going back to the roots of the Macpreneur show about helping you streamlining your business, but also leverage my pre-EasyTECH expertise when I was working in lean product development for a large Fortune 500 company. If you run your solopreneur business efficiently on a Mac and you want to share tips and tricks on how you do that, I would love to have you on the show. Just visit macpreneur.com/apply. If you are already on Podmatch, just click on the button, and you will land on the show profile. Otherwise, fill the application form, and I will get back to you within a few days. Once again, it’s macpreneur.com/apply. And until next time, I’m Damien Schreurs wishing you a productive day.
Outro (26:48)
Nova AI Thank you for listening to the Macpreneur podcast. If you’ve enjoyed the show, please leave a review and share it with a friend right now.