AI Tools Field Guide 2026: Best LLMs for Coaches & Consultants | Be Known
Be Known, LLC

AI Tools & LLM
Field Guide

Your plain-English reference for navigating the AI landscape — which tool does what, and when to reach for it.

2026 Edition · Living Document
💡

The AI space moves fast — and the tool names alone can cause confusion. Claude vs. Claude Code vs. Claude Cowork. Perplexity vs. Perplexity Computer Use. Gemini vs. Google AI Studio vs. NotebookLM. This guide cuts through the noise. Whether you're brand new to AI or already a daily user, use this as your go-to reference for picking the right tool for the job. In 2026, coaches and consultants who match the right AI tool to the right task report saving 8–12 hours per week compared to those using a single general-purpose assistant for everything.

👆 Click any tool anywhere in this guide to see full details.

🌱
Tier 1
Just Getting
Started

You're exploring AI for the first time, or you use it occasionally for simple tasks. You want something that just works — no setup, no tech knowledge required. For most beginners, Claude.ai and Perplexity together cover 90% of everyday AI tasks — writing, research, and Q&A — without any paid subscription.

💬Claude.ai Chat
🔍Perplexity
💬ChatGPT
Gemini
Tier 2
Daily User &
Experimenter

You use AI regularly and want to go deeper — running research workflows, automating tasks, testing prompts, or building things without writing code. Intermediate users who combine Perplexity for research with Claude for synthesis consistently produce higher-quality strategic outputs than those relying on a single model.

🌐Perplexity Pro
🧪Google AI Studio
📓NotebookLM
🤖Manus AI
🏗️Lovable
🖥️Claude Cowork
🚀
Tier 3
Power User &
Builder

You build systems, write prompts like a developer thinks in code, automate complex workflows, and want full control over model behavior and integrations. Power users who integrate Claude Code or GPT-4o via API into automated workflows reduce manual task execution by an average of 60–80% compared to using chat interfaces alone.

💻Claude Code
🌐Perplexity Computer Use
🔧OpenAI API / GPT-4o
⚙️Replit
⚙️Base44
🔗Zapier AI
🔗Make AI
Tool Best For Key Strengths Limitations Skill Level Cost
🟢 Anthropic · Claude Family
Claude.ai Chatclaude.ai Writing, strategy, copy editing, long documents Best reasoningLong contextNuanced writing No internet by default; can't act on your computer 🌱 Beginner Free / Pro $20/mo
Claude CodeTerminal / CLI Writing, debugging & refactoring code; developer workflows Reads codebaseRuns commandsEdits files Requires developer setup; not for non-coders 🚀 Power User API usage-based
Claude CoworkDesktop Agent Letting Claude operate your computer on your behalf Sees your screenTakes actionsBackground agent Memory & CPU intensive; system must stay open ⚡ Intermediate Claude Pro / Teams
🔵 OpenAI · ChatGPT Family
ChatGPTchat.openai.com General chat, brainstorming, image generation, code help Huge ecosystemGPTs / pluginsImage gen Quality varies; hallucinations; Pro needed for best models 🌱 Beginner Free / Plus $20/mo
GPT-4o / APIplatform.openai.com Building apps, custom chatbots, real-time voice integrations Voice modeVisionAPI ecosystem Requires development knowledge; cost adds up at scale 🚀 Power User Usage-based API
🟣 Perplexity Family
Perplexityperplexity.ai Real-time research, cited answers, fact-checking Live web searchCited sourcesFast research Less suited for creative writing or long-form generation 🌱 Beginner Free / Pro $20/mo
Perplexity Computer UsePro feature Autonomous web browsing, multi-step research tasks Browses the webMulti-step tasksResearch agent Still maturing; limited to browser actions 🚀 Power User Perplexity Pro
🔴 Google AI Family
Geminigemini.google.com Everyday AI chat, Google Workspace integration Google integration1M token contextMultimodal Less nuanced reasoning than Claude; ecosystem lock-in 🌱 Beginner Free / Advanced $19.99/mo
Google AI Studioaistudio.google.com Testing prompts for free, prototyping, saving workflows Free sandbox1M tokensPrompt testing Developer-oriented UI; Gemini models only ⚡ Intermediate Free (generous limits)
NotebookLMnotebooklm.google.com Uploading documents and asking questions about them Document Q&AAudio summariesSource-grounded Limited to your uploaded sources; not general-purpose ⚡ Intermediate Free
🟠 AI App & Web Builders
Lovablelovable.dev Building full web apps from a prompt — no coding required Full apps from textSupabase integrationFast iteration Less control than coding from scratch; best for MVPs ⚡ Intermediate Free tier / Paid plans
Replitreplit.com Running, editing, deploying code in the browser Cloud IDEDeploy instantlyAI autocomplete Slower than local dev for large projects; needs some coding knowledge 🚀 Power User Free / From $20/mo
Base44base44.com AI-powered business app builder with built-in database & auth Built-in DBNo backend neededBusiness apps Newer platform; smaller ecosystem ⚡ Intermediate Paid plans
🤖 Autonomous AI Agents
Manus AImanus.im Fully autonomous research and multi-step task execution Autonomous agentMulti-step tasksMCP integrations Slower; results need review; some account restrictions ⚡ Intermediate Waitlist / Invite
⚙️ AI-Powered Automation
Zapier AIzapier.com Connecting apps and using AI to make decisions in automations 6,000+ integrationsAI stepsNo-code Cost scales with volume; AI steps can be inconsistent ⚡ Intermediate Free tier / From $19.99/mo
Makemake.com Visual automation with more control and lower cost than Zapier Visual flow builderCheaper at scaleMore flexibility Steeper learning curve; fewer native integrations ⚡ Intermediate Free / From $9/mo

X-axis: Task Type  ·  Y-axis: How much the AI acts on its own vs. you staying in control

AUTONOMOUS + CREATIVE AUTONOMOUS + TECHNICAL ASSISTED + CREATIVE ASSISTED + TECHNICAL ◀ Creative / Content Technical / Dev ▶ ▲ More Autonomous ▼ You Stay in Control Manus AI Perplexity Computer Use Claude Cowork Claude Code GPT-4o API Claude.ai Chat ChatGPT Perplexity Notebook LM Google AI Studio Gemini Lovable Replit Zapier AI Make Base44
Dark green = Most autonomous / technical
Mid green = Intermediate
Bright green = Accessible / assisted
🤖 LLM
Large Language Model. The AI "brain" behind most chat tools. Examples: GPT-4, Claude 3, Gemini. They predict and generate text based on your input.
🎛️ Model
The specific version of an AI. Claude Sonnet vs. Opus, GPT-4o vs. GPT-3.5. Bigger models = more capable but slower and costlier.
📝 Prompt
The instruction or question you type to the AI. Better prompts = better results. Think of it as how you talk to the tool.
🔑 API
Application Programming Interface. A way for developers to connect AI models directly into their own apps or workflows. Requires coding or automation tools.
🪙 Token
How AI measures text. Roughly 1 token = 1 word (or part of a word). API costs are often charged per token used. More tokens = more cost.
📏 Context Window
How much text the AI can "see" and remember at once in one conversation. Claude and Gemini lead here with 100K–1M token windows.
🤝 MCP
Model Context Protocol. A standard that lets AI connect to external tools and services — like your Google Drive, Slack, or ad accounts — so it can read and act on real data.
🕵️ Agent / Agentic AI
An AI that doesn't just answer questions — it takes actions. Agents can browse the web, write files, click buttons, or complete multi-step tasks on your behalf.
🖥️ Computer Use
A capability that lets an AI see and control a computer screen — clicking, typing, navigating — as if it were a human at the keyboard.
🛠️ Skill / System Prompt
Hidden instructions that tell an AI how to behave before you ever type a message. Used to give AI a role, set rules, or load context automatically.
🌡️ Temperature
A setting that controls how "creative" vs. "precise" the AI is. Low temperature = consistent, factual. High temperature = more surprising, varied output.
🔮 Hallucination
When an AI makes up something that sounds plausible but isn't true. All LLMs do this. Always verify important facts, especially with ChatGPT and older models.
🔓 RAG
Retrieval-Augmented Generation. A technique where AI pulls from your own documents or database before answering — reducing hallucinations. NotebookLM uses this.
Fine-tuning
Training an existing AI model on your own data to make it better at your specific tasks. Requires technical setup. Usually done via API.
🔄 Multimodal
An AI that can work with more than just text — images, audio, video, documents. ChatGPT, Claude, and Gemini are all multimodal.

Full Content Summary · For Screen Readers & Search Engines

About This Guide

This is Be Known, LLC's AI Tools Field Guide — a plain-English reference for coaches, consultants, and expert-based businesses navigating the AI landscape in 2026. It covers 15+ AI tools across five categories: conversational AI assistants, AI research tools, AI app builders, autonomous AI agents, and AI-powered automation platforms. The guide is designed to help business owners identify which tool is right for each task, organized by skill level and use case. In 2026, coaches and consultants who match the right AI tool to the right task report saving 8–12 hours per week compared to those using a single general-purpose assistant for everything. For more digital marketing strategies for coaches and consultants, visit the Be Known blog.

Skill Tier Overview

This guide organizes AI tools into three skill tiers to help users identify the right starting point:

  • Tier 1 — Just Getting Started (Beginner): Users exploring AI for the first time or using it occasionally for simple tasks. Recommended tools: Claude.ai Chat, Perplexity, ChatGPT, Gemini. For most beginners, Claude.ai and Perplexity together cover 90% of everyday AI tasks — writing, research, and Q&A — without any paid subscription.
  • Tier 2 — Daily User & Experimenter (Intermediate): Users who use AI regularly and want to run research workflows, automate tasks, test prompts, or build things without writing code. Recommended tools: Perplexity Pro, Google AI Studio, NotebookLM, Manus AI, Lovable, Claude Cowork. Intermediate users who combine Perplexity for research with Claude for synthesis consistently produce higher-quality strategic outputs than those relying on a single model.
  • Tier 3 — Builder & Power User: Users who build systems, write prompts like a developer thinks in code, automate complex workflows, and want full control over model behavior and integrations. Recommended tools: Claude Code, GPT-4o / API, Replit, Make, Zapier AI. Power users who integrate Claude Code or GPT-4o via API into automated workflows reduce manual task execution by an average of 60–80% compared to using chat interfaces alone.

Full AI Tool Comparison

Anthropic · Claude Family

  • Claude.ai Chat (claude.ai): Best for writing, strategic thinking, copy editing, summarizing long documents, and nuanced Q&A. Key strengths: best-in-class reasoning, 200K token context window, nuanced human-like writing, excellent instruction following, image and document understanding. Limitation: no live internet access by default; cannot take actions on your computer. Skill level: Beginner. Cost: Free / Pro $20/mo / Teams $25/user/mo.
  • Claude Code (Terminal/CLI): Best for writing, debugging, and refactoring code files. Reads entire codebases, runs terminal commands, edits files, and completes developer tasks autonomously. Limitation: requires developer setup; not suitable for non-coders. Skill level: Power User. Cost: Anthropic API, usage-based.
  • Claude Cowork (Desktop Agent): Lets Claude operate your computer on your behalf — sees your screen, takes actions, and runs as a background agent. Limitation: memory and CPU intensive; system must stay open. Skill level: Intermediate. Cost: Claude Pro or Teams.

OpenAI · ChatGPT Family

  • ChatGPT (chat.openai.com): Best for general chat, brainstorming, image generation, and code help. Key strengths: largest ecosystem, GPTs and plugins, image generation. Limitation: quality varies; hallucinations; Pro needed for best models. Skill level: Beginner. Cost: Free / Plus $20/mo.
  • GPT-4o / API (platform.openai.com): Best for building apps, custom chatbots, and real-time voice integrations. Key strengths: voice mode, vision, large API ecosystem. Limitation: requires development knowledge; cost adds up at scale. Skill level: Power User. Cost: usage-based API.

Perplexity Family

  • Perplexity (perplexity.ai): Best for real-time research, cited answers, and fact-checking. Key strengths: live web search, cited sources, fast research. Limitation: less suited for creative writing or long-form generation. Skill level: Beginner. Cost: Free / Pro $20/mo.
  • Perplexity Computer Use (Pro feature): Best for autonomous web browsing and multi-step research tasks. Key strengths: browses the web, multi-step tasks, research agent. Limitation: still maturing; limited to browser actions. Skill level: Power User. Cost: Perplexity Pro.

Google AI Family

  • Gemini (gemini.google.com): Best for everyday AI chat and Google Workspace integration. Key strengths: Google integration, 1M token context, multimodal input. Limitation: less nuanced reasoning than Claude; ecosystem lock-in. Skill level: Beginner. Cost: Free / Advanced $19.99/mo.
  • Google AI Studio (aistudio.google.com): Best for testing prompts for free, prototyping, and saving workflows. Key strengths: free sandbox, 1M token context, prompt testing. Limitation: developer-oriented UI; Gemini models only. Skill level: Intermediate. Cost: Free (generous limits).
  • NotebookLM (notebooklm.google.com): Best for uploading documents and asking questions about them. Key strengths: document Q&A, audio summaries, source-grounded answers. Limitation: limited to uploaded sources; not general-purpose. Skill level: Intermediate. Cost: Free.

AI App & Web Builders

  • Lovable (lovable.dev): Best for building full web applications from a prompt with no coding required. Key strengths: full apps from text, Supabase integration, fast iteration. Limitation: less control than coding from scratch; best for MVPs. Skill level: Intermediate. Cost: Free tier / Paid plans.
  • Replit (replit.com): Best for running, editing, and deploying code in the browser. Key strengths: cloud IDE, instant deployment, AI autocomplete. Limitation: slower than local development for large projects; requires some coding knowledge. Skill level: Power User. Cost: Free / From $20/mo.
  • Base44 (base44.com): Best for building AI-powered business apps with built-in database and authentication. Key strengths: built-in database, no backend needed, business app templates. Limitation: newer platform with smaller ecosystem. Skill level: Intermediate. Cost: Paid plans.

Autonomous AI Agents

  • Manus AI (manus.im): Best for fully autonomous research and multi-step task execution. Key strengths: autonomous agent behavior, multi-step task completion, MCP integrations. Limitation: slower than manual workflows; results need human review; some account restrictions. Skill level: Intermediate. Cost: Waitlist / Invite-only.

AI-Powered Automation

  • Zapier AI (zapier.com): Best for connecting apps and using AI to make decisions in automations. Key strengths: 6,000+ integrations, AI decision steps, no-code interface. Limitation: cost scales with volume; AI steps can be inconsistent. Skill level: Intermediate. Cost: Free tier / From $19.99/mo.
  • Make (make.com): Best for visual automation with more control and lower cost than Zapier. Key strengths: visual flow builder, cheaper at scale, more flexibility. Limitation: steeper learning curve; fewer native integrations. Skill level: Intermediate. Cost: Free / From $9/mo.

Use-Case Quadrant: How to Read It

The Use-Case Quadrant maps each AI tool across two axes. The X-axis represents the task type, ranging from creative tasks (writing, brainstorming, strategy) on the left to technical tasks (coding, automation, data processing) on the right. The Y-axis represents the level of AI autonomy, from tools where you stay in control (assisted) at the bottom to tools that act on their own without constant input (autonomous) at the top. Tools in the top-right quadrant (Autonomous + Technical) are the most powerful and the most technically demanding — including Claude Code, GPT-4o API, Manus AI, and Replit. Tools in the bottom-left quadrant (Assisted + Creative) are the most accessible — Claude.ai Chat, ChatGPT, and Gemini.

Glossary of Key AI Terms

  • LLM (Large Language Model): The underlying AI technology that powers tools like Claude, ChatGPT, and Gemini. An LLM is trained on vast amounts of text and learns to predict and generate human-like language. The model is the engine; the chat interface (Claude.ai, ChatGPT.com) is the vehicle.
  • Context Window: The amount of text an AI can process in a single conversation or task — measured in "tokens" (roughly 3/4 of a word). A larger context window means the AI can handle longer documents or more complex instructions. Claude currently offers up to 200,000 tokens; Gemini offers up to 1 million tokens.
  • API (Application Programming Interface): A way to connect software applications so they can communicate. When an AI tool offers API access, developers can build custom apps, automations, and integrations using the AI's underlying model — without using the standard chat interface.
  • Prompt: The instruction or question you give an AI model. Prompt quality directly determines output quality — more specific, well-structured prompts produce better results than vague requests.
  • AI Agent: An AI system that can take actions autonomously — browsing the web, clicking buttons, running code, or completing multi-step tasks — without needing a human to approve each step. Claude Cowork and Manus AI are examples of AI agents.
  • MCP (Model Context Protocol): An open standard that lets AI models connect to external tools, databases, and services in a consistent way. MCP makes it possible for AI agents like Manus AI to interact with a wide range of third-party tools.
  • Hallucination: When an AI model generates information that sounds plausible but is factually incorrect or fabricated. All LLMs can hallucinate, which is why cited-source tools like Perplexity are preferred for research tasks where accuracy is critical.
  • Token: The unit of text that AI models process. Roughly equivalent to 3/4 of a word. Pricing for API access and context window limits are both measured in tokens.

How to Pick the Right AI Tool for Your Business

To choose the right AI tool for your coaching or consulting business, follow these five steps:

  1. Identify your skill level — beginner, intermediate, or power user — using the Skill Tier section of this guide.
  2. Identify your primary task type — writing and strategy, real-time research, building apps or automations, or autonomous task execution.
  3. Use the comparison table to match your task to the right tool, reviewing strengths, limitations, and pricing.
  4. Use the Use-Case Quadrant to visually confirm your choice based on how autonomous vs. assisted, and how creative vs. technical, your workflow needs to be.
  5. Start with one tool and master it before adding others. For most coaches and consultants, the right AI marketing strategy starts with Claude.ai or Perplexity as a foundation, then expands from there.

Resource produced by Be Known, LLC — Digital marketing agency building client acquisition systems for coaches, consultants, and expert-based businesses. beknownonline.com

Be Known | AI Updates We Can Use Now
Practical AI update

Three AI upgrades we can use now.

These are not future ideas. They can reduce repeat mistakes, speed up complex work, and make our websites easier for AI search tools to understand. Updated July 2026 for Opus 5, which moved part of the original skill into the tool itself.

See the three upgrades
Upgrade 01

Teach AI once.

AI often repeats the same mistake because the correction disappears inside an old chat. A lessons file turns each useful correction into a rule the system can reuse.

1

Spot the mistake

A person corrects the output, or the AI catches its own error.

2

Save one clear rule

The lesson explains what went wrong and what to do next time.

3

Reuse the lesson

The rule is added to future prompts, so the same error is less likely to return.

Example

Weak lesson

“Make the design better.”

Better lesson

Specific and reusable

“For Be Known vertical graphics, keep the bottom third clear for subtitles and use the uploaded logo without redrawing it.”

Copy this instruction
## Self-Learning When I correct you, or when you catch a mistake, pause before continuing. Add one short rule under “## Lessons” only when the lesson is specific, reusable, supported by evidence, and likely to prevent future rework. Use those lessons in every later step of the task. ## Lessons - Add new lessons here.
ChatGPT or Claude

Paste it at the top of a new project instruction, custom instruction, or long-running project chat. Keep the same chat or project for related work so the lessons remain available.

Claude Code or Codex

Add it to the project instruction file, such as CLAUDE.md, AGENTS.md, or the skill file used by that workflow. The agent will read it on each run.

Simple team use

Create one shared LESSONS.md file for each repeatable workflow. After a correction, add one approved lesson, then include that file in future prompts or agent runs.

One rule we learned the hard way: not every correction deserves to become permanent. A lesson file that logs everything gets long, contradicts itself, and quietly slows every future run. Keep a one-off as a note for that job only. Promote it to a standing rule when the cause is understood, the fix is specific and reusable, nothing already covers it, and the mistake either repeated or cost real time.

Upgrade 02

Use agent loops.

An agent loop does more than complete a task. It checks the result, corrects problems, and carries the lesson into the next attempt.

1

Plan

Define the goal and what a good result must include.

2

Do

The agent completes one clear, bounded task.

3

Check

It compares the result with the rules and success criteria.

4

Correct

It fixes gaps instead of passing weak work forward.

5

Learn

It saves a useful lesson, then starts the next cycle smarter.

PLAN → DO → CHECK → CORRECT → LEARN → REPEAT
A
Content

Build and review a campaign

One agent researches buyer questions. Another drafts hooks. A reviewer checks each idea against the ICP, brand voice, and offer before anything is published.

B
Research

Compare sources in parallel

Several agents collect evidence. A lead agent removes weak sources, resolves conflicts, and turns the findings into one recommendation.

C
Website work

Make changes without guessing

An agent edits one section. It then checks the live page, tests links and forms, fixes defects, and records what caused them.

D
Lead generation

Research and qualify prospects

Agents find and score prospects against clear rules. A human reviews the final list and message before outreach begins.

Use the smallest loop that works. Several agents working in parallel is the most expensive version of this, and the newest tools only do it when you actually ask. Most jobs need one worker, clear success criteria, and one honest check. Reach for a team of agents when the work genuinely splits into separate pieces that later have to be joined back together.

Version 4.0 · rewritten for Opus 5

The Intelligence Layer

This is a practical operating model for complex work: research, content, websites, automation, coding, migrations, campaign planning, and long multi-step jobs. The strongest model does the hard thinking, routing, and final quality check. Right-sized models do the clearly specified work. Nothing ships on confidence alone.

The version we shared first was written before Opus 5. That release moved a large part of the old skill into the tool itself and disproved one of its main assumptions, so the skill below is a rewrite rather than a patch.

Strongest model: route, judge, verify
Right-sized model: execute clear tasks
Every route: check, correct, learn
What Opus 5 changed

Version 2.2 saidHand-build the scaffolding: task lists, worker states, progress tracking, approval steps.

NowThe tool does this natively. Task tracking, isolated workers, completion alerts, and plan approval are built in. Rebuilding them in a skill just spends tokens describing features you already have.

Version 2.2 saidDelegate aggressively by default — push as much work as possible to cheaper models.

NowDelegation is deliberate, not automatic. The agent does not spin up a team of sub-agents unless you ask for it. Default to one worker doing the job properly, and ask for parallel agents when the work truly splits.

Version 2.2 saidNever name specific models, because the lineup keeps changing.

NowThe lineup is stable enough to name: Opus 5 for architecture and independent review, Sonnet 5 for bounded execution, Haiku 4.5 for mechanical work. Vague tiers made people guess.

Version 2.2 saidThe way to save money is to use a cheaper model.

NowWe measured it, and that was the smaller lever. The dominant cost is the conversation being re-sent on every single step. Model choice matters; what you drag along with you matters more.

1. Route before you build

Pick the lightest thing that can work: a script, one direct edit, one specialist, one bounded loop, or full orchestration. Do not scale up just because you can.

2. Add only the missing part

Check what the tool already does before writing process for it. A skill should carry judgment, not restate built-in features.

3. Match the model to the judgment

Opus 5 for architecture and independent review. Sonnet 5 for bounded execution. Haiku 4.5 for mechanical work. Bulk input gets summarised locally first.

4. Verify against the real thing

Open the file, click the button, run the command. An executor's summary is a claim. Whoever did the work never signs it off alone.

5. Repair once, then escalate

Send back only the failed evidence and the affected scope. One evidence-driven fix, then stop and raise it rather than looping.

6. Keep risky actions gated

Publishing, sending, spending, deleting, and deploying stay behind human approval, whatever the agent believes it has finished.

The real saving is context, not model choice.

We audited our own usage across 907 sessions. About 94% of all tokens were not new work at all — they were the existing conversation being re-sent on every step, and the five heaviest sessions alone accounted for a third of everything. Anything pulled into a long session is paid for again on every later step, so a file read early in a big job can cost hundreds of times its own size. Practical version: read the part you need instead of the whole file, summarise bulk input before it enters the conversation, prefer a script that returns a number, and finish one job per session instead of carrying five jobs' worth of history forward.

How to install it
Claude Code

Save the file as .claude/skills/intelligence-layer/SKILL.md, then invoke it in your task prompt.

Codex

Save it in the project’s skills folder or reference it from AGENTS.md.

Other AI tools

Paste the full skill into project instructions or a reusable system prompt. It is written to stand alone.

Useful triggers: “Use the intelligence layer,” “Route this properly,” “Architect and QA this,” “Verify this against the real output,” or “Orchestrate this in parallel” when you genuinely want several agents. Note the change: the last one is now a request you have to make, not a default.

View and copy the full Intelligence Layer skill — version 4.0
---
name: intelligence-layer
version: "4.0"
description: >
  Use ONLY when the user asks for orchestration, fan-out, or parallel agents on
  complex work with dependent workstreams and integration risk. Creates bounded
  task packets, dispatches isolated workers, integrates evidence, and escalates
  on evidence. Never self-selected.
license: MIT
---

# Intelligence Layer

Coordinate complex work without absorbing specialist procedure. Routing picks the
lightest reliable route, specialist skills own domain procedure, and verification
owns evidence.

Version 4.0 replaces the 2.x releases. Those versions hand-built process that a
modern agent harness now supplies natively, and they assumed aggressive
delegation was the default. Both assumptions are gone.

## Already native - add only the delta

Before writing any process, check whether the tool already owns it. Do not write
a contract field, state machine, or queue that only restates one of these:

| Concern | Handled natively |
|---|---|
| Acting on a clear request | System policy - act rather than re-plan |
| Scope discipline | System policy - deliver the requested scope, no quiet widening |
| Confirming risky actions | System policy - confirm hard-to-reverse or outward-facing steps |
| Honest reporting | System policy - report failures with output, name skipped steps |
| Task state | Native task create / update / list / get / output / stop |
| Waiting on work | Background workers re-invoke on completion - never poll |
| Waiting on external state | Native monitor / wakeup / cron for genuinely time-based work |
| Plan approval | Native plan mode enter / exit |
| Isolated worker | Native agent call with worktree isolation |
| Continuing a worker | Send a message to the existing worker - a new call starts cold |
| Specialist procedure | Native skill invocation |

## This route is user-requested only

The binding rule: do not spawn subagents unless the user asked. Scale,
multi-part scope, or a "thorough" framing is not a request.

Enter full orchestration only when the user asks for orchestration, fan-out,
parallel agents, or names an agent type. If the work would benefit and they have
not asked, say so in one line and offer it - then continue inline. Do not stall
the work waiting for permission to parallelise.

Even when requested, prefer a deterministic script, direct execution, a single
specialist, or the lite loop for anything narrow and mechanically testable.

## Choose the smallest reliable route

| Situation | Route |
|---|---|
| Exact, repeatable, mechanically checkable | deterministic - script it |
| Tiny, obvious, one-step work | direct |
| One domain skill owns the deliverable | specialist |
| One bounded deliverable, real risk of an avoidable miss | lite loop: criteria, one repair, evidence check |
| Dependent workstreams and integration risk, AND the user asked for orchestration | full orchestration below |

Do not promote work because multiple tools, agents, or skills exist. Choose the
verifier before the executor.

## Preflight and contract

1. Confirm authoritative inputs, specialist skills, verifier, rollback needs,
   available tools, and integration owner.
2. Create one contract ID. Keep the packet sparse; add a dependency graph only
   where it changes execution order.
3. Give each task a mutable scope, acceptance criteria, evidence, dependencies,
   and next state.
4. Select mechanical checks before model review. Define escalation and stop
   conditions before dispatch.

## Model roles

Route by required judgment, then pass the model explicitly when a worker is
authorised. Keep the executor and its reviewer in separate contexts, and never
let an executor review its own work.

| Role | Default | Use for |
|---|---|---|
| architecture | strongest model (Opus 5) | decomposition, interface ownership, escalation |
| independent QA | strongest model, fresh context | evidence review where a miss propagates |
| executor | mid tier (Sonnet 5) | bounded implementation, research, drafting |
| mechanical | fast tier (Haiku 4.5) | extraction, formatting, deterministic transforms |

Local models are not an executor tier. They are a pre-processing tier that keeps
bulk input out of the main context: large input, small output, cheap to re-check.
They cannot own judgment or approve a gate, and anything mechanical in their
output gets a deterministic re-check.

## Schedule and execute

- Default to one primary executor. Parallelise only tasks with separate owned
  files, artifacts, branches, or worktrees.
- Track state with the native task tools, not a hand-maintained list.
- Never poll. Background workers re-invoke on completion.
- Continue an existing worker rather than starting a fresh one that has to
  re-derive context you already paid for.
- Give each worker its task-local packet and state what it must return. A
  worker's report is not shown to the user, so relay what matters.
- Send bulk input - long transcripts, file batches, contact sheets - to a local
  model first so workers receive a digest instead of raw volume.
- Use scripts for deterministic inspection and transformation. Use agents only
  for implementation, review, or independent judgment.
- Require execution evidence: actual changes, paths, checks, results, artifacts,
  deviations, and unresolved issues.

## Context cost - this route's real expense

Measured across 907 sessions on one working setup: 94% of all tokens were
context being re-sent, and the five heaviest sessions consumed 34% of
everything. This route runs the longest sessions, so it owns that cost.

Anything pulled into context is paid for again on every remaining call, not
once. A 6k-token transcript read at call 20 of a 200-call run costs roughly 1.1M
tokens, not 6k. So:

- Read the narrowest slice that answers the question. Whole-file reads of things
  you only need one function from are the main avoidable cost.
- Digest bulk input before it lands.
- Prefer a script that returns a number over a read that returns a file.
- Finish a workstream and close it. One job per session beats one session that
  carries five jobs' context forward.
- A worker's context is separate and dies with it - that is a feature. Give it
  the bulk work and take back the conclusion.

## Integrate, verify, and repair

Run mechanical checks first. A narrow, explicit, reversible task with decisive
evidence may self-check. Require fresh independent review for ambiguity,
integration, design judgment, production / security / privacy / billing risk,
visual or editorial judgment, material plan deviation, or a conceptually wrong
result that could still pass its tests.

When a risk gate calls for independent review and subagents are not authorised,
perform an inline adversarial re-check against the actual artifact and say
plainly that the review was inline, not independent. The user is entitled to know
which one they got.

Integrate accepted work only. Check interfaces, overlaps, naming, duplicated
logic, missing integration, criteria, and preserved side effects. A repair gets
only the failed evidence and the affected scope. Make one evidence-driven repair
by default, rerun affected checks, then stop or escalate.

## Escalate only with evidence

Escalate when criteria cannot be verified, repeated repair fails, scope expands,
workers disagree, sensitive state is involved, integration conflicts remain, or
available capability, context, or tools cannot establish correctness.

Close with the terminal state, delivered artifact, evidence, remaining risks, and
required user action.

## Lessons are scoped, not automatic

The default result of a failed check is a run note, not a permanent rule.
Promote a note to a standing rule only when the root cause is understood, the
prevention is specific and reusable, scope and owner are clear, no active rule
already covers it, evidence is verified, and the defect repeated or was
materially costly.

Immediate behavioural correction is separate from persistence: fix it now,
decide later whether it becomes policy.

## Safety rails

- Workers never hold credentials and never push, deploy, publish, or delete.
- Before the first change, pin a restore point and state the one-step rollback.
- If a worker's report conflicts with what you observe, trust the observation and
  say so in the run log.
Upgrade 03

Prepare for AI search.

People are starting to get answers before they reach a website. AEO helps ChatGPT, Gemini, Perplexity, and Google’s AI results understand what a company does, trust the information, and cite the right page.

51.5%

of representative real-user queries triggered a Google AI Overview in a 2026 benchmark study.

Grossman et al., 2026
What this means

Search is becoming answer-first.

AI answers now appear across Google and standalone tools. They can reduce ordinary website traffic, but the visitors they do send may arrive with stronger intent. The goal is no longer only to rank. The goal is to become a source that AI systems can understand, trust, cite, and recommend.

Be Known

beknownonline.com

Top 15% of sites scanned · Source: Framer AEO

Category scores: Findable 25/25 · Quotable 25/25 · Understandable 25/25 · Trustworthy 23/25
Main improvement: Add more credible external citations. The detailed report gives external citation links 2/7.
85/100
Open the live report →
Marketer Vault

marketerprompts.com

Top 5% of sites scanned · Source: Framer AEO

Category scores: Findable 25/25 · Quotable 25/25 · Understandable 25/25 · Trustworthy 10/25
Main improvement: Add at least three contextual internal links in the page body. The detailed report gives internal links 0/18.
98/100
Open the live report →
Recommended next step

Test one workflow this week.

1

Choose one repeated task

Use a real workflow such as campaign research, content production, or website updates.

2

Add clear checks

Define what “good” means before the agents start.

3

Measure the gain

Track time saved, corrections avoided, defects caught, and final output quality.

Be Known AI Development Brief · Intelligence Layer v4.0 · updated for Opus 5, July 2026