Lesson 6 of 6 · Agents
Where this is going
Three things are moving fast and measurably. A lot of other things are being shouted. Here is how to tell them apart, and how to stay current on a few minutes a month.
This is the only lesson in the platform that is mostly about the future, so a rule up front: every forward-looking sentence here is either tied to a source you can check or is flagged as a guess. When you read about agents anywhere else, hold the writing to the same standard. The list of sources is at the end. The habit of asking for them is the point.
Three things that are actually changing
Longer tasks. The best-measured trend in agents comes from METR, an independent research group that tests how long a task an agent can finish on its own. Their method: give agents software and research tasks that take skilled humans anywhere from seconds to many hours, and find the task length at which the agent succeeds half the time. In their January 2026 update the leading model handled tasks of roughly five hours of human work at that 50 percent mark, and the task length had been doubling roughly every four to seven months depending on the period measured. METR's own caveat, which matters for you: the tasks are software and research tasks. Nobody has shown the same curve for running a front counter. But the direction is real, it is measured, and it has held for years.
Computer use. Agents that click, type, and read a screen the way a person does, rather than through purpose-built tools. As of August 2026 this is shipping, not promised: Anthropic offers Claude in Chrome, an extension that reads, clicks, and navigates websites alongside you, on its paid plans; OpenAI launched ChatGPT Work in July 2026 as an agent that takes an outcome, gathers information across connected apps, and works for hours. Both makers describe the same posture you learned in lesson 5: it pauses before sensitive actions, and it still needs a person nearby. What is changing is how much of the web an agent can operate without anyone building a connector first.
Agents calling agents. One agent handing a sub-task to another and combining the results. Claude Code's docs already describe a lead agent coordinating several sub-agents on one job. Across companies, Google published the Agent2Agent protocol in April 2025, a standard for agents from different makers to find each other and delegate work, and it is now governed by the Linux Foundation with backing from Google, Microsoft, AWS, Salesforce, SAP, and ServiceNow among others. The plumbing exists. Evidence of ordinary businesses running it day to day does not yet, which puts it in the "watch" column, not the "buy" column.
What is hype
Hype is easy to spot once you know the tells. It uses the word "replace." It has no task length attached ("agents can do anything" versus "agents can finish five-hour software tasks half the time"). It skips the failure rate. And it never mentions permissions, because permissions are where the "go fishing" story falls apart.
The specific claims to be skeptical of this year: that an agent can run a role, not a task (no one measuring this has shown it); that you can set it and forget it (every maker's own documentation says the opposite); and that the doubling curve applies to your job (METR themselves say their tasks are software and research). Sort the six claims below and see how your instincts hold up.
Sort it: real now, too early, or hype?
Claim 1 of 6
An agent can work on a well-defined software task for hours without a person stepping in.
Verdicts reflect what could be verified from primary sources on August 24, 2026. Sources are listed at the end of the lesson. This will age; the habit of asking for evidence will not.
What will probably stay true
Predictions, flagged as such. These are guesses grounded in the structure of the thing rather than in any announcement.
- Permissions will still be the safety story. Models get more capable; the fact that they only write requests does not change. Whatever the products look like in two years, the question "what can it do without asking?" will still be the one that matters.
- Accountability will not move. No vendor is going to sign for your quotes. The email still goes out under your name.
- The ladder will still be the way in. Read, draft, approve, routine. Longer tasks and better computer use raise the ceiling of what a rung can do; they do not skip rungs.
- The gap between demo and daily use will stay wide. A 50 percent success rate is a headline in a benchmark and a problem at a front desk. Expect the impressive numbers to arrive in your business a year or two after they arrive in the news.
Staying current without living on the news
You do not need a feed. You need three questions and a calendar reminder.
- Once a quarter, check the measurement. METR's time-horizon page is the one number worth tracking. If the doubling holds, the rung-four jobs get bigger. If it stalls, you lose nothing by having waited.
- Once a quarter, reread the maker's permission page. Claude Code and Codex both document exactly what runs without asking in each mode. Those pages change, and a change there matters more to your business than any launch announcement.
- When something is announced, ask the four questions. What tools does it have? What runs without asking? What stops it? What is the measured success rate on a task like mine? A product that cannot answer those is a demo, not a tool.
That is fifteen minutes a quarter. Everything else, the model races and the launch videos and the threads predicting the end of work, you can skip without losing a thing that will matter to a business in Washington County.
Try this yourself
Turn the four questions into a reflex. Next time you see an agent announcement, paste the link or the text into your AI app with this.
Here is an announcement about an AI agent product: [paste]. Answer only from what the text actually says, and say "not stated" where it does not. 1. What tools does the agent have, and which of them write or pay? 2. What can it do without asking a person? 3. What stops it: a step limit, a spend limit, a list of actions that always ask? 4. What measured success rate does it claim, on what kind of task, from whom? 5. Which of my rollout ladder rungs would this fit on, and what would I need to see before moving it up one? Then give me one sentence: is this a tool I could use at rung one today, or a demo to watch?
Sources
Everything factual in this lesson was checked against these on August 24, 2026. They are the pages to go back to.
- METR, Time Horizon 1.1 (January 2026) and the time-horizons page: metr.org/time-horizons
- Anthropic, Claude in Chrome help pages: support.claude.com, and the Claude Code permission modes page: code.claude.com/docs
- OpenAI, ChatGPT and Codex documentation and changelog: learn.chatgpt.com/docs
- Model Context Protocol history and the Agentic AI Foundation: the Wikipedia entry for Model Context Protocol, and anthropic.com/news/model-context-protocol
- Agent2Agent protocol: the Linux Foundation press release of June 23, 2025, and a2a-protocol.org
That is the Agents module. If you started here, the earlier modules, beginning with AI Basics, fill in the model side of the story. And if you want a hand deciding which job in your own business belongs on rung one, that is the kind of conversation B-Squared has with local businesses every week.
Last updated August 24, 2026