A business can give a human employee a purchasing target without giving that person unlimited access to the business bank account. That is because the skills required to pursue that goal and the authority to act are separate decisions.
I believe it's even more important to make this type of separation clear for Agents. A request such as “resolve this customer's problem” describes a useful outcome. However, the agent has a lot of options e.g. the agent may issue a refund, change an account, contact another person or make a commitment on the company's behalf. It all depends on 'what' they can do!
In my experience, enterprises often approach AI as a way to make existing people and workflows faster. Thinking seriously about agents requires another conversation: which decisions and actions are we actually delegating?
A goal describes the outcome. Authority defines the actions available in pursuing it.
That distinction should be visible in the product design. We need to define the systems around the agent and clearly outline the way leaders evaluate success.
Agent Capability is Expanding into Ordinary Work
On September 28, 2026, Anthropic introduced Sonnet 5.5, describing stronger everyday work capabilities and additional safeguards. These are vendor claims that require testing in a particular application, but they reinforce a practical direction: increasingly capable models are becoming available for routine workflows. Anthropic's announcement provides the release details and limitations.
As capabilities increase, a task (or instruction) can become operationally ambiguous in new ways. A system that previously drafted a recommendation may now be able to carry it out. A connection that was useful for reading information may also expose actions that change the same information.
Organizations therefore need to revisit what delegation means. A better model can expand the range of actions the system can take, even when the business objective remains the same.

A conceptual distinction: what the business wants, what the agent can do and what the system permits it to do.
Define the Authority in Terms of Business Actions
Consider a hypothetical agent helping with customer returns. It can inspect an order, check a policy and prepare a response. Those capabilities do not automatically authorize it to refund a payment or promise that a replacement will arrive tomorrow.
The relevant boundaries should use language the business understands. The agent might be allowed to read an order and calculate an eligible refund. Issuing the refund could require an additional check. An exception to policy could require a named person's decision. Changing the policy itself would be outside the task.
This is more useful than a broad instruction to “be careful.” It creates a set of decisions that product and technology teams can implement and test together.
An authority record for the workflow should answer five questions: what information can the agent access, which actions can it take, what evidence must exist before acting, when must it stop, and who owns the result?
The details will vary by workflow, but the important feature is that the answer can be checked against an actual action rather than inferred from a general expression of intent.
Put Enforcement Around the Agent
Instructions help explain a task. The surrounding system also needs to enforce the limits of that task.
NVIDIA's September 28 announcement of its Open Agent Safety Platform describes runtime and separate monitoring controls around agents. It is an architectural signal rather than proof that a platform eliminates failures. The original announcement sets out the vendor's approach.
The broader product principle is straightforward. A component pursuing an objective should not be the sole mechanism deciding whether its own next action is permitted.
In the returns example, a request to issue a refund should pass through the application's actual policy and permission checks. Access should be limited to the information needed for that job. An unrelated customer record should not become available simply because a search tool can technically retrieve it.
This requires several kinds of judgment. A platform team can enforce an access rule, but a business owner must help determine what the rule should mean. Product teams need to decide what the customer experiences when a request reaches a boundary. A silent failure and a useful escalation are very different outcomes.

A proposed operating pattern. Technical controls and clear business policy must work together; neither establishes perfect safety.
Completion Should Need Evidence from the System that Changed
There is a second boundary that deserves attention: the difference between attempting an action and completing it.
An agent can produce a convincing account of what it intended to do. A business process needs evidence of what happened. If a refund failed, the customer should not receive a message saying it was issued. If a record update succeeded but its confirmation was lost, retrying blindly could create another problem.
A useful workflow therefore separates the proposed action, the attempt and the confirmed result. It checks the target system's state or an appropriate independent record before announcing completion.
This is familiar engineering discipline, but it becomes easier to overlook when a fluent assistant presents the interaction as one continuous conversation. The interface should preserve uncertainty when the underlying operation is uncertain.
For a leader reviewing an agent demonstration, a useful question is: what evidence allows this system to say the work is done? The answer should point to more than the agent's own explanation.
Recovery Should be Part of the Original Design
An agent workflow should have a response to mistakes before it gains a larger scope of action.
Some changes can be reversed, while others need a compensating action. E.g. a message sent to a customer cannot be made unread, and a promise can create expectations even when no database fields were changed.
The returns agent might need to stop further actions, identify affected records and bring the issue to an owner who can resolve it. That person needs enough context to understand what the agent attempted, which checks passed and what the target systems actually recorded.
Recovery also has a cost. A workflow that looks inexpensive during successful runs can become expensive if every exception requires a lengthy investigation. That burden belongs in the business case.
Testing should include partial completion, missing information, rejected actions and a failed escalation. The team should also test whether a person can actually stop the workflow and whether it stays stopped. A recovery procedure that nobody has exercised remains an assumption.

Recovery may involve reversal or compensation. The appropriate response depends on what changed and whether its consequences can be undone.
Expand Authority Only After Collecting Evidence
Start by giving agents a narrow set of actions with well-understood consequences. An agent might investigate an issue and recommend a resolution while a person authorizes the external action. As reliability improves, routine decisions can be delegated while unusual or high-risk cases are escalated.
The goal is not to keep humans in every loop, which would preserve the very bottlenecks automation is meant to remove. Instead, move the boundary deliberately based on evidence from real operation: whether the agent consistently achieves the intended outcome, stays within its authority, escalates the right cases, creates errors that can be detected and recovered from, and reduces rather than increases the amount of human supervision required.
This also makes organizational change tangible. Teams can see what agents handle, what still requires human judgment, and when people can intervene.
Humans do not need to remain in every loop. They need to remain where judgment, accountability, and the consequences of failure justify it.
The Business Should Own the Delegation
Agent authority is a joint product, engineering and leadership decision. The product defines a useful outcome. The business determines acceptable commitments and exceptions. Engineering makes the boundaries enforceable and the results observable.
The agent may become better at carrying out the work. The organization still owns the choice to give it access and authority.
Before expanding an agent's role, ask for a clear account of the actions it may take, the evidence required to confirm completion and the response when something goes wrong. Those decisions make autonomy a capability the business can operate, rather than a promise contained in a demonstration.