When designing an app, the useful questions are specific. Does the interface follow the design system? Does the data model support what the admin area needs to do? Are translations consistent? Has a feature been completed on the web while its mobile counterpart is still missing?
Those questions suggest a more useful way to organize agents than assigning one to imitate a product manager, another a developer, and another a QA department. An agent can investigate a particular problem, produce something another agent can use, or verify a part of the application from a different direction.
I am building an app for creating communities that interact and benefit each other. Thinking about this work makes the coordination question concrete: how should the work be divided so that the parts fit together?
In my experience, enterprises often try to fit AI into existing processes and focus on making current workflows faster. Copying those processes into a collection of agents can preserve the same unnecessary handoffs. Designing the work around the application gives us a better starting point.
An agent role should earn its place through the work it improves.
Start with an app and a clear assignment
Imagine an app with web and mobile clients, an admin area, and support for multiple languages.
The assignment might be to add a way for administrators to archive a content item and ensure that ordinary users no longer see it in the active list. That sounds like a small feature. It raises questions about data states, permissions, confirmation messages, shared components, translations and behaviour on each platform.
A useful first pass could divide the review like this:
Agent task What it examines A useful output Review the design system Existing components, spacing, typography, interaction states and accessibility expectations Components to reuse, inconsistencies to resolve, and missing states Review the data model Item states, relationships, validation and permissions Proposed data changes and the rules every client must follow Review translations Message keys, missing languages, repeated strings and usage context A list of missing messages and possible duplicates, with their locationsThese agents can inspect the same agreed version of the app in parallel. Each has a distinct question and a concrete output. None needs to begin by rewriting the shared components, schema or translation files.
That distinction matters. Parallel investigation is easier to coordinate than parallel changes to the same shared foundation.
Agree the interfaces before building the pieces
If you tried to follow your org chart and created a PM agent, a developer agent and a QA agent, then things may not turn out as well.
Suppose the PM-agent proposes an archived state, while the Dev-agent assumes archiving means deleting the record. Both agents could produce internally consistent work that becomes incompatible when combined.
Before implementation, resolve what archiving means, which users can perform it, what the API returns, and how the change reaches the clients. Confirm which existing component provides the action and what its confirmation and error states should say.
Those decisions become shared inputs. The admin-area agent can design the management screen around the agreed states and permissions. An agent optimizing the UI can improve navigation, loading, empty and error states using the existing component system. Translation work can proceed against agreed message meanings, with a final check against the rendered screens.
Some of that implementation may happen in parallel once the dependencies are clear. Shared files still need ownership: decide who integrates a schema change, a component change or an update to the translation catalogue. Otherwise, several agents can each make sensible edits that conflict or overwrite one another.

Review the same baseline independently, resolve shared decisions, then build and check the integrated application. The diagram describes a proposed workflow, not a measured result.
The plan can change as the agents learn. A translation audit might reveal that a supposedly missing message already exists under a different key. A UI review might find that a requested new component duplicates one already in the design system. The useful response is to revise the work, not complete every original assignment regardless of what has been discovered.
Give review agents something specific to find
I find that “Review the app” is a large, ambiguous instruction that may or may not give me the right outcome. I try to provide more focused assignments that make it easier to judge whether a reviewer contributed something useful.
Find redundant code. Ask an agent to locate repeated validation, similar data-fetching logic, or components implementing the same interaction differently. The output should identify the files, explain the overlap and propose what could be shared. Similar-looking code is not automatically interchangeable; platform requirements or different business rules may justify the difference. Confirm that before consolidating it.
Find duplicate translations. Ask another agent to inspect keys, values and references. Two keys may express the same message, while the same English word may need different translations in different contexts. A useful finding explains the meaning and where each key is used. Blindly merging identical strings can erase an important distinction. After a change, check both references and the interface: a valid translation can still overflow a mobile button.
Find incomplete platform coverage. Ask a reviewer to trace the feature through web and mobile. Has the action been wired up in both places where it is required? Do the clients interpret the same state correctly? Are loading, failure and permission-denied cases handled? Can an archived item still appear in one client because its list has not refreshed?
For the illustrative archive feature, the admin action might intentionally be available only on the web. The required feature parity is that both user-facing clients respect the archived state (and don't show it). Equivalent product behaviour does not mean identical screens. We don't need to build every administrative action on every device.
A simple coverage table can distinguish completed, missing, not tested and intentionally unsupported behaviour, with links to the relevant code or test evidence. “Not tested” should remain visible. An agent that could not run the mobile app cannot establish that the mobile experience works by inspecting the web implementation.
Integration is where the app becomes the product
Individual assignments can all be marked complete while the user’s journey remains broken. The admin screen may update the correct field, the API may accept it, and the translation file may contain the right message, yet the mobile app may still display the item as active.
Review the connected behaviour. In the example, an administrator archives an item; the permitted state change is stored; the appropriate lists reflect it; and users receive understandable feedback in each supported language. Verify denied actions as well as successful ones. A hidden button alone does not demonstrate that an unauthorized request is rejected.
The evidence can include relevant automated checks, screenshots of rendered states and a record of the actual flows exercised. Each finding should say what was tested and what remains uncertain. An agent reviewing another agent’s summary is less useful than one checking the changed code and observable behaviour.
Agreement between agents does not establish correctness if they share the same mistaken assumption. Give the reviewer a specific requirement to challenge, such as whether the mobile client respects the new state or whether a duplicate translation is genuinely equivalent.
Someone also needs to accept the integrated result. That could be me reviewing whether the app does what I intended, with agents helping assemble the evidence. The acceptance decision remains explicit as the execution plan changes.

Compare arrangements using the working feature and the effort needed to verify it. More completed agent assignments do not automatically mean a better application.
Test whether the extra coordination is worthwhile
For a small UI correction, one capable agent may be enough. For a change spanning the schema, admin area, translations and two clients, separate reviews may expose problems that a broad assignment misses. These are choices to test, not a reason to use a large agent team for every change.
Give one appropriately equipped agent a fair baseline, then compare an alternative under the same task, spending ceiling, deadline and acceptance criteria. Count the total cost, time to an accepted result, defects found and effort spent resolving conflicting changes. A team that finishes its individual tasks quickly but leaves extensive integration work may not be helping.
This is consistent with the architectural questions in Anthropic’s December 2024 guidance on simple workflows, parallel work and dynamically assigned subtasks. Its tooling examples are historical; the useful principle here is to justify additional complexity through the outcome. Source: Anthropic, December 19, 2024.
Ethan Mollick’s October 1 essay considers agents taking on more of the coordination themselves, while leaving open how well this extends to sustained organizational work. That is a useful direction to explore, with concrete assignments like these providing a bounded place to test it. Source: Ethan Mollick, October 1, 2026.
Keep the outcome explicit as the plan changes
For this app, keep four commitments visible: the feature’s intended behaviour, the constraints on changes, the evidence required for acceptance, and the owner of the decision. Within those limits, the breakdown of work and the agent assignments can adapt.

A proposed way to delegate execution while retaining a clear agreement about what the app must do.
Domain knowledge, experience, taste and judgment still matter here: deciding whether the admin flow makes sense, whether the interface is coherent, and whether a difference between web and mobile is intentional. Currently, these are best done by humans. I expect more of these tasks to become delegable over time. The useful approach is to keep examining where that delegation works.
The business result is an application whose pieces fit together. Reusable components can reduce future inconsistency. Clear data rules help clients behave predictably. Deliberate translation and platform checks can prevent a feature from being declared complete for only part of its audience.
Start the next assignment with the application’s unfinished work: review the design system, check the data model, design the admin area, improve the UI, audit translations and verify platform coverage. Give each agent a question it can answer and evidence it must return. That is a more practical basis for an agent team than a set of familiar job titles.