When AI Can Do More, What Should Count as Progress?
I took my first Waymo ride in San Francisco, where I was staying to attend OpenAI DevDay in person on September 29, 2026. Since returning home, I have kept connecting that experience of traveling with nobody in the driver's seat to the direction for AI presented at DevDay, as I reflect on my work and the operation of my firm.
At DevDay on September 29, OpenAI described agents taking on ongoing responsibilities. Its official recap makes that direction explicit. For me, the question shifted from what AI could now do to what I wanted that capability to change.
Two parts of my practice came together: supporting decisions in banks and manufacturing companies, and trying to make the operation of Fragment Practice, my one-person firm, more delegable to AI. They differ in scale and subject, but lead back to the same questions. What counts as an outcome? Who makes the judgment? Under what conditions may the work proceed? And what can people stop carrying themselves?
More material does not necessarily move a meeting forward
Ask AI to research and prepare for the next meeting, and it can produce an issue list, comparison tables and slides. Preparation takes less time.
Yet if the meeting's decision remains unclear, the result may simply be more material.
Consider a meeting about the scope of AI use. Is the decision to start a trial, define the information it may handle, or adopt it more widely? Who has authority to decide, and which uncertainties can remain unresolved? The answers change what evidence the meeting needs. Unless somebody can carry out the decision afterwards, even agreement can turn into another meeting.
This is an illustrative situation, not an account of a particular client engagement. In decision-support work, I try to turn a request to prepare a document into the conditions that would make a judgment possible.
The same change helps with AI delegation. Instead of starting with “give me more options,” try “make this decision possible.” That changes when research should end. If an uncertainty could change the decision, investigate it. If another document will make no difference, there may be no reason to produce it.
Counting decisions is not enough either. Establishing why a decision should be deferred, or deciding that work should not be done, can be valuable outcomes. Before counting completed tasks, I want to know what the work was supposed to change.
A large organization and a one-person firm share a design problem
When delegating to a person, we discuss the expected result, the information they may use, what they may decide and when they should consult someone. Capability does not resolve conflicting assumptions about the assignment.
These questions remain when delegating to AI. But shared atmosphere and tacit agreement cannot carry the whole arrangement. Decision criteria need to be written down, permitted actions constrained through access controls, and states such as awaiting approval or stopped made legible to the agent.
In operating Fragment Practice, I am trying to externalize work and its decision conditions in a form AI can use. By externalize, I mean making the purpose and current state available to the next participant, so they can proceed within an agreed scope. I do not mean documenting every thought in my head.
For example, an agent can prepare an article, a separate participant can review it, and I can decide whether it should be published. Execution then needs to establish that the version being applied is the version approved. A clean review does not itself authorize publication.
In a bank or manufacturer, those assumptions need to be shared across departments and roles. In a one-person firm, assumptions concentrated in my head need to become accessible to people and AI. The settings differ, but the design problem is continuous: what must be made explicit for delegation, and which judgments stay with whom?
Delegation can expose ambiguity in the original work. It can also accelerate work toward an ill-defined goal. Greater execution capability amplifies the quality of the design behind it.
What the Waymo ride made tangible
What struck me was a function associated with somebody's physical presence and skill being delivered through a vehicle and an operating system around it.
In my work in Japan, I have encountered situations where performance depends on a particular person knowing the history, coordinating the participants and connecting the final steps. That includes expertise and attentiveness that are not easily formalized. My first Waymo ride made me ask how much of the work sustained by that person's capability could be carried outside the individual.
To me, it looked like one symbol of an American pursuit of productivity through systems. That is an impression from a single ride, not evidence for dividing Japanese and American work into two national types or ranking them.
An empty driver's seat also does not mean an operation without people. Waymo's official account of fleet response says people may provide additional context while the Waymo Driver retains control. A vehicle may stop, resume autonomously, or require manual retrieval by roadside assistance. I am describing the published operating approach, not something that happened during my ride.
Autonomous driving and software agents have different technologies and risks. Still, both raise a useful design question: what conditions allow delegation, and what should happen outside those conditions?
For an agent doing business work, that could mean unavailable evidence, conflicting records or a changed target after approval. Should it fill the gap with an assumption and continue, or stop and return the decision? Designing autonomy includes deciding how it handles that boundary.
A continuing role is different from a personality
OpenAI describes dots, introduced at DevDay, as agents with their own cloud environments that pursue ongoing goals between conversations. Rollout depends on plan and market; specialist dots with dedicated organizational roles are in pilots.
Giving a continuing agent a name can make it feel like an extension of one's capabilities, acting alongside the person. I see a shift from a tool answering isolated requests toward something sustaining a role over time.
But familiarity does not establish the conditions for delegation. A continuing assignment needs continuing clarity about available information, permitted actions, evidence of completion and when judgment must return to a person. Remembering a conversation cannot provide all of that.
OpenAI's own safety description includes access limits, action checks, read-only proactive research and mechanisms to stop work, while acknowledging that mistakes remain. Whether a product has safeguards and whether my work can be delegated under those safeguards require separate assessment.
What I hope to gain from a persistent role is less repeated explanation. I do not want to evaluate that prospect through the appearance of personality alone. The burden of reconnecting work is a question I explored in Faster Tasks Do Not Mean the Work Moves On.
An AI that knows my interests also shapes my attention
When AI works continuously with the context of the person assigning the work, as well as the state of the work itself, delegation raises a wider set of questions.
My expectation of personal AI goes beyond remembering previous conversations. It might come closer to the questions I keep returning to, or a discomfort I have not yet found words for, and offer material that helps me notice something at a useful moment.
While I am thinking about delegation, for instance, an example of operations or failure in another field might change how I see my own problem. That is a possibility I find valuable, not a claim that today's AI accurately understands my cognition.
Choosing what to surface also affects what receives my attention. An agent that repeatedly follows past interests might reinforce them when I hoped it would widen my perspective. Inferring a preference does not establish permission to keep prioritizing it.
Google's description of Personal Intelligence acknowledges excessive personalization and errors in interpreting context. More memory and more connected information do not complete an understanding of the person.
I want to choose what information the AI can use, but also to inspect why it brought something forward and which interest it inferred. I need room to correct that inference, decline personalization, stop assistance and consider other options. Repeatedly relying on an agent's framing could also weaken the habit of forming my own questions.
So I want to ask not only what personal AI makes easier for me, but what it might make harder to choose.
Externalization creates maintenance work too
At Fragment Practice, I want to reduce the need to hold every task's progress and context in my head, or repeatedly reconnect agents myself. When existing work resumes, I want its purpose and unresolved points to be recoverable from current records, with only the necessary judgments brought back to me.
This remains an experiment under development. Fully autonomous operation has not been established. Recording a boundary, observing an execution once and knowing that work can be reliably delegated over time are different achievements.
Recovery from failure and interruption belongs to the work too. If it is unclear what finished and what remains unverified, resuming may duplicate an action. An artifact can exist without having been accepted or applied. Until the resulting state is checked, some of the burden stays with me.
There are limits to this approach. Describing decision conditions in detail may take more work than it saves. Trying to define every exception in advance may constrain the ability to notice a new question. I do not assume that the value of human experience can be replaced by the parts that fit into a record.
I would start with recurring judgments and conditions whose misunderstanding would have significant consequences. I would look beyond the number of approval requests to whether the evidence arrives ready for a decision, and whether rework and explanation have also decreased. If maintenance costs more than it removes, the delegated scope should become smaller.
Before adding the next assignment
AI makes research and creative work possible that previously had to be left aside. I welcome that possibility. But taking on everything that becomes feasible can also multiply the work people must review, decide on and answer for.
Beyond productivity, I want a way to keep doing worthwhile work without continually taking on more tasks. Making judgment and responsibility explicit should leave people able to choose what they take on and delegate what they have chosen, with the system serving that choice.
For the next assignment I give AI, I want to start with one piece of work. What change would count as an outcome? Who judges it, what can be delegated, and what should stop execution? Once it is finished, is there less that I must continue carrying?
Decision support in banks and manufacturers, and AI operation in a one-person firm, are practices approaching those questions from different settings. DevDay and the Waymo ride helped me see their connection again.