Abstract
Codex made it possible for one person to create many specialized AI collaborators and attempt work with a breadth, depth and pace that no individual could reasonably achieve alone. That leverage had value only if the AI team could carry the expanding research, technical and operating burden while returning the human’s attention to the purpose, people, values and consequential decisions that AI could not supply.
As the work expanded, however, the advantage of specialization produced its own constraint. Knowledge became distributed across conversations. Individual tasks accumulated context, compressed, drifted or needed replacement. Decisions and limitations traveled without their full reasoning. The human became the only dependable connection among specialists and repeatedly had to reconstruct and operate the team.
This paper describes a four-day field experiment in building a user-controlled operating layer around Codex desktop tasks. The model combines durable organizational roles, file-first institutional knowledge, a database-backed current-state registry, provenance-aware STORE PLAN and RECALL, role-based training, applied qualification, targeted recovery, independent evaluation and versioned operating procedures. It was developed and tested inside CivicSS™, a long-running civic-information platform whose work spans many technical and human domains.
The result is not a mature autonomous organization. It is an advanced proof of concept with bounded live evidence. A fresh task recovered two nuanced bodies of work in 3 minutes 14 seconds without reading the originating chats or broadly searching the underlying library. Tasks were trained, failed, retrained and retested. A coordinating role began managing dependencies across specialists while preserving distinct ownership and evaluation authority. The experiment suggests a practical third option between a long-running conversation with accumulated bias and a fresh conversation with no institutional memory: a fresh mind with durable, sourced institutional memory. Its primary success test is human: whether the system returns Sherie’s time to the CivicSS work that requires her presence and judgment.
The purpose: expand one person’s reach without replacing the human
The operating model was not created because Sherie wanted to spend her time administering an AI organization. It was created because CivicSS had become too ambitious for one human or one AI conversation to research, build, operate and communicate alone.
Only months earlier, Sherie had never used an AI platform and had little experience with the ordinary pathways through which residents participate in Town affairs. One local decision made her want to understand more. The questions quickly expanded: How does the Town work? What do residents’ taxes support? What has already been promised? What major investments are coming? Who knows which part of the answer? How do residents see the combined effect of several reasonable choices? How can a community recognize the moment when analysis has done all it can and a genuine values decision begins?
What began as a public-information website grew into databases, evidence libraries, resident reports, scenario models and a decision framework. That growth was not technology for its own sake. Every new layer answered a problem encountered while trying to make the Town’s whole picture understandable. A spreadsheet could not hold the data. A database could hold it but could not tell the human story. A report could explain one subject but not the relationships among many decisions. Scenario analysis could show what might be possible but could not determine which values a community should choose.
Codex made it possible for Sherie to keep following those questions rather than stopping at the boundary of one person’s time, technical skills or domain knowledge.
The specialist team made that ambition possible. Different tasks could learn municipal finance, property, affordability, capital assets, school decisions, land use, public records, database architecture, reporting, scenario modeling, testing and writing. They could work in parallel and combine capabilities far beyond Sherie’s individual time or technical range.
But the team created value only if it also protected the work Sherie could not delegate to AI. She needed time to:
- define why CivicSS exists and the ethical code it must follow;
- decide which human and civic questions matter;
- meet, observe and listen to residents and Town participants;
- understand how real people experience trust, complexity, participation and public decisions;
- test what helps residents understand their role and feel prepared to attend and vote;
- interpret lived, historical and relational context that documents cannot establish safely;
- identify the point at which evidence ends and competing human values begin;
- explore whether the Town sees value in aligning its own information, responsibilities and decisions more clearly; and
- make the consequential choices about risk, representation, publication and direction.
CivicSS does not decide for residents. It helps them understand how a decision develops, what is known, what is assumed, whether the information is ready for the next decision, what choices remain, what each may provide or require, who has authority and where the facts can no longer determine the answer. Residents remain free to weigh the values and choose for themselves.
Codex can collect the sources, trace the process, perform the calculations, model the alternatives, expose the gaps and help explain the result. It cannot attend a Town meeting as Sherie, build trust with a neighbor, know what a pause or expression meant, decide which community value should prevail or determine what another person ought to believe.
The operating model is therefore infrastructure for human leverage. It should take routine task administration, recovery, coordination and evidence management away from Sherie so one person can lead the combined capacity of many specialists while remaining present for the work only a human can do.
The model succeeds only if it helps one person accomplish what she could not accomplish alone without displacing the human contribution that makes the work meaningful.
The human process the capacity is meant to support
The hoped-for CivicSS process begins and ends outside the AI system.
- Sherie encounters a real civic question through residents, public meetings, Town participants or lived experience.
- She determines why it matters, whose understanding is missing and which ethical or human questions the platform must protect.
- Specialized Codex tasks assemble the public record, financial and operational context, assets, authorities, options, dependencies, uncertainties and consequences across domains.
- Sherie adds the lived, relational and historical meaning that cannot be inferred reliably from the sources.
- CivicSS presents a connected decision story: what is known, what is not, what each path may provide or require, who is responsible and where facts stop determining the answer.
- Sherie tests that explanation with residents and appropriate Town participants—not to secure agreement, but to learn whether people can understand the choice and participate on their own terms.
- What she learns returns to the platform, improves the analysis and reporting, and may reveal ways the Town’s own communication or operating process could become more connected.
- After a decision, the system follows implementation and results so the community can learn rather than beginning the next choice without institutional memory.
This is a hopeful but demanding human process: understand together, decide together and learn together without using information to manipulate the answer. It requires domain expertise, public-record discipline, data engineering, analytics, scenario modeling, design, writing, user research, civic relationships, ethical judgment and organizational learning. Ordinarily, that is a large multidisciplinary team.
Using fewer tasks would make the operating model simpler, but it would not make the civic problem simpler. The specializations existed because the mission required them. The operating layer was built because the resulting team then needed continuity, coordination, recovery and a way to keep its human leader at the purpose-and-judgment layer.
The hopeful result is not that every resident reads every report or reaches the same conclusion. It is that someone can begin with a personally meaningful question, understand how it connects to the larger Town picture, recognize what is known and still uncertain, see where and when participation matters and make a more informed choice. A resident may still support the school, preserve the building, oppose the expense or rank affordability above another value. CivicSS succeeds when the person can see the choice more clearly—not when the person chooses the answer Sherie, Codex or anyone else prefers.
Codex helped reveal the human role itself
This was not only automation of a role Sherie had already defined. The collaboration helped her discover a new role at the intersection of:
- municipal finance, property, affordability, schools, assets, land use and public process;
- civic systems thinking and organizational alignment;
- resident experience, comprehension, trust and motivation;
- public relationships and real-world observation;
- ethical boundaries, role capacity and neutral decision support;
- translation between technical evidence and human meaning;
- competing-value and decision-readiness design; and
- learning from what happens before, during and after a public choice.
No single person could ordinarily sustain that breadth, depth and pace alone. It would more commonly require a multidisciplinary team of researchers, analysts, engineers, designers, writers, project leaders and subject-matter specialists. Codex supplied much of that supporting capacity. The continuing dialogue also helped Sherie see connections among those disciplines and conceive a higher-order human role: learning from the community, setting the ethical purpose, directing the system and carrying what the system learns back into real civic decisions.
The AI did not replace the human role. It helped make a much larger human role imaginable and possible.
1. The problem appeared only after Codex worked well
Sherie began using AI in April 2026 while becoming involved in local civic affairs. What started as a public-information website grew into CivicSS, a platform intended to help residents understand the Town’s financial position, property and tax mechanics, affordability, public assets, capital needs, schools, land-use choices and major decisions as one connected picture.
The work quickly exceeded what one spreadsheet, one application or one conversation could hold. Codex made specialization possible. Separate tasks developed expertise in database architecture, data loading, public records, property reporting, affordability, capital management, school scenarios, resident communication, quality assurance, writing and program coordination.
The arrangement produced real value. Specialists could go deeper. Several lines of work could advance at once. A reporting task could refine a resident explanation while a research task preserved public records and a data task improved the underlying structure.
Then the second-order problem emerged.
The more capable and specialized the tasks became, the more difficult it was to maintain one coherent project. Important knowledge lived inside separate chats, temporary working files, code worktrees, downloaded artifacts and the memory of whichever task had been present for the original decision. Tasks sometimes believed they owned the same thing. A product could have a newest package that was not the accepted current release. A finding could survive while the reason for rejecting an earlier interpretation disappeared. A new task could inherit a product name without inheriting the semantic rules that kept its data safe.
Long-running tasks also had a natural limit. As their conversations accumulated history, they needed to carry current instructions, superseded decisions, exceptions, open obligations, source distinctions, working relationships and day-to-day coordination in the same active context. Compression could preserve the broad story while losing the exact rule, current pointer or unresolved limitation. A task could remain fluent and helpful while becoming less reliable.
Starting fresh solved the accumulated-context problem but created a reconstruction problem. Sherie had to explain what had happened, which source controlled, why one version had been rejected, who owned the next step and what the new task must not assume. The cost increased with every specialist and every month of work.
The operating problem was no longer whether an individual Codex task could do excellent work. It was whether many excellent tasks could become one recoverable team without making the human their permanent memory, scheduler and message bus.
2. The third option
Persistent AI work often appears to require a choice:
- keep the experienced conversation and retain continuity along with its context load, assumptions and drift; or
- begin with a fresh conversation and lose the accumulated understanding.
The experiment tested a third option:
A fresh mind with durable, sourced institutional memory.
The goal was not to make one AI remember everything. It was to make the work recoverable, attributable, current enough to use and connected to deeper evidence.
That reframed memory as an organizational capability rather than a property of one conversation. A future task should be able to determine:
- which role it is filling;
- what outcome that role owns;
- which decisions and limitations are current;
- where the substantive evidence lives;
- what changed and why;
- what remains uncertain or blocked;
- who owns adjacent functions;
- what authority the task does and does not have; and
- what applied evidence shows that it is ready to work.
The individual task may be finite. The role, knowledge, obligations and learning history should be able to survive it.
3. The operating layer
The resulting design has seven connected parts.
Durable roles, replaceable tasks
An exact Codex task is a communication endpoint. A role is the durable organizational job. A display name may change. The task filling the role may eventually be replaced. Product ownership, management accountability and service routing are related but different relationships.
This distinction matters in daily work. A reporting specialist may own a calculation. A library role may preserve its evidence. A manager may be accountable for the specialist’s operating continuity. An independent evaluator may decide whether the work passed. A human may retain the consequential decision. Treating all of those actors as “the owner” creates unsafe handoffs.
At one structural snapshot, the registry contained 41 current tasks, 41 durable roles, 41 primary task-to-role assignments and 41 accountability routes, with no structurally unmanaged task. Those counts demonstrate organizational coverage, not universal competence.
File-first substantive knowledge
Important reasoning must survive outside chat. That does not mean saving every conversation. It means preserving the knowledge a future task would otherwise have to reconstruct: purpose, decisions, semantic distinctions, rejected interpretations, limitations, unresolved questions, current release routes and the smallest safe next step.
The durable record points to deeper owner files and source evidence. A compact plan reduces orientation cost without pretending to replace the underlying record.
Database-backed current state with separate history
Files are well suited to narrative knowledge, evidence and explanation. They are not an ideal shared current-state table for many writers. The model therefore separates substantive files from a queryable operating registry.
Normal operations read current identity, role, reporting, ownership, qualification, activation and status from current records. Changes are preserved separately in history and evidence. This follows a familiar database principle: current state should be easy to query; audit history should explain how it changed.
The database does not replace the evidence library. It tells a task which role or record is current and where the governed evidence can be found.
Selective institutional memory through STORE PLAN and RECALL
STORE PLAN captures a bounded body of important work while the originating context is still fresh. The record includes the subject, key decisions, distinctions, current state, ownership, limitations, next actions and links to deeper sources. It also includes current release pointers for code-producing roles.
RECALL begins with a small discovery surface rather than a broad search of every file. The retrieving task selects relevant segments, verifies identity and opens only the linked records needed for the question.
Retrieval does not silently transfer ownership, convert a plan into truth or make an old conclusion current. Provenance and staleness remain visible.
Training, qualification and professional development
Reading instructions is not proof that a task can apply them. The model separates:
- training;
- teachback;
- applied qualification;
- independent evaluation; and
- activation for a bounded operating scope.
Later development can add a new competency without discarding everything the task already demonstrated. A writing specialist, for example, can retain proven voice capability while receiving a separate school-domain module. Only the new layer remains provisional until it is applied and tested.
Failure becomes a learning event. The failed result is preserved. The smallest missing competency is identified. Training is focused on that gap. A fresh applied test determines whether the capability recovered.
Executable operating procedures
Some instructions are too important to depend on conversational memory. Routine procedures such as task registration, training administration or plan storage can be written as versioned programs that the responsible task rereads and verifies before acting.
The prototype also explored exact command contracts. The phrase “store plan” had once been interpreted as a topic for discussion when Sherie intended the full procedure to run. The response was a versioned distinction between review language and the exact execution form STORE PLAN — <description>, with a machine-readable registry, input schema and local classification tests.
Five classification cases passed. That proves only the local classifier. It does not prove native Codex invocation or completion of the full workflow.
Human authority as an explicit layer
The model does not aim to remove the human. It aims to stop using human attention for work the organization can perform reliably itself.
Sherie supplies purpose, values, priorities, meaning, material judgment and consequential authority. She also contributes diagnostic questions that cross specialist boundaries. AI roles translate that direction into architecture, research, implementation, tests and communication.
The intended movement is upward: from personally routing every message and reconstructing every task toward intervening where judgment changes the meaning or direction of the work.
4. What happened during the concentrated experiment
The operating model did not begin with an abstract schema. It emerged from failures in real work.
One trigger occurred when older releases were moved to verified archive storage. Nothing had disappeared, but a current process continued treating the artifacts’ former locations as part of their identity. Admission of a new plan became blocked by unrelated historical files that had moved. Sherie identified the essential distinction: this was a process failure, not data loss.
That led to deeper organizational questions. What is the immutable identity of a task? What is a durable role? Who manages a role? Who owns a product? How does a task find database, library, training or reporting help without asking the senior Partner? Which state belongs in files, and which belongs in a database? Who is permitted to update it?
Sherie and the senior Partner shaped the operating principles. Specialist tasks translated them into role structures, database candidates, lifecycle programs, tests and evidence. Independent reviewers rejected deficient versions. The failures remained preserved and informed successor designs.
Once the organizational backbone existed, integrated operating accountability moved to an Operating Model PMO. The PMO met its direct reports, discovered misunderstandings, routed focused retraining, preserved an unfinished obligation attached to a dormant predecessor and restarted several workstreams without absorbing the specialist authority of the roles doing them.
The most important transition was not that every component became complete. It was that Sherie and Partner could begin handing routine coordination to a recoverable operating function while retaining strategy, architecture and human authority at the appropriate levels.
5. What the evidence supports
The experiment produced three especially useful demonstrations.
First, a fresh-task retrieval test recovered two nuanced subject areas in 3 minutes 14 seconds. The task did not read the originating chats or broadly search the Knowledge Base. It began from a 14-row discovery view, selected five segments, verified relevant hashes and opened thirteen linked files. Both subject tests passed. Later review found that fast retrieval did not guarantee complete decision coverage, and that limitation remains part of the result.
Second, training and recovery became observable. One management exercise failed at 16/20, remained preserved, received focused retraining and passed a new 20/20 exercise. Multiple technical or operating candidates were rejected before use, corrected and retested. A lifecycle program progressed through rejected versions before an accepted version passed 48/48 bounded tests.
Third, coordinated team behavior appeared in live work. A senior role defined architecture and exceptions. PMO coordinated dependencies. Window Operations maintained lifecycle state. Library Maintenance preserved evidence. Schema Control designed technical changes. UAT retained independent evaluation authority. Domain specialists retained substantive meaning. Sherie corrected plan-level meaning without taking back the detailed work.
These cases do not establish unattended autonomy. They establish that specialization, recoverable knowledge, role boundaries, applied testing and accountable coordination can produce capability greater than the reliable active memory of one conversation.
6. What was most valuable to the human
The immediate benefit was not the number of files or database tables. It was protection of attention.
Before the operating model, a small mistake could pull Sherie away from civic analysis and into an extended investigation of what else a task might have forgotten. She had to determine whether the problem reflected one missed instruction or broader drift, find the right source, retrain the task and ensure the original work resumed.
With management, training and evidence routes in place, she could sometimes describe the problem once. The system could translate it into a bounded correction, targeted learning, applied review and preserved result. She could observe the work without becoming its trainer, custodian, evaluator and coordinator.
The larger human value is not simply reduced workload. It is a better division of work.
Codex can carry more of the expanding information and operating burden: finding and preserving sources, tracing versions, separating facts from assumptions, performing calculations, identifying missing evidence, building reports, maintaining current state, testing work and coordinating repeatable procedures.
Sherie’s attention can remain on the questions that require a person embedded in the real civic environment:
- Why does this work matter?
- What ethical principles must the platform protect?
- What do residents need to understand before they can participate meaningfully?
- Where does the evidence end and a personal or community value choice begin?
- What are officials, committees, boards and residents actually trying to accomplish?
- Which legal, procedural, relational or trust boundaries affect how the work can be used?
- Is there interest in improving how the Town itself connects information, roles and decisions?
- What may the platform publish, deploy or represent externally?
This division is visible in CivicSS’s central ethical promise. The platform does not decide for residents. It helps them understand how a Town decision works, which statements are facts or claims, whether the information is ready for a decision, which choices remain, what each choice may provide or require, who has authority and where the facts can no longer determine the answer.
Facts can narrow the responsible options. They cannot tell every resident how to weigh affordability, services, education, efficiency, continuity, growth, character, preservation or change. CivicSS should make that transition visible rather than hiding a value judgment inside an apparently objective report.
Sherie also performs the human user research Codex cannot manufacture from documents. She observes where people become confused, overwhelmed, disengaged or mistrustful. She talks with residents and Town participants. She tests whether different people can use the same information and still reach different responsible conclusions. She asks whether the reporting helps someone understand their role well enough to attend, ask questions and vote, whatever that person ultimately believes.
Her interaction with Town officials, employees, boards, committees and residents supplies lived and relational context. It helps reveal what each participant knows, controls, fears, recommends or still needs; where a public process is open but not understandable; and where better alignment may be useful. CivicSS can surface that opportunity. It cannot impose a new Town process, claim sponsorship or become a shadow authority.
The human therefore does not disappear at the end of automation. Sherie defines the purpose and ethical code, researches the human experience, interprets civic meaning, maintains real relationships, sets acceptable risk and makes consequential authority decisions. Codex makes the information and process sufficiently organized that she can spend more time doing those things.
Codex carries the complexity of the information. Sherie carries the purpose, people, values and decisions that give the information meaning.
7. What is genuinely new here
No individual component is unprecedented. Organizations use roles and reporting lines. Databases separate current state from audit history. Knowledge systems preserve source links. Software teams test releases. Training programs distinguish instruction from demonstrated competence.
The contribution is their integration around long-running Codex use by one human:
- conversational AI tasks are treated as bounded specialists rather than one universal assistant;
- institutional memory is selective, sourced and external to the conversation;
- a fresh task can recover the role and reasoning without inheriting the whole old context;
- task development includes training, applied qualification, drift recovery and successor transfer;
- organizational current state is queryable rather than inferred from task names or chat history;
- failures remain part of the memory instead of being overwritten by the corrected answer;
- operating commands are moved from remembered prose toward rereadable, testable procedures; and
- the human remains the source of purpose and authority without remaining the permanent integration layer.
The most important emergent finding is this:
The same system that makes AI work recoverable can make AI workers developable.
That capability may remain useful even if the larger multi-task organization proves too complex for sustained operation in the desktop product.
8. The maturity boundary
The work is best described as an advanced proof of concept or early bounded operating pilot.
It has progressed beyond a paper design. It includes real specialist tasks, durable records, current-state structures, applied tests, independent rejections, recovery events, successor training and live coordination.
It has not demonstrated:
- long-term unattended reliability across dozens of persistent tasks;
- universal recovery after compression, restart or task replacement;
- native Codex enforcement of the user-built command and registry conventions;
- guaranteed role-addressed delivery and acknowledgment among tasks;
- automatic continuation of every dependency chain to its terminal outcome;
- complete task-level resource isolation or working-set control;
- elimination of human correction; or
- equivalence between any task and the senior Partner role.
Some mechanisms may belong in a purpose-built orchestrator rather than Codex desktop. A credible evaluation should allow that conclusion.
9. Why it may generalize
CivicSS is unusually demanding, but the underlying problem is common to any long-running project that combines many domains, evolving evidence and consequential human judgment.
Possible applications include complex research, litigation support, public policy, organizational transformation, product development, consulting, due diligence, family-office work and long-running creative production.
The model is probably unnecessary for a short or self-contained task. Its value rises when:
- work lasts longer than one useful conversation;
- specialists need to work in parallel;
- decisions depend on provenance and changing evidence;
- experienced tasks will eventually be replaced;
- the cost of reconstructing context is high;
- failed reasoning should remain visible; and
- the human’s scarce attention belongs on judgment rather than coordination.
10. Implications for Codex
The experiment suggests that the useful unit of work can become larger than one task. A user may not want a single assistant with an indefinitely expanding memory. The user may want a team whose members can remain bounded while the organization preserves continuity.
Native product support could reduce much of the user-built machinery:
- durable roles and team views above individual task identities;
- selective, source-aware recovery after compression or replacement;
- visible task health, qualification and current instruction state;
- role-resolved cross-task messaging with acknowledgment and artifact fallback;
- workflow continuation that distinguishes handoff from terminal completion;
- small versioned command contracts loaded before ordinary interpretation;
- resource visibility and the ability to unload working context without losing durable continuity; and
- clearer guidance about when Codex desktop should yield coordination to an API-based agent system.
OpenAI’s current developer platform includes reusable agents and optional multi-agent configuration, making the broader direction directly relevant. This specific experiment, however, was built as a user-space operating model around desktop tasks. It should be evaluated as field evidence about human needs, not as a claim that the current desktop product or agent platform already supplies every required guarantee.
Conclusion
Four days did not produce a finished autonomous company. They produced something more credible and more useful: evidence that one person could begin turning many finite AI conversations into a recoverable organization.
The organization worked because the tasks did not all know everything. Each could hold a bounded specialty. Shared records exposed identity, ownership, current state and evidence. Training and evaluation made competence visible. Failure became recoverable learning. A coordinating role held the integrated outcome. The human retained purpose and authority.
The central research question remains open at production scale. But the bounded result is real:
A fresh task can inherit institutional memory without inheriting an entire old conversation. Specialized tasks can develop and coordinate without becoming one undifferentiated assistant. And a human can lead the system without remaining its permanent message bus.
That is the opportunity: not simply a more capable AI conversation, but a durable human–AI team.