← WilbrahamCenter.org CivicSS™How it was built

How CivicSS was built · Shared with the ChatGPT Codex development team

From Many Capable Codex Tasks to One Recoverable Human–AI Team

A live field experiment in institutional memory, specialist development and team-scale coordination

One human. A town-sized question. A team of specialized Codex tasks learning how to remember, recover and work together.

Purpose + ethical codeUnderstand the decision. Never steer the answer. Invite what we may have missed.Capacity: Sherie Schaefer, as CivicSS™ creator.See how this promise is built into the system

CivicSS exists to improve the quality of civic decisions—not to determine their outcome. It brings together public records, numbers, uncertainties, choices and tradeoffs so people can understand what is possible and see clearly where evidence ends and personal or community values begin.

It helps people understand both the decision and their own role in it. It does not choose for people, manufacture agreement or promote a hidden outcome. It welcomes corrections, missing context and evidence that could change the picture. The goal is not agreement. It is a fair, informed process in which people with different values can understand what they are choosing and retain ownership of that choice.

Purpose

People cannot meaningfully participate in a process they cannot understand. CivicSS helps residents and Town participants see how a decision works, what is known, what remains uncertain, which choices exist, what each path may require and where their own participation belongs.

CivicSS does not exist to convince a community to adopt a preferred answer. It exists to improve the quality of the decision package and create conditions in which people with different values can make the wisest decision possible together.

Values

Understanding before judgment. Verification before conclusion. Accuracy over persuasion. Original sources over unsupported claims. Transparency over speculation. Curiosity before certainty. Fair treatment of responsible alternatives. Dignity through disagreement. Accessibility without oversimplification. Long-term trust over short-term influence. Stewardship of what future generations will inherit.

How the ethical code is built into CivicSS

Good intentions are not enough. A human can unconsciously frame a question, favor confirming information, use language that subtly steers or miss a conflict between roles. AI can also introduce bias, overstate certainty or miss a boundary. CivicSS therefore uses its human–AI operating model to question the work as it develops—and neither the human nor AI is treated as automatically objective.

That means CivicSS must:

  • ask whether the work helps people understand or subtly steers them toward an answer;
  • distinguish sourced fact, direct observation and calculation from interpretation, assumption and working theory;
  • trace material claims to original sources whenever possible;
  • preserve uncertainty, conflicting evidence, disagreement and missing information;
  • say clearly when the available information is not yet decision-ready;
  • test competing explanations and ask what information could change the picture;
  • treat responsible alternatives, competing values and affected people fairly;
  • show who may benefit, who may carry the cost and what another choice could displace;
  • avoid implying authority, endorsement or certainty that does not exist;
  • protect private, personal and student-level information;
  • invite new information, corrections and perspectives that may be missing;
  • correct material errors visibly; and
  • leave consequential decisions with the people and public bodies authorized to make them.

AI does not determine what is ethical, guarantee neutrality or replace human accountability. It provides a persistent second set of eyes—questioning assumptions, identifying possible bias and creating a deliberate pause when a boundary is unclear. Sherie Schaefer, as CivicSS™ creator, remains responsible for judgment, authority and publication.

Invite what we may have missed

See something we missed—or information that could change the picture? Share it through Wilbraham Center. New information will be reviewed and, where appropriate, verified before it is incorporated. Submission does not guarantee inclusion, endorsement or public display.

A growing platform with a durable foundation

CivicSS is still growing, and the ways these principles are built into its research, analysis, review and publication process will continue to evolve. The methods will improve; the commitments will not. These are the standards CivicSS begins with, returns to and works to apply as consistently and transparently as possible.

No process can eliminate every bias or error. The promise is to keep questioning, invite correction and make material changes visible.

Originator and human operating authority: Sherie Schaefer
Real-world proving ground: CivicSS™
Observed build period for the concentrated operating-model experiment: approximately four days in September 2026
Maturity: advanced proof of concept / early bounded operating pilot
Sharing status: curated external-review candidate; not a claim of production autonomy

Why we built it

The operating model exists for one practical reason: to free Sherie to do the CivicSS work that requires a human being.

Codex had already given one person extraordinary reach. Sherie could create specialist tasks for research, databases, public records, calculations, reporting, scenarios, writing and testing. Together, that team could attempt a civic-information platform with a breadth, depth and pace that no individual could reasonably achieve alone.

But as the team grew, Sherie became responsible for running it. She had to remember what every task knew, notice when context drifted, train replacements, reconcile ownership, route handoffs, find missing work and keep stalled processes moving. The machinery that expanded her capacity began consuming the same human attention it was meant to multiply.

That was especially costly because the work displaced was the part AI cannot perform for her:

  • defining the purpose, values and personal ethical code of CivicSS;
  • deciding which civic questions matter and why;
  • interacting with real residents, Town officials, employees, committees and boards;
  • observing what people understand, trust, fear or need before they can participate;
  • testing what information helps residents feel prepared to show up and vote, whatever they ultimately believe or choose;
  • interpreting lived, historical and relational context that cannot be established from files alone;
  • making clear where facts, calculations and decision readiness end and human values begin;
  • exploring whether Town participants see value in a more connected and understandable operating model; and
  • making consequential decisions about risk, representation, publication and direction.

The operating model is therefore enabling infrastructure, not the end product. Its value is not the sophistication of its organization chart, database or controls. Its value is whether those structures let one person lead a capable AI team without becoming its permanent administrator, and then use the resulting capacity to accomplish work that neither the person nor AI could accomplish alone.

The human role itself also emerged through the collaboration. Sherie did not begin with a complete job description for someone who would combine municipal-domain learning, civic systems thinking, resident experience research, public relationships, ethical stewardship, organizational-change exploration and decision-process design. Codex helped her connect those disciplines, test the connections in real work and recognize the higher-order role they formed. Without Codex, attempting that role would ordinarily require a large multidisciplinary human team. With Codex, one person can lead that supporting capacity while remaining responsible for the human work at its center.

Codex did not only help Sherie perform the role. Working together helped her discover what the role could be.

If the operating model does not return Sherie’s time to the uniquely human work of CivicSS, it has failed.

The hopeful human process

The intended result is not an AI organization talking to itself. It is a human being able to move continuously between the community and a depth of analysis that was previously out of reach.

Sherie can listen to a resident, attend a public meeting, speak with an official or notice where a major choice is becoming confused. The Codex team can then assemble the public record, financial context, physical assets, legal process, institutional responsibilities, uncertainties, options and prior commitments surrounding that choice. Sherie can add the history, relationships, values and human meaning the records cannot safely supply. CivicSS can return the whole picture in a form that residents and decision-makers can understand, question and use. Sherie can take it back into real conversations, learn what still does not work and improve it again.

Done well, that loop could help people understand their own role, see what a choice would protect or require, recognize where facts end and values begin, and arrive at public decisions better prepared—even when they still disagree.

That hopeful process is why the work became so large. It requires knowledge across municipal finance, property, affordability, education, public assets, capital planning, land use, law, public records, governance, data architecture, modeling, reporting, design and communication. It also requires resident research, ethical judgment, relationship awareness and organizational understanding. Reducing the number of tasks would not reduce the underlying problem. It would put an inherently multidisciplinary mission back onto one overloaded human or one overloaded conversation.

Codex made the needed specialization possible. The operating model became necessary to keep that specialization coherent, recoverable and in service of the human mission.

The question

How can one person use Codex to pursue work far beyond one person’s practical capacity while preserving the human purpose, relationships, values and judgment that make the work worth doing?

The obvious answer is to divide the work among many Codex tasks. That creates specialist depth and parallel capacity. It also creates a new problem: the knowledge, decisions and obligations become divided with the work.

The human can preserve continuity by keeping an experienced task alive, along with its accumulated assumptions, drift and context load. Or the human can start a fresh task and pay the cost of reconstruction, missing nuance and lost history.

We tested a third option:

A fresh mind with durable, sourced institutional memory.

What we built

Around ordinary Codex desktop tasks, we built a user-space operating layer with:

  • durable roles that can survive the task currently filling them;
  • specialized workers, managers, service providers and independent evaluators;
  • file-first knowledge with provenance and deep routes to evidence;
  • a database-backed registry for current identity, role, ownership, status and accountability, with separate history;
  • STORE PLAN and RECALL for selective cross-task institutional memory;
  • role-based training, teachback, applied qualification and bounded activation;
  • targeted recovery after drift, failure, compression or task replacement;
  • versioned procedures that a task rereads before performing routine operations; and
  • explicit human control over purpose, judgment, priorities and consequential authority.

We did not try to make every task know everything. We tried to make the organization know where its knowledge lived, who owned it, how it could be recovered and what remained unresolved.

What made the experiment different

This was not a hypothetical agent architecture or a demonstration built around toy tasks. The operating model emerged inside CivicSS, a real, continuing civic-information platform spanning municipal finance, property, affordability, assets, capital planning, schools, land use, public records, law, ethics, scenario modeling, reporting and resident communication.

The work was already too broad for one person or one task to carry. The operating failures were real. Context compressed. Tasks filled up. Important distinctions became difficult to find. New specialists needed training. Rejected work needed to remain visible. Handoffs could succeed mechanically while failing organizationally. Sherie repeatedly became the only person who knew how the whole system connected.

The prototype converted those failures into operating structures, tests and recovery procedures. Then the structures were used in live work.

Three results worth examining

1. A fresh task recovered nuanced work in 3 minutes 14 seconds

A separate evaluator received a compact discovery view rather than the originating chats or the entire Knowledge Base. It selected five relevant segments, verified hashes, opened thirteen linked files and correctly recovered two nuanced subjects, including important limitations. Both subject tests passed.

The result did not prove universal memory recovery. It demonstrated that a fresh task could recover useful, sourced institutional knowledge without replaying the original conversations or inheriting all of their accumulated context.

2. Failure became professional development rather than silent drift

One management exercise scored 16/20. The failed result remained preserved. The task received focused retraining on the exact missed competency and passed a new 20/20 exercise. Other deficient candidates were independently rejected before live use, corrected and retested rather than relabeled as complete.

The system began treating an AI task as a developable worker: preserve proven capability, isolate the gap, teach the missing layer and test it through applied work.

3. The human stopped being the only coordination system

In a bounded live sequence, a senior integrator set the architecture and authority boundary; an operating PMO coordinated the outcome; lifecycle operations maintained current state; evidence custodians protected the artifacts; specialists retained subject ownership; and independent evaluators could stop deficient work.

Sherie still supplied purpose, judgment and occasional high-value corrections. She no longer had to perform every technical, administrative and routing step herself.

The human role became clearer, not smaller

The operating model created a more valuable separation between human and AI work. Codex could carry more of the evidence, calculations, source control, testing, recovery and routine coordination. Sherie could concentrate on the civic purpose, platform ethics, human experience, real relationships and consequential decisions.

CivicSS does not decide for residents. It helps people understand how Town decisions work, what the facts and uncertainties are, whether the available information is ready for a decision, which choices remain and where the evidence ends and personal or community values begin. Sherie defines and protects that promise. She observes how residents and Town participants actually understand the material, supplies lived and relational context, tests whether the reporting helps people participate on their own terms and decides what the platform may publish or represent.

Codex carries the complexity of the information. Sherie carries the purpose, people, values and decisions that give the information meaning.

The central finding

We did not merely preserve AI memory. We began building the conditions under which specialized AI tasks could learn, coordinate, fail, recover, transfer responsibility and continue operating as a team.

That is the difference between several capable conversations and a human-led AI organization.

The larger human proposition is equally important:

Codex can let one person assemble enough specialized knowledge and execution capacity to attempt work no individual could complete alone. The operating model protects that leverage by keeping the human focused on the work only a human can do.

This is not a claim that one person has become a substitute for every expert or public institution. It is evidence that one person can use Codex to assemble, direct and learn from capabilities that would otherwise demand a large human team—while preserving the authority, expertise and lived judgment that must still come from real people.

Why this may matter to Codex

The prototype exposed product opportunities that individual prompting cannot fully solve:

  • durable role and team views above individual tasks;
  • selective recovery after compression or replacement;
  • source-aware shared memory without loading an entire archive;
  • dependable cross-task addressing, acknowledgment and continuation;
  • visible task health, training and recovery state;
  • versioned commands that bind user intent to a complete procedure;
  • task-level working-set and resource controls; and
  • explicit support for keeping the human at the judgment layer rather than making the human the permanent message bus.

The work also raises an honest platform question: which parts belong natively in Codex desktop, which are better implemented through the OpenAI agent platform, and which should remain in a user-controlled external system of record?

  1. 01-RESEARCH-BRIEF.md — the complete, self-contained argument, including the human story and purpose.
  2. 02-SYSTEM-AT-A-GLANCE.md — the compact structural model, boundaries and two connected operating cycles.
  3. 03-THREE-DEMONSTRATIONS.md — the strongest bounded evidence.
  4. 04-CODEX-PRODUCT-OPPORTUNITIES.md — capabilities the experiment suggests.
  5. 05-EVIDENCE-LIMITATIONS-AND-NEXT-TESTS.md — what the evidence does and does not establish.

The much larger v1.1.0 research archive remains preserved separately. It contains the full thesis snapshot, detailed case studies, technical programs, SQL, tests, audits and manifests. It is supporting evidence, not the first-reader experience.

What we are asking OpenAI

We would value a product or research conversation about four questions:

  1. Which parts of this operating model align with the intended direction of Codex?
  2. Which user-built mechanisms point to capabilities that should exist natively?
  3. What would be the smallest credible joint evaluation of recovery, coordination and human attention savings?
  4. Is this best understood as a Codex desktop workflow, an agent-platform architecture, or a useful bridge between the two?

The goal is not to claim that a finished autonomous organization has been created. It is to share evidence from a demanding real use case in which the unit of value became larger than one excellent AI conversation.

Next

Full brief