Structured-work: giving coding agents a place to resume

September 28, 2026

A coding session can end with working code and a poor handoff. The next session opens the repository, finds a plan, reads a summary, and still has to work out which decisions were accepted. A passing test appears in the notes, but the notes do not say which version it tested. Somewhere in the conversation is the reason a seemingly useful feature was left out.

I want the next session to begin with enough information to make a responsible next move. That is the problem behind structured-work, a project operating model for work that spans coding-agent sessions. It gives current state a designated home and connects it to the decisions and evidence that support it.

The repository contains a written guide and a Python tool for checking workspace structure and generating navigation graphs. The guide uses Codex as its reference runtime. Its central idea is that each kind of project information needs a clear owner, so an agent can tell which document governs the work.

Start with the working agreement and current context

Consider a small example: an agent is adding CSV export to a reporting page. During the session, someone proposes moving large exports to a background job. The team keeps the first version synchronous, gets the formatting tests passing, and stops before checking the browser download.

When work resumes, the next agent needs to recover that boundary. Otherwise, it might spend the session building the queue that was only proposed, or report the feature complete because the formatting tests passed.

Structured-work makes the starting point explicit. The agent reads the canonical AGENTS.md, which holds the working agreement, followed by the designated CONTEXT.md, which holds the current operating state. The two-entry rule is meant to establish the goal, permitted scope, and next action before the agent follows deeper references.

The agreement contains durable instructions: which files are authoritative, where the application lives, and which actions require permission. Current context contains information that changes during the work. Mixing these together makes the working agreement accumulate yesterday's progress, while today's next step becomes harder to find.

An illustrative context entry for that task could read:

## Goal Finish CSV export for one report. ## Scope Use the existing synchronous endpoint. Background jobs remain a proposal. Follow the export specification. ## Status Formatting tests pass for this version. Browser download: not yet checked. ## Open decisions None blocking the synchronous version. ## Next action Check the CSV download in a local browser. Compare it with the export specification. Record the result before closing the task.

An actual project would link the specification and test evidence beside those claims. The entry should give the next agent a route into the work, with enough detail to avoid repeating discovery. Source files and detailed records remain available when the task calls for them.

Give decisions a home of their own

Current context is a poor place to preserve every explanation a project has accumulated. It needs to remain short enough to read at the start of a session. Structured-work puts durable reasoning in MEMORY.md, while specifications and architecture decision records retain the decisions they already own.

An architecture decision record, often shortened to ADR, describes a design choice and its reasoning. If an ADR owns the decision to keep exports synchronous, context points to it. Copying the full decision into memory, context, and a handoff would leave several versions to reconcile after the next change.

The same discipline applies to evidence. SOURCE-MAP.md records where a source came from, what was inspected, and what that inspection supports. A test result belongs to a particular implementation and set of checks. Recording those limits gives a future reader a way to judge whether the result still applies.

This is especially useful when a proposal sounds convincing. In the CSV example, a later document might explain why background jobs would help with large reports. Its date and level of detail do not make that design accepted. The supersession rules require an accepted replacement to identify the scope it replaces.

Even an accepted decision about background jobs would leave unrelated export requirements in force. The required columns and access rules would still have their own owners. That narrower treatment of change helps preserve decisions that were never being reconsidered.

Old evidence also keeps its meaning. A test that passed against an earlier version remains evidence about that version. It needs a fresh check before it can support a claim about changed behavior.

Use graphs to find the supporting material

Once decisions and evidence have clear owners, a graph can help an agent find the relevant material without loading every document. Structured-work distinguishes a knowledge graph, which connects explicit records, from a code graph, which points into the implementation.

The included graph generator has a deliberately limited view. Its knowledge graph uses registered records and their declared relationships. It does not read a paragraph and decide that a proposal has become accepted architecture.

The code graph provides a module index. It extracts Python imports using Python's syntax parser and uses text-based extraction for relative JavaScript and TypeScript imports. It does not resolve a complete call graph, package aliases, or runtime wiring. An import relationship can point you toward a file worth inspecting; checking whether the export button reaches that code still requires other evidence.

Generated graphs include fingerprints of their inputs so changes can be detected. When the underlying records or code change, the affected graph can be regenerated. If a graph disagrees with an accepted specification, the specification keeps its authority. The graph's job is to help the reader reach the source.

A structural pass leaves work to review

The reference tool exposes separate check and generate commands. check reads the declared workspace inputs and prints a JSON report. generate writes the declared graph outputs, with checks that prevent it from replacing an existing file it does not recognize as its own.

One detail I like is the checker's handling of success. The repository's minimal workspace example can produce these two fields together:

{ "status": "needs_review", "structural_status": "pass" }

The structure can be valid while review remains open. The checker implementation retains separate obligations for reviewing meaning, privacy, fresh-reader orientation, and product behavior. A structurally valid check with those obligations returns exit code 2.

For the export task, valid paths and consistent record references would tell us little about the downloaded file. We would still need to check that the browser receives the expected CSV and that the feature follows its specification. Similarly, a source record pointing to a test report does not prove that the test ran.

The guide also proposes a useful orientation test: give a fresh reader only the working agreement and current context, without the previous conversation or expected answers. Ask them to recover the active scope, blockers, authority, and next action. If they cannot, the starting documents need attention. If they can, the result supports the quality of the handoff; the underlying claims still need their evidence.

Start small enough to maintain

All of this creates maintenance work. Someone has to update context when the next action changes and keep links to decisions accurate. Stale summaries can send an agent in the wrong direction even when the folder structure looks orderly.

The guide addresses that cost by separating setup weight from work mode. A light setup requires a working agreement, current context, and a source map. Standard adds durable memory and a knowledge graph, with a code graph required when substantial project-owned code exists. The current task can involve research or a prototype without forcing a new directory layout.

For an existing repository, adoption begins by finding the owners already in use. The guide preserves the application's Git root, build paths, and established conventions. A project with a useful decision record should keep it. The role mapping can point to that document instead of creating a competing version in a preferred folder.

I would start by looking at one interrupted task. Write down the accepted scope, link the evidence behind its current status, and name the next permitted action. Then read that handoff without relying on the conversation that produced it. Any answer that still requires remembering the chat is a specific gap to fix.

The repository's minimal workspace example is a practical place to try the checker. It uses a synthetic light workspace with no application or Git root, so you can inspect the report before deciding how to map the model onto an existing project.

GitHub
LinkedIn
X