I started building Staves while working on an agent-supported report generation feature with Relae, an intelligent CRM for non-profits. I was having trouble holding the workflow in my head.

I needed to follow the internal APIs the agents called and understand how the information they returned ended up in a board report; I also needed to decide where a person should give direction or review the result. With AI-assisted coding, we could change the implementation faster than I could update my understanding of it.

The snapshots I was taking weren't enough. I wanted a view of the work that I could inspect with collaborators, including people who wouldn't be reading the code.

Staves grew out of that need. It draws the people, agents and systems involved in a workflow, with the jobs they do and the information they pass between them. You can ask your coding agent to describe something you've already built, or work through an idea before writing code.

I was also talking with design-driven product teams about who does design and when it happens. If the next round of implementation is already underway while we're still trying to understand the last one, how do we make room for a useful conversation about it?

For me, that conversation often needs to start with someone who knows the work.

An expert can explain why a routine case sometimes takes an afternoon, or what they check before trusting a result. A researcher can bring accounts from the people doing the job. I want that knowledge involved in deciding what to build, while there's still room to change the design.

The interview works through what the agents may do and where a person needs to decide.
The interview works through what the agents may do and where a person needs to decide.

The report workflow shown here is an illustrative example, with a scripted interview captured in the real Staves editor. It follows a problem similar to the one that led me to build it.

Sales and finance have different figures for revenue. One counts signed contracts; the other recognises revenue over the delivery period. An agent choosing the newer figure would miss the reason they differ.

The interview works through what the finance reviewer needs: source records for the same entities and cut-off date, with the definitions kept beside the figures. That gives the evidence agent a specific job. It can retrieve records and prepare the discrepancies for review. Finance decides what those figures support. A reporting agent can then draft from the reviewed claims, preserving the caveats, before the owner approves the report.

Report owner → evidence agent → finance review → reporting agent → owner approval.
Report owner → evidence agent → finance review → reporting agent → owner approval.

That division of work is worth discussing before implementing it. Perhaps the expert spends most of their time gathering evidence that an agent could prepare. Perhaps the difficult part is a judgment we haven't understood yet. Those lead to different designs.

I'd like Staves to help teams find where they can extend someone's expertise: give a reviewer more of the relevant evidence, or give a less experienced colleague enough context to handle a case and know when to ask for help. Whether that saves time depends on the work, including how much checking the assistance introduces.

On the board, we can propose a different arrangement and examine it alongside the original. In this example, the alternative adds an evidence checker before finance review. It checks for missing sources and dates; the financial decision stays with the reviewer.

A proposed evidence checker prepares missing-source and definition issues for finance review.
A proposed evidence checker prepares missing-source and definition issues for finance review.

Each agent job can open into its tasks and the tools it uses. That lets us get specific about what an API returns, what it leaves out and what should happen when the information is missing.

Inside the evidence job: retrieve records, compare definitions and prepare exceptions.
Inside the evidence job: retrieve records, compare definitions and prepare exceptions.

An unexpected benefit was the effect on conversations with the coding agents. Asking them to describe the human jobs brought up questions I wanted in the development process.

In one review of OwnedBy, research could continue after a shopper received an answer. But the code we inspected stopped the screen from checking for updates once it had resolved. The answer could improve without the person already reading it ever seeing the change.

Another review found a gap in the correction workflow. The description said an operator could reject a correction or request further research. The screen offered only Reject. What could the operator do if they believed the correction?

A third challenged the use of a company chooser to handle contradictory ownership evidence. Choosing between similarly named companies is a reasonable thing to ask a shopper to do. Deciding between conflicting claims about one company's ownership needs further research. The proposed interface had put that problem in front of the wrong person.

These observations don't establish that Staves makes an agent better at finding defects. They do show the kinds of questions that came up when we described the work from the person's position. Sometimes the description itself needed correcting.

I want collaborators to be able to have those discussions as the software changes. Someone who understands the job should be able to point to a misplaced responsibility or a missing decision without first turning it into a coding request.

Staves keeps questions alongside the work they concern. Descriptions can link to source files so changes can be flagged, and an agreed design can go back to a coding agent for assessment or implementation. Where execution records are connected, we can also compare what ran with what we thought should happen. An interview, a code inspection and a recorded run may disagree; each gives us somewhere to investigate.

I've made the Staves format, editor, MCP server and CLI open source. You can use them locally with your own agent or model key, and there's a hosted service at staves.io. I want small teams and individuals to have access to this way of working, and to keep the descriptions they make.

If you're building something whose behaviour is getting difficult to explain to the people involved, I'd like you to try it. That's where I started.

Try Staves · Source and local setup