Developer Workflow: When the Agent Says It Is Done

A finished ticket is the agent's claim. Verified means someone checked and dated it. How AuricIDE keeps the two apart.

&&01->#~&&DEVELOPER WORKFLOWAuricIDE · Blog

Day 1, 09:12. An agent finishes a ticket, and the ticket flips to Done.

Day 2, 10:00. Someone opens the project. AuricIDE archives the ticket.

Day 40. A teammate asks whether the feature behind that ticket still works. The requirement it served says active, and its Last Verified field reads Never. By then the ticket is archived, and nothing in the project can say what anyone checked, or whether anyone did.

What did the ticket record?

The agent's claim. That is all a Done ticket says: the agent finished its work and said so, and nobody has to look at the result for the status to change.

A parcel can be marked delivered while nobody has opened the box. The ticket is the tracking page. The requirement is what was in the box, and nothing shows that anyone looked.

A digital illustration showing a developer workflow automated by a robot handling task management and archiving.

The two records also run on different clocks. AuricIDE tidies tickets on its own. A requirement's status stays put until somebody changes it. That gap, more than any single bug, is where a developer workflow around coding agents leaks.

Isn't a green test run enough?

The obvious answer is yes: if the agent finished and the tests pass, the work is verified. Passing tests matter. But a green run shows that the checks somebody wrote succeeded. It doesn't show that those checks covered the requirement, that the requirement was read correctly, or that anybody accepted the result.

So the workflow around a coding agent doesn't end when the agent reports done. It ends when somebody states what was checked, and the record puts a date on that statement. "Done" is a claim the agent makes. "Verified" is a claim someone stands behind.

A test result answers one question and a verification answers another. They support each other, but one can't stand in for the other. AI agent testing covers what a good check should assert. Even a good check is an input to the statement, not the statement.

What do the four Mission Control numbers show?

Mission Control is the overview screen of a project in AuricIDE, an open-source desktop IDE for running CLI coding agents. It has four number buttons in a row: Spec, Plan, Execute and Verify. On a wide pane, faint arrows sit between them, so it looks like a pipeline.

Mission Control is a summary. Nothing moves through it. No number advances another. Each button counts something and jumps to the place behind it.

  • Spec counts the files under specs/.
  • Plan counts open tickets, meaning anything not done, archived or discarded.
  • Execute counts running agents.
  • Verify shows how many requirements hold out of all requirements that are active, implemented or verified. A requirement whose check is out of date drops out of the held count.

Screenshot from https://auric-ide.tech

Read together, the numbers say nothing about progress. Spec can be high because files exist, Execute can be zero because no agent is running, and the amber button can still be up. The Verify button is the one that matches the scene. Ticket status and requirement status are two separate lists. A ticket can go Done and be archived while the requirement it served stays unproven.

AI agent monitoring covers what a live process and an exit code can and can't tell you about a single run. Verify sits after all of that. It asks whether a requirement has a current recorded check, not whether a run looked healthy.

Mission Control
Spec
3
files under specs/
Plan
1
open tickets
Execute
1
running agents
Verify
1/3
requirements hold
2 stale/unverified

Each number counts one thing. None of them moves the next.

Files under specs/
Agents running
Ticket status
Ticket 1
Ticket 2
Ticket 3
Requirement status
Requirement 1
Requirement 2
Requirement 3

Stale means the last verification is more than 30 days old.

What does Verify Now do, and what runs its clock?

Verify Now sits next to Last Verified on a requirement. Clicking it sets the status to verified and stamps the current time. That is everything it does.

It runs no test. It doesn't read the code. It doesn't execute the ticket's test cases either. Those are text handed to the agent. The app never runs them.

The click changes what Mission Control reports, because the requirement now counts as held. Then the clock starts. Once the last verification is more than 30 days old, the requirement is stale and stops counting as held. Mission Control shows an amber button reading "N stale/unverified", and the count includes requirements that are still active or implemented, so not yet verified. The amber button doesn't say the feature failed. It says nobody has stood behind it lately, or at all.

Tickets have their own, shorter clock. A ticket marked Done is archived after 24 hours. AuricIDE checks that when the project loads. Archived tickets leave the default board and show up only in the archive view. That clock describes the ticket record and says nothing about the requirement.

One optional check sits on the ticket side. The Conductor works through unblocked tickets on its own, and it has a Judge review checkbox. When it is on, an independent judge looks at each finished ticket before it counts as done. That makes "Done" harder to earn. It doesn't stamp a requirement as verified.

Day 0
Ticket: Done on day 0
Done

Under 24 hours since Done.

Requirement it served
active
Last Verified: Never

Verify Now stamps the current time. It runs no test.

Mission Control

Verify 0/1

Amber: 1 stale/unverified. The requirement is still active and nobody has verified it.

The ticket clock does not touch the requirement.

What does this design cost?

AuricIDE makes verification a recorded statement, not an automatic test result. Someone claims a requirement was checked, and the record shows when. The price is that Verify Now stamps a date and nothing checks whether the statement was true. Somebody has to look at the requirement, decide what counts as enough evidence, and click with that decision made. If nobody does, it stays unverified. The amber count stays up. The auric-pm MCP server also has a tool that marks a requirement verified, so an agent can stamp one as well, and the record doesn't say who did. A fresh date therefore means "marked verified at this time", and no more.

Test cases work the same way. They give an agent written instructions for checking. The app doesn't run them. That avoids pretending text is execution. It also leaves the actual run, the result and the judgment to the rest of the workflow.

The 30-day window is a fixed default. It isn't tuned per project. A safety-critical service, a slow internal tool and a short experiment may want different review habits. The stale signal only points at the question. Whether 30 days fits the risk is the team's call.

An automatic green light would be tidier. Think of a sign-off sheet on a workshop wall instead. It shows when somebody signed, and it can't prove the inspection was careful. The gain is a dated statement that doesn't pretend to be proof. The cost is the inspection itself, and on a small project that can feel like ceremony. Automatic verification would remove the click, but only a check tied to the requirement could make the resulting claim stronger.

What still holds in any other tool?

Any agent tool has to answer two questions separately: was the task finished, and does the requirement still hold? The second answer needs its own place, apart from "finished", and it needs a date that can expire. The ticket keeps what the agent completed. The requirement keeps what somebody marked as verified, and its timestamp tells the next engineer whether that is still current. A green check is evidence for a judgment. It isn't the judgment. And if a tool offers only one status, the second answer has to be written down somewhere that survives ticket cleanup.

A related problem is what an agent hands to the next step, covered in agent context management.

The record of what was checked has to survive that handoff, or the next person inherits a finished ticket and no answer.

AuricIDE is open source

AGPL v3, alpha, and built in the open. If the loop above sounds like the way you want to work, the code is the fastest way to judge it.

★ Star on GitHub