Finding out why an answer got worse
Staff can change what the Sustentus assistant says by publishing text, without a release. That is a fast way to improve an answer and an equally fast way to spoil one — and until now the only way to tell the difference was to wait for somebody to complain. Three changes close that gap.
What a turn actually did
Agent runs in the console lists the assistant’s turns for a chosen month, failures first. Each one shows which assistant answered, which model served it, how it ended and why, which workspace it was for, and every tool it called — with the workspace each tool was bound to and how long the step took. Beside it is the exact set of prompt versions the answer was composed from, and each one links straight to that text in the prompt editor.
The page can also show every turn, not only the failures, so it can be opened to confirm that nothing is wrong rather than only to find out what is.
Whether a prompt change helped
Each assistant now has a page of its own showing every version of every file it reads, newest first, with who published it and why, and a comparison against the version before it.
Beside each publish are the answer-quality figures either side of it: how many turns ran, how many failed, and how many were refused. Those are matched by the versions each turn actually ran on rather than by the clock, so the comparison holds even after a restore. Where nothing has run yet the page says so instead of showing zero — “nobody has asked yet” and “nothing has gone wrong” are very different facts about a change you are deciding whether to keep.
Error and refusal rates by month sit underneath, with the models that served them, so a change in quality can be read against a publish, against a change of model, or against neither.
Checking a prompt before it goes live
The same page runs a set of checks against an assistant’s prompt: that figures come back as claims with their definitions attached, that it still names the tools its answers depend on and uses them at the right scope, that it still says what the assistant is for, and that answers which owe the reader a caveat carry one.
The checks can be run against what is live or against unpublished drafts — so a prompt can be scored before anyone sees it. They are advisory: running them changes nothing and blocks nothing, and a publish goes ahead whatever they say. They are there to be read, not to be obeyed.