The full view of an output or input gains a Copy button - Copy JSON for
a JSON value - and the outcome payload a copy button in its corner, which
stays put while a long payload scrolls. JSON is copied indented, whether
it came as an object or as a model's text in a fence, so it pastes as
clean JSON elsewhere; text is copied as shown.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sized to its content and centred on the page, the modal grew and shrank
with every node opened or closed, its top edge jumping each time. Showing
a JSON tree it now has a fixed height, and only its content scrolls.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Outputs, intermediate inputs, loop rounds and outcome payloads that are
JSON - objects, arrays, or JSON a model wrote as text, fenced or not -
are shown with the JSON tree instead of their text. A preview opens the
first level and a few entries, summing up the rest; the full view opens
the whole tree. The JSON viewer gains the depth and entry limits this
takes.
In a read-only graph, such as an execution's, connections can no longer
be picked or shown as selected: there is nothing to do with one there.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
For an execution with a loop, the side panel gains a Rounds tab: each
loop, which round it is on out of how many, and every round with what
its steps were given - the draft sent for review, the verdict handed on
- and how it ended: went round again, left the loop, or stopped at the
limit. Earlier rounds come from the execution's step history, the last
from its steps as they stand.
From the second round of a loop on, a step's inputs are shown as they
are now rather than as prepared at the start, which is only the first
round's value. This also corrects the Intermediate tab and the inputs
shown on a node in a loop.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The panel scrolled itself to the newest line on every poll, so a reader
who scrolled up to follow something got yanked back five seconds later -
exactly when a log is worth reading, and exactly the moment it became
unusable.
Following now stops as soon as the reader leaves the end, and resumes on
its own when they come back to it, so the common case needs no button.
The button is there to say which of the two is happening and to override
it: a log that stops moving on its own otherwise looks like a log that
stopped.
The threshold lives in the utils with the rule, not in the component: not
zero, because a list that grew by a line between the scroll and the
handler would otherwise read as the reader having walked away.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The log showed each event's sentence and threw away everything behind it.
A pruning warning saying "6332 characters over budget" could not tell you
what the total actually was, a tool call could not say which tool: those
numbers reached the browser and were dropped, leaving the API as the only
way to read them.
Rows that carry details now offer an eye, which opens them read-only.
Rows that carry none offer nothing - a button opening an empty box is
worse than no button - so the decision of what is worth opening lives in
the view model, where it can be tested, rather than in the template.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Mechanical: three comment lines at the top of every .ts, .html and .css
under src, and nothing else. Split out from the licence commit so the
files that carry an actual change stay readable in the history, and kept
to its own commit because it moves the blame line on 390 files.
The short SPDX form rather than the full GNU notice - it is
machine-readable under REUSE, it satisfies the requirement to keep the
licence notice intact, and it points at LICENSE-ADDENDUM instead of
restating the attribution term in every file.
The template and stylesheet headers do not reach the bundle: Angular
discards template comments and the production build strips CSS ones. The
one in index.html survives, since that file is served as written, which is
no loss.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The report now reads the per-subject sections the service produces: a row per
iterated subject with what its labelled numbers did (Score 7 -> 4), the two
texts behind it a click away, and the changed ones listed first. The band at
the top leads with what the run actually did - whether the final decision
changed, how many subjects moved, the largest delta - because the counts of
changed nodes that used to open the report were the least actionable thing in
it. A container's iterations are listed with the inner node that changed,
which is what a per-subject iterator run needs and the accumulated list could
never show. Reports produced before any of this exists still render, from
their raw outputs.
"Evaluate impact with LLM" sits next to the report it is about, in all three
places one is mounted, and opens the provider and model picker the interaction
simulator uses - extracted so both call the same dialog rather than two of
their own, with temperature 0 offered by default because a verdict that reads
differently every time it is asked for is worse than none. The assessment runs
as a job, polled like the isolated experiment, and is stored on the report, so
reopening it later shows the same verdicts and the model that produced them.
It is labelled an assessment throughout, and a pair the model could not answer
for is marked without hiding that pair's own figures.
A rerun of a simulated run now opens the Simulate dialog on the simulator it
inherited, with the inherited sampling out where it can be seen - a seed
carried over is the reason the two runs are comparable, and behind a closed
section nobody would find it. Before it is started, a run says which simulator
the run it repeats used; afterwards, both the bias report and the run-to-run
comparison say so when the two sides were not answered the same way, since
that difference is not the intervention's doing and nothing said it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A green play and a flask said nothing about the one thing that separates them:
whether you answer the flow's chat and decision steps, or an LLM answers them for
you. The flask was actively misleading - "experiment" is the bias experiments'
word, so it pointed at a different feature.
Both buttons now keep the play, and a small badge carries the difference: a
person, or a robot. A gear was the first thought and is the wrong glyph - it is
the settings icon everywhere else in the product, so on a run button it reads as
"configure this run" rather than "a machine runs it".
They also sit next to each other now. Stop and resume used to separate them, so
each had to be understood alone - which is precisely the job a 17px badge cannot
do. Side by side they are one choice, and each is what makes the other legible.
The tooltips stop repeating the button and say who answers: "Run - you answer the
human steps", "Run simulated - an LLM answers the human steps". That is the
sentence the icon can only gesture at.
557 frontend tests green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nothing in the run history said whether a run was a bias variant. Worse, the
viewer header claimed "Bias variant" for *every* execution: it tested for the
presence of biasExecutionContext, which the backend sets on all of them,
defaulting to NORMAL. Only the mode distinguishes them. The same faulty test also
offered the bias comparison action on any plain rerun.
Each history row now carries a kind badge - Run, Rerun, Bias, Mitigation, or Bias
+ mitigation - coloured and with a tooltip saying what it means for the result. A
bias variant reads as a variant even when it is also a rerun, because carrying
probes is what changes how its output should be read.
The rows also stop showing raw uuids. A rerun names the run it came from by
number ("from #1") instead of repeating a 36-character id, the execution id is
gone from the row entirely - the viewer header owns it - and the group header no
longer prints the source flow id. "1 runs" reads "1 run".
The viewer header keeps only what changes how a result should be read: Simulated,
the bias variant, Subflow. Execution id, simulator descriptor, experiment id,
baseline and probe internals moved behind a Details toggle, closed by default.
Two Italian strings in that block are now English, like the rest of the app.
Cards are flat here too, matching the flows list: no gradients, no lift on hover,
a left accent for the selected run.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An outcome payload is often a generated document - the local example is a full
rejection email - which pushed the execution graph down the page. The section now
collapses.
It starts open, because this is the flow's answer and it was invisible until a
moment ago. Collapsed, the header keeps the outcome codes, so the conclusion is
never hidden entirely: only the payload folds away.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A flow result is built only from *unconnected* outputs, so wiring the last block
into an End node moves its value out of the result and into the outcome payload -
which nothing in the UI read. The run then looked like it produced nothing at all.
There is already an execution in a local database whose outcome payload is a full
generated rejection email that was invisible for exactly this reason.
The run view now has an Outcomes section above the graph, listing each End the run
passed through: its code, its label, the step it came from, and its payload. An
End reached with no value says so, rather than showing an empty box - the two
cases mean different things.
A text payload renders as text rather than through the JSON tree. The tree would
have kept the line breaks but wraps strings in quotes, and the common case here is
a generated document, not a data structure.
This is the smallest of the options for the underlying gap: the data was already
persisted and already in the API payload, so nothing in the engine changed. The
gap itself remains - End cannot be attached as a pure ordering dependency
(canDependOnOtherNodes is false), so a single-exit flow still has to choose
between a labelled end and a value in `result`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Lower rete editor min zoom (0.35 -> 0.1) so wide flows fit on screen.
- Remove the Export action from the read-only subflow preview dialog,
since it's only ever opened from the execution view.
- Add a fullscreen toggle to the flow editor and task execution viewer,
placed next to the title; disables the "flows" sidebar tab and hides
the right assistant/validation panel while the flow editor is
fullscreen, and collapses the left sidebar if it was open.
- Allow dragging nodes in the read-only execution graph and persist
custom positions client-side in sessionStorage, keyed by execution id,
restored on reload/refresh.
Finishes the bias impact experiments plan (docs/bias-impact-experiments-plan.md
steps 7-12):
- Side-effect policy selector: reuse the .llm-warning visual language for the
external side-effects banner, add a REQUIRE_CONFIRMATION note.
- Full-flow compare ("Compare with baseline"): new bias-compare-dialog
(service + host) triggering compareBiasExecutions and opening the shared
report viewer; inline errors read from errors[].message/detail.
- Persisted reports: new bias-impact-report-list (list + detail in one view)
wired into a new "Bias impact reports" tab in task-execution-viewer;
404/403 on report detail show the same inline message on purpose.
- Canvas: annotation badge on generic-node (count, executable-probe
indicator, severity from the backend catalog); new
BiasComparisonViewStateService driving bias-active / downstream-changed /
routing-change highlighting on task-step-node and custom-connection, fed by
a highlightOnCanvas event from bias-impact-report-viewer wired in all three
places that render it; legend + "back to normal view" action in the canvas
toolbar.
- Fixed a bug where the bias variant context badge only rendered for
simulated executions.
- Fixed "Measure bias impact" to stay visible-but-disabled with an
explanatory tooltip while the baseline hasn't reached a final state,
instead of being hidden outright, per the §12 checklist.
- Added a Retry action to the compare dialog's error state for parity with
the report list.
- Added the end-to-end facade flow test (annotation -> capability -> isolated
experiment -> report) plus coverage for all new components/services.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>