Where a file comes from (INPUT or GLOBAL) and where an MCP server comes
from (CATALOG or CUSTOM) are two options to pick between, not a value to
go back to. A new row still starts on INPUT or CATALOG, and blank still
reads as it; only the button goes. It stays on the model parameters,
where clearing a knob really does hand it back to the model's own default.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A node listed the tool names it had to have used successfully before its
answer was believed. Nobody could tell what to write there without
knowing each server's tools, and listing them from the editor would take
the server's credentials, which may only exist at run time.
Each bound server now says instead whether it must be used: the node
fails, with LLM_REQUIRED_MCP_SERVER_UNUSED, when the model finishes
without one successful call to a required server, and the message says
which one was never called or only ever failed. A made-up tool name
counts for no server. The per-tool list is gone; a flow that still
carries it reads as before, without the check. The fixture that used it
now requires the development server and the browser.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When the answer did not carry the fields the node declared, the step
failed saying the answer was on the response port, unchanged - but a
failed step publishes no port, so the reply was gone and the message
said otherwise. The failure now quotes it instead: the end of the reply,
where the object was asked to be, or the object the model wrote when a
field is missing from it. It also carries its own code,
LLM_STRUCTURED_OUTPUT_MISSING.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only the three most recent fields had an order, so they came first and
the model and prompt last. Every field now has one: the model, the
prompt and the files sent with it, then what the node may use - skills,
MCP servers and the tools it must use successfully - and last the
structured outputs it hands back.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The tools a node must have used successfully are checked by the loop
that bound MCP servers start; a node without one runs a single call and
never looks at them. The editor showed the field regardless, inviting a
setting that could do nothing. It now appears once the node has an MCP
server, through a visible-when rule that can say "once this holds
something" - the same present rule enabled-when already had.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Headed "Attached document", the text of a PDF read for a model that
takes only text made that model answer it could not open files: it read
the heading as a file it had not been given and never got to the text
under it. The heading now says the full text is included, between two
markers, to be read directly.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An LLM node gains the upload rows MCPAgent has: each is a file port of
the step, or a global input of the flow, holding images (PNG, JPEG,
WebP) or PDF documents. They are sent to the model with the prompt.
- Messages carry files, and each provider says which kinds it takes.
Ollama takes images, the OpenAI family images and PDFs (a compatible
gateway images only), Gemini both. A provider that takes none refuses
a message with files rather than flattening them away.
- Whether a model sees images is asked of Ollama per model, from the
vision capability of /api/show, remembered like the thinking one.
- A PDF goes as a document where the provider reads documents.
Elsewhere it goes as its text, added to the prompt, and - for a model
that sees images - as its first pages rendered with PDFBox.
- With MCP servers bound, the files ride on the opening message of the
tool loop, which pruning never touches.
When the files cannot reach the model the step fails with a code and a
message that says which model, which node and which file:
LLM_MODEL_CANNOT_SEE_IMAGES, LLM_DOCUMENT_UNREADABLE (a PDF with no
text for a model that sees no images), LLM_ATTACHMENT_INVALID, and
LLM_ATTACHMENTS_REJECTED when the provider itself refuses them, with
its own words. What was sent is logged as LLM_ATTACHMENTS_SENT.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Files sent to a model are declared, turned into ports, resolved and
checked the same way whichever node sends them. What MCPAgent had for
itself moves to shared places, with its behaviour and messages kept:
- UploadInput, the upload row (was MCPAgentUploadInput); its JSON is
unchanged, so saved flows read as before
- UploadInputs, the FILE ports and their picker constraints
- attachments/UploadAttachments, the files a row names, from a port or a
global input
- attachments/AttachmentValidation, the counts, byte budgets and
signatures, with the node named in each message
- attachments/FileAttachments and UploadedFileKind (were in mcp/)
This is what the LLM node needs to take files too.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Both columns hold 255 characters, and nothing checked it before the
database did: a longer description made saving a flow fail with a 500
and the SQL error as its message. Creating, updating and importing a
flow now refuse it with a validation error naming the field.
The assistant writes the name and description itself, so rather than
failing its own flow over a wordy description it shortens them to fit.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An execution's view carried each earlier round's steps but not whose
rounds they were: nodes do not carry their capabilities, so a page could
not tell a loop's guard from any other step. The view now lists the
loops the engine runs - guard, entry, the output that goes round, the
limit and the steps of each - the same after a reload as before it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every step's completion emptied the execution's variable maps and filled
them again, while steps running on other threads were reading them. Two
branches running side by side could fail with a
ConcurrentModificationException copying the map, or read it in the
moment it was empty.
Only what differs is now changed. Variables are declared up front, so
usually nothing is written at all.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A step hands its outputs on before marking itself completed. In a loop,
that let the guard run, go round and reset a step still finishing its
round - which then marked itself completed over its reset. And a guard
going round was marked completed for a moment before the loop reset it:
long enough for a check from another branch to find every step done and
end the execution in SUCCESS with the loop still going.
A step now settles - outputs, status, notification - under the
execution's lock, the one the loop's reset holds, and a guard going
round stays RUNNING until it is reset. A suspended container settles the
same way, taking the execution's lock before its own like everything
else; it used to take them the other way round.
LoopFuzzTest wires small flows at random and runs every one the
validation accepts, a person answering at random. Each must end in
SUCCESS with every step done, or in ERROR because a loop reached its
limit. Before this change about one accepted flow in a thousand did not.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An input holds one value. With two connections into it, whichever
arrived last silently won; the editor never let anyone draw that, but
nothing refused it in a flow imported or generated. A loop's entry is the
one input that legitimately takes two - from before the loop, then from
the way back - so the editor has to allow it there, and the server now
says no everywhere else.
A connection's loop settings are left out of its JSON when absent, so a
flow without loops serialises exactly as before.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Execution views now carry the round each step is on and the steps of
every earlier round, which is what a page needs to show a loop's history.
A node in a loop asks the same question once per round. The interaction
and evaluation endpoints take an optional iteration: an answer for a
round the loop has already left is refused with 409 instead of being
taken as the answer to a draft its author never saw.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When a loop's guard chooses the output that leads back, the execution
resets the loop's steps and starts the next round from the entry with
the value the guard produced. Every step of the loop leads to the guard,
so all of them have finished by then; the reset and the delivery happen
under one lock.
The back edge is not wired like other connections, which push their
value the moment it is produced, before the loop has been reset. Left
unwired, the entry's input still counts as open when nothing else feeds
it, so the first round takes it from the person starting the run. Inputs
from outside the loop keep their values from round to round. The guard's
other outputs are not marked as branches not taken while it goes round,
since a later round may leave through them.
Each step knows its round, events of loop steps carry it, and each
round's steps are kept in the execution's step history, so earlier
rounds and earlier verdicts stay readable after a reload. Choosing to go
round with no rounds left fails the execution with
LOOP_ITERATION_LIMIT_REACHED rather than leaving the loop.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An execution waiting for a person is read back from storage whenever it
leaves the in-memory cache or the service restarts. Its steps were rebuilt
with no executor, and resuming a waiting execution never gave them one. A
person's answer then made the next step ready, the step had nowhere to
run, and the execution sat in RUNNING for good.
Resuming a waiting execution now attaches the executor to its steps,
scheduling only a step that was already ready to run.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The execution-order check refused every cycle as a deadlock, and the
branch check refused a router whose branches meet again - which, going
round a loop, they do. Both now set a loop's back edge aside.
A cycle that cannot run gets the reason in its own code instead of a
generic deadlock: nothing to decide when it stops, a side way out, no
way out, loops inside loops, or a limit out of range. The assistant
treats them like the deadlock it already re-planned for.
The engine does not run loops yet; that is the next change on this
branch.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A loop can run when it goes round through a block whose type routes
exclusively: that block decides each time whether to go again or leave,
and the connection it goes round by is the back edge. Nothing marks that
connection; FlowLoops finds it from the graph, so validation and the
engine agree on it and imported or generated flows need no marking.
Only shapes the engine can run safely are accepted: one back edge per
loop, and nothing leaving the loop except through the guard's other
outputs. The author's iteration limit rides on the back edge, from 1 to
100, defaulting to the Loop container's 10.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Which blocks pick one output and leave the rest as branches not taken was
written out twice by name: in branch validation and in the engine's
not-selected marking. It is now a capability the type declares, so a
routing block added later is treated as one without either learning its
name - and loops can use the same capability to find their guard.
The loop container's default of ten iterations moves to a shared
constant, for connections that lead back to reuse.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Rebuilding dependency state on reload re-asked every step's activation
policy. A step that had already run has all its inputs, so the policy
answered READY, and a completed step came back from the database as one
waiting to start. Its dependents were then never told it had completed.
Activation now only decides for steps that have not run yet.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The catalog hides the IO lists the runtime derives, but keeps structural
configuration that happens to be called inputs or outputs. An LLM node's
declared output fields are that kind: the author writes them and each one
becomes a port. The test predated them and counted them as hidden IO.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An agent builds something, runs it and looks at it in a browser; a person then
opens the running preview and judges it, with the agent's own report hidden
until they have. That last part is the reason the evaluation node exists, and
until now there was nothing running for anyone to judge.
Checked here rather than only by being imported, because the two halves are
wired to each other by name. The agent declares previewUrl and agentReport; the
evaluation node names those same two in its target and reads one of them as its
blind reference. A typo in either imports perfectly cleanly and then hands the
evaluator an empty address at the moment they are asked to open something.
The agent node must prove it used start_service, browser_navigate and
browser_snapshot before its report is accepted. That is not caution in the
abstract: this node has already reported building and testing something without
making a single tool call, and nothing in the report said so.
It reports previewUrl and not publicUrl, and the prompt says why they differ -
one is reachable from inside the container network and the other by a person.
Handing over the wrong one produces an address that works for everybody who
tests it and for nobody who is asked to use it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A node could only ever produce one port: its whole answer, as text. That is
enough until it has two things to say to two different readers - the address of
something it started, and the report about it - at which point whoever is
downstream has to pick them apart by guesswork, or a parser block has to be
wired in to split on a delimiter the model may or may not honour.
Declaring fields turns the answer into ports the editor can wire. The mechanism
is only asking: the prompt gains an instruction to close with a JSON object, and
that object is read back out of the reply. No provider feature is involved, so
it works the same for a single call and for a whole tool loop, on every provider
in the catalogue.
It refuses rather than guesses. A missing object, or a declared field the model
left out, fails and says which - the same choice the loop already makes for a
tool call typed as prose. A port that silently arrives empty is worse, because
everything downstream reads the emptiness as an answer. The whole reply stays on
the response port either way, so a node whose structure disappointed is still
readable.
The closing object is found by scanning forwards, tracking strings: the last
balanced one wins, because an agent that has just written a JSON file tends to
quote it before answering. Scanning backwards was the first attempt and cannot
be made correct - a quote cannot be told from an escaped one without counting
the backslashes before it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A turn with no tool calls is how the loop knows it has the final answer. A model
that writes "<function=list_services>" into its content satisfies that test
exactly, so the step completed and handed the flow a call that had never run -
a success carrying nothing, which is worse than a failure because everything
downstream believes it. Seen with qwen3-coder on Ollama, which advertises tool
support and then answers this way when it is left to decide about reasoning for
itself: the trace comes back mixed into the content rather than as tool calls.
The guard reads only markers a chat template emits and prose never does, so an
answer that merely talks about tool calling still comes back untouched. The
message says what to try, because the cause is the model and not the flow.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A flow could write software but never run it: the coding agent edits files and
that is where it ended. These two entries close the loop - start the code the
execution wrote, then open it in a browser and read what the page says.
Both are streamable-http, so an LLM node binds them with no new Java: the
retriever already filtered the catalogue by transport and the draft normalizer
already dropped unknown servers. The development server's preview is pinned to
the execution with x-preview-key, the way the coding agent's workspace already
is, so the scope travels in a header the service fills in rather than in an
argument a model could choose for itself.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nothing could answer it. Authentication is a stateless JWT, so there is no
session table to count; nothing streams to the browser, so there is no open
connection to count either. LoginEntity.currentSessionStartedAt looks like the
answer and is not - it only clears on an explicit logout, so anyone who closed
a tab stays "in session" indefinitely. Fair input to an average session length,
useless as a list of who is here.
ActiveUserTracker records a sighting from the authentication filter, the one
place every authenticated request passes. In memory rather than a column: a
write per request would put the busiest path in the application on the database
to record something nobody reads between restarts, and after a restart "nobody
is connected" is not lost state but the correct answer. Per instance therefore,
which is what this is deployed as; behind replicas each would report its own
callers, and the fix then is to ask each of them.
GET /stats/activity returns that alongside the runs a restart would interrupt -
RUNNING, SUSPENDED and WAITING, every state that is neither initial nor final.
WAITING belongs there precisely because it looks idle from the outside: someone
is mid-task on the other side of it, and it is the most expensive thing to lose.
The two are reported separately because they fail differently - an active user
loses an unsaved edit and comes back, an in-flight run is gone.
POST /auth/heartbeat exists only to give the editor something to call: it
answers nothing, the filter having already recorded the sighting. Without it
anyone reading a flow or typing a prompt makes no request for minutes and would
be invisible, which is the person a restart interrupts worst.
Logout was excluded from the authentication filter, so its @AuthenticationPrincipal
was always null and recordLogout never ran - sessions only ever closed on the
same user's next login. Removed from the exclusion list: SecurityConfig still
permits it, so it keeps working with a bad or missing token, and it now both
records the logout and drops the person from the presence map immediately
rather than letting them age out.
Window is five minutes, PRESENCE_WINDOW_SECONDS, wide enough to survive a
couple of missed heartbeats.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every Ollama call carried think:true unconditionally, on the belief - stated
in a comment - that a model without the capability would ignore it. It does
not: Ollama answers with 400 "gemma:7b" does not support thinking, so every
flow pointed at gemma, llama3 or any other non-reasoning model failed outright.
Resolved per model rather than per flow, in two steps so that neither an old
Ollama nor one behind a proxy is left broken:
- /api/show reports a capability list, and "thinking" in it is the answer.
Cached per endpoint and model, so it is asked once, not per call.
- When it cannot be asked - no capabilities field, /show unreachable - the
flag is sent anyway and the rejection itself is the answer: the identical
call is repeated without it and the result remembered. One wasted request
per model per process, and none after that. Matched on the server's own
wording, never the bare 400, so a misspelled model stays the error it is.
Off means the key is absent, not think:false - absence is the only state every
version of Ollama reads as "no opinion".
ModelParameters gains "reasoning" to overrule that per-model answer: ON sends
it whatever the model is and lets an incapable one fail rather than quietly
answering without it, OFF never sends it - the case detection cannot guess,
a model that does reason but where the tokens are not worth it. Empty, which
is what every existing flow has, is the per-model default.
An enum and not a boolean because the editor collapses a boolean to true or
false as soon as its group is opened (schema-driven-fields, next[key] =
rawValue === true), which would have turned reasoning off for the models that
have it. It follows MCPAgentUploadKind, lenient @JsonCreator included, since
a select nobody chose from posts "" and an enum cannot be coerced from that.
THINKING is deliberately outside LLMProvider.supportedParameters()'s default,
which is no longer allOf: only the Ollama protocol has the switch, and a
provider inheriting a claim to it would leave the run log silent about a
setting that reached nothing.
The MCP agent path carries the same setting but cannot detect anything - the
bridge is not Ollama and exposes no capability endpoint - so there it is
whatever the flow says, defaulting to on as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The attribution term of section 7(b) names the source, so the URL it names has to
be the one that actually serves it: anyone redistributing this must be able to
reach the original. Eight references moved together - the licence addendum, the
NOTICE, the README attribution, the citation metadata and the four in the POM,
including the SSH developer connection.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three things that belong to one machine rather than to the project: the README
named the build host and the account on it, the Ollama default URL pointed at an
internal deployment, and docs/ held thirty working notes written for this team.
The Ollama default is now the address Ollama listens on out of the box, so a
clone runs against a local one without editing anything.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
local.env carries the database password and the provider key of whoever runs the
service: it is one person's machine, not the project. .claude/settings.json is the
same kind of thing, a local tool's permission list. Both stay on disk and are now
ignored.
This removes them from future commits only. Their contents remain in the history
that is already on origin, so the credentials they held must be treated as known
to everyone with access to the repository.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The budget is in characters and a model's window is in tokens, and the
ratio between them is not a constant: prose runs about four characters
per token, a conversation of JSON, file paths and UUIDs closer to two and
a half. 60000 was sized for the first. An orchestrator reading its own
registry reached 66332 characters - roughly 26k tokens with the answer
still to generate - and a 32k window dropped the opening message, which
the provider reports as "no user query found in messages".
Halving the per-result cap attacks the same failure from the other end:
the current iteration's results are the ones pruning can never shrink, so
what one turn reads is what decides whether the next call fits. A model
that reads the same file twice in a turn spent 24000 characters of a
60000 budget on one file.
Every call now records what was sent and what was allowed back. Without
promptChars and maxTokens, a context failure could only be reconstructed
by reading the node's configuration beside a warning about characters -
which is how this one was diagnosed, slowly.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A verdict off the criterion's scale, or naming a criterion the node does
not have, is the caller getting it wrong. Reported as a 500 it reads as
an outage, and the message saying exactly which criterion is at fault -
the one thing that makes it fixable - gets buried in a stack trace.
Found by submitting one against a running service, not by reading the
code: the executor was right to refuse, the endpoint was wrong about
whose mistake it was.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The human nodes could ask a question or collect a paragraph, so a flow
that wanted a judgement about a running piece of software got back prose:
not comparable between two runs, and nothing a later node could branch on.
HumanEvaluationBlock asks instead for a verdict per named criterion, with
an optional script so two testers exercise the same thing, and evidence
files. It points at the target rather than hosting it - provisioning
environments or driving browsers belongs outside the engine, and the flow
already knows the URL.
blindUntilSubmitted keeps a reference verdict, usually an agent's, out of
sight until the person commits to their own, and the execution records
which came first. That turns the node from a place to rubber-stamp an
agent into something that measures the automation bias the catalogue
already names.
The engine needed only a way to submit several fields as one act, which
Step.interact already accepted; and the interaction contract now comes
from the block type, so the editor and the assistant stop keeping their
own table of block names.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A skill on a node is only an id, and the editor grew a rule of its own to
show the instructions behind it: it looked for the literal retriever name
"Skills" and built the endpoint by hand. That is a special case inside
machinery that is otherwise entirely schema-driven, and the next binding
that wanted the same view would have needed another one.
@FieldRetriever now takes a definitionUrl, emitted as
x-retriever-definition-url, and the editor offers the view wherever it is
declared. The MCP server binding deliberately declares none: its
definitions endpoint answers with a JSON schema rather than readable text.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The bridge coerces a tool argument from string to object whenever the string
happens to be valid JSON, even though the model emits it correctly and the
tool's own schema declares it as a string. The coercion happens inside the
bridge, between the model's response and the downstream MCP call: verified by
calling the model directly (arguments.content stays a string, byte for byte)
and by comparing a JSON-valid value (rejected) against a malformed or plain
text one (passed through untouched). The failing call never reaches the MCP
server, and the agent's retries under a different encoding are what turn an
unfinished operation into one that returns status: completed with the model's
last preamble as if it were the answer.
There is no fix available on our side for the bridge itself, so this removes
it from the path instead. An LLMBlock can now bind MCP servers straight from
the catalog and run its own tool-calling loop in this service:
- LLMProvider gains chatWithTools/supportsTools; only OllamaProtocolProvider
implements it for now. Tool arguments stay JsonNode end to end - never a
string, never re-parsed - which is the one change that actually closes the
bridge's bug rather than working around it.
- A native streamable-http MCP client (mcp/client/) talks to a server without
the bridge: initialize, tools/list, tools/call, session header handling,
both response shapes the spec allows.
- MCPToolServerBinding is a narrower binding than MCPAgent's, restricted to
catalog servers reachable over streamable-http - the ones this service can
call directly, not the stdio ones the bridge still hosts a process for.
- LLMToolLoop runs the model/tool/model cycle with real budgets: a wall-clock
deadline and iteration cap that fail the block explicitly rather than
return a partial answer, and a character-based context budget that replaces
older tool results with a placeholder once the conversation - plus the tool
schemas sent on every call, which do not appear in the conversation but are
not free either - grows past it. The iteration just completed is never
pruned, and a result under ~500 characters is left alone: shrinking it would
cost about as much as it saves.
- Ollama's done_reason now travels back as ToolChatResult.finishReason, so a
turn that answers nothing can say whether the model chose silence or
num_predict cut it off mid-thought - two different problems with two
different fixes, previously indistinguishable from the error alone.
- A new skill, mcp-context-economy, carries the operating rules a real run
against a 24-task plan exposed the hard way: write_file to create a file,
apply_patch only to edit one that exists, and never read a file straight
back after writing it or re-pull an already-inline document into the
conversation - each halves the context a node needs for the same work.
MCPAgent and the bridge are untouched: this is a second path, not a
replacement, for the one transport (streamable-http) this service can reach
without it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The editor derived the control from the field being optional, which put it on
nearly every field in every dialog. "You may leave this blank" and "leaving this
blank means something specific" are different claims, and only the second is
worth a control.
The claim is now made per field with @DefaultsWhenEmpty, published as
x-ui-defaults-when-empty. It goes on the five sampling parameters - where empty
means the provider decides, and no typed number gets that state back - and on the
three fields that declare a concrete default, which the editor already names
alongside. The value itself still comes from JSON Schema's own `default`: a
parameter has no value to name, only an absence to return to, so declaring
`default: null` would have said something false to every other reader of the
schema.
Providers also now report which sampling parameters they actually apply. All five
were offered to every provider and the unsupported ones were dropped at run time,
reported in a warning on an execution that had already happened - Gemini applies
no seed, the OpenAI-protocol providers no top_k.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nobody checked that descriptor.provider actually named a registered
LLMProvider bean: a removed or misspelled provider surfaced only at
runtime, with a bare "Provider not found" and no indication of which node
named it.
FlowDataValidator.validateLlmDescriptorProviders runs from validateBlock,
so the existing subflow recursion in validateContainerSubFlow already
covers every container and Loop guard subflow for free. Only the provider
is checked, never the model (that's step 12, and a hosted provider's
catalogue isn't known here anyway), and a templated provider name is
skipped defensively even though nothing in the codebase ever writes one.
Since this constraint backs ValidFlowStructure, the new
LLM_PROVIDER_NOT_FOUND error surfaces two ways: saving a flow with one
still succeeds, as DRAFT, with the error in its validation list - the same
treatment every other not-yet-executable state already gets - while
creating an execution from one is rejected outright, since there would be
nothing such an execution could ever do.
This surfaced a pre-existing, widespread test convention: seven structural
validation tests used a placeholder provider name ("testProvider") that
was never a real bean, only ever exercised through flow save/execution
creation, never through an actual provider call. Renamed to "InternalOllama"
in all seven, the one provider name always registered in a full Spring
context.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A wrong model name only ever failed on the first call to it, which can be
minutes into a run for a step deep in a flow - by which point the endpoint
and credential that would have let it fail immediately were already known.
AuthorizationRequirementResolver.resolveAllDescriptors mirrors the existing
requirement-collecting walk (blocks, containers, Loop guard subflows) to
list every LLMDescriptor in a flow instead. ExecutionsService verifies each
one against its provider's own catalogue, for a provider whose
canListModels() is true, before starting - today that is only our own
Ollama, since every hosted provider declares canListModels() false
precisely because it cannot be asked without a credential the check does
not have. A model that is empty or still a template placeholder is
skipped, since its real value is only known at call time; a catalogue
that cannot be listed just now does not block the run either - this is a
defense in depth, not a gate a transient network failure should be able
to close.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
BiasJudgeRequest carried no way to supply a credential: a judge whose
provider required one only ever worked by coincidence, when the baseline
execution it judges happened to already carry a saved credential for that
same provider. There was nowhere to pick a credential for the judging
itself, unlike the interaction simulator and the assistant.
BiasJudgeRequest now carries an optional credentialId, mirroring
AssistantLlmSelection and the simulator's own field from the previous
commit. It flows through the whole asynchronous path -
BiasExperimentsController, BiasImpactJobService.createJudgeJob (a new
judge_credential_id column on BiasImpactJobEntity, so a job recovered after
a restart keeps it), BiasImpactService.judgeReport - down to
BiasImpactJudge.resolveAuthorization, which resolves it via
UserSecretService.resolveCredential when present and falls back to the
existing baseline-authorizations lookup otherwise, unchanged.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
startSimulationExecution resolved the simulator's provider authorization
lazily, inside whichever step happened to reach it first - a credential-
requiring simulator would start a RUNNING execution that then failed deep
inside a step, instead of being refused outright. ExecutionSimulationRequest
now carries an optional credentialId (mirroring AssistantLlmSelection), and
the service validates or falls back to an already-provided credential for
the same provider before starting simulation at all.
Fixing this surfaced a second, previously silent gap: a simulated
container's child never inherited the simulator's credential, since it is
never part of any execution's requiredAuthorizations and so the ordinary
per-container authorization propagation loop never touched it. Every
simulated container subflow would have started failing the same upfront
check once it went in, so propagateSimulatorAuthorizationToChild copies the
already-validated credential down to the child before it starts simulating.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
resolveProviderAuthorization used to name "InternalOllama" as the only
provider that could be used from the assistant without a credential, and
reject every other one with 409 - including a provider that plainly
declares requiresAuthorization() false, such as a credential-free remote
Ollama. The rule is now exactly that capability: !requiresAuthorization()
means no credential is asked for, whatever the provider is called.
The now-dead INTERNAL_PROVIDER_NAME constant goes with it - nothing else
referenced it.
AssistantSelectionResolverTest is new: this method had never been tested
in isolation, only indirectly through AssistantControllerTest, which
never exercised a credential-free provider under any name but the one
that used to be hardcoded.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Gemini is the provider that verifies AbstractHttpLLMProvider is really
transport only: it authenticates through a "key" query parameter, not a
header, so if the base assumed a bearer header this migration would have
had nowhere to go. It did not need to - the base has no opinion on
authentication mechanics at all, so query-string auth needed no hook,
just its own call site, same as before.
GeminiLLMProviderBodyTest was written first, against the
still-unmigrated provider, specifically to survive this move: it pins the
retry policy (ten attempts, 30s backoff), the two-minute timeout, the
role mapping (system and user both become "user", only assistant becomes
"model"), and the deliberate exclusion of seed from supportedParameters -
every one of which the interface's own defaults or the base's own
defaults could have silently replaced if an override were dropped by
accident. All thirteen assertions pass unchanged after the migration.
The retry policy and timeout are now the explicit overrides
AbstractHttpLLMProvider expects (retryPolicy(), timeout()) rather than
being built inline in the one method that used them - same values, same
pinned constants, just named as what the base already knows how to ask
for.
Gained for free, the same way Ollama did: 4xx responses now carry their
body, and every failure is logged - Gemini previously had no logging of
its own at all.
A live end-to-end HTTP test was attempted and dropped: Gemini's base URL
is a private constant with no way to redirect it to a local test server
without either changing production code or wrapping WebClient.Builder in
a test double fragile enough to break on the next Spring release. Given
that AbstractHttpLLMProviderTest already proves the shared HTTP mechanics
generically, repeating that proof through Gemini's specific, unreachable
URL would not have added real coverage.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
InternalOllamaLLMProvider carried its own client, its own error mapping
and logging, and its own request bodies and parsing, all mixed together.
OllamaProtocolProvider pulls out everything that is genuinely the
protocol - nested options, think:true, top_k, the JSON-path defaults the
flow assistant depends on, response parsing - onto the transport base,
leaving InternalOllamaLLMProvider about fifty lines: its own URL, its own
key, and nothing else. getName() still returns "InternalOllama" - that
string is persisted in seventeen places in workflow-editor-init/flows.json,
in every existing flow, and in the vault's provider column, so it could
not change even in a refactor this size.
Not extended from OpenAIProtocolProvider, even though Ollama also exposes
an OpenAI-compatible endpoint: the native shape differs enough - nested
options, think, top_k, none of which OpenAI has - that a subclass would
override every method the parent provides, which is not a subclass, it is
a different implementation wearing one.
The base ended up with two hooks instead of the OpenAI family's one,
because there is a real asymmetry here that family does not have:
resolveApiKey() lets InternalOllamaLLMProvider ignore whatever credential
a caller passes and always use its own server-configured key, while
RemoteOllamaProvider - the new provider, for an Ollama instance other than
our own - requires the caller's. baseUrl() has the same shape as
OpenAICompatibleProvider's: a constant for the internal instance, read
from the credential (and validated through OutboundEndpointGuard) for the
remote one. RemoteOllamaProvider cannot list its models either, for the
same reason OpenAICompatibleProvider cannot: listing would run from the
editor, with no credential and therefore no endpoint to ask.
InternalOllamaLLMProviderBodyTest - the existing test pinning every
request body byte for byte - passes unchanged, which is what "extraction"
is supposed to mean here. InternalOllamaLLMProviderHttpTest is new: no
test before this one exercised the actual HTTP round trip, only the
bodies, so there was no way to confirm the 4xx-carries-its-body upgrade
(the whole reason for building the shared base) actually reached Ollama
until now.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>