Commit Graph

282 Commits

Author SHA1 Message Date
Lucio Lelii 4eb962ed9b Offer "Use default" only where a field declares one
The editor derived the control from the field being optional, which put it on
nearly every field in every dialog. "You may leave this blank" and "leaving this
blank means something specific" are different claims, and only the second is
worth a control.

The claim is now made per field with @DefaultsWhenEmpty, published as
x-ui-defaults-when-empty. It goes on the five sampling parameters - where empty
means the provider decides, and no typed number gets that state back - and on the
three fields that declare a concrete default, which the editor already names
alongside. The value itself still comes from JSON Schema's own `default`: a
parameter has no value to name, only an absence to return to, so declaring
`default: null` would have said something false to every other reader of the
schema.

Providers also now report which sampling parameters they actually apply. All five
were offered to every provider and the unsupported ones were dropped at run time,
reported in a warning on an execution that had already happened - Gemini applies
no seed, the OpenAI-protocol providers no top_k.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 11:52:00 +02:00
Lucio Lelii f4dbb866c0 Reject a flow that names an LLM provider nothing registers
Nobody checked that descriptor.provider actually named a registered
LLMProvider bean: a removed or misspelled provider surfaced only at
runtime, with a bare "Provider not found" and no indication of which node
named it.

FlowDataValidator.validateLlmDescriptorProviders runs from validateBlock,
so the existing subflow recursion in validateContainerSubFlow already
covers every container and Loop guard subflow for free. Only the provider
is checked, never the model (that's step 12, and a hosted provider's
catalogue isn't known here anyway), and a templated provider name is
skipped defensively even though nothing in the codebase ever writes one.
Since this constraint backs ValidFlowStructure, the new
LLM_PROVIDER_NOT_FOUND error surfaces two ways: saving a flow with one
still succeeds, as DRAFT, with the error in its validation list - the same
treatment every other not-yet-executable state already gets - while
creating an execution from one is rejected outright, since there would be
nothing such an execution could ever do.

This surfaced a pre-existing, widespread test convention: seven structural
validation tests used a placeholder provider name ("testProvider") that
was never a real bean, only ever exercised through flow save/execution
creation, never through an actual provider call. Renamed to "InternalOllama"
in all seven, the one provider name always registered in a full Spring
context.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 15:08:59 +02:00
Lucio Lelii 11c70115c1 Check a flow's configured models against their provider at start
A wrong model name only ever failed on the first call to it, which can be
minutes into a run for a step deep in a flow - by which point the endpoint
and credential that would have let it fail immediately were already known.

AuthorizationRequirementResolver.resolveAllDescriptors mirrors the existing
requirement-collecting walk (blocks, containers, Loop guard subflows) to
list every LLMDescriptor in a flow instead. ExecutionsService verifies each
one against its provider's own catalogue, for a provider whose
canListModels() is true, before starting - today that is only our own
Ollama, since every hosted provider declares canListModels() false
precisely because it cannot be asked without a credential the check does
not have. A model that is empty or still a template placeholder is
skipped, since its real value is only known at call time; a catalogue
that cannot be listed just now does not block the run either - this is a
defense in depth, not a gate a transient network failure should be able
to close.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 15:08:48 +02:00
Lucio Lelii ed128ac014 Let the bias judge take its own credential, not just inherit one
BiasJudgeRequest carried no way to supply a credential: a judge whose
provider required one only ever worked by coincidence, when the baseline
execution it judges happened to already carry a saved credential for that
same provider. There was nowhere to pick a credential for the judging
itself, unlike the interaction simulator and the assistant.

BiasJudgeRequest now carries an optional credentialId, mirroring
AssistantLlmSelection and the simulator's own field from the previous
commit. It flows through the whole asynchronous path -
BiasExperimentsController, BiasImpactJobService.createJudgeJob (a new
judge_credential_id column on BiasImpactJobEntity, so a job recovered after
a restart keeps it), BiasImpactService.judgeReport - down to
BiasImpactJudge.resolveAuthorization, which resolves it via
UserSecretService.resolveCredential when present and falls back to the
existing baseline-authorizations lookup otherwise, unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 14:42:51 +02:00
Lucio Lelii 058918f4d7 Require the simulator's credential upfront, not on first use
startSimulationExecution resolved the simulator's provider authorization
lazily, inside whichever step happened to reach it first - a credential-
requiring simulator would start a RUNNING execution that then failed deep
inside a step, instead of being refused outright. ExecutionSimulationRequest
now carries an optional credentialId (mirroring AssistantLlmSelection), and
the service validates or falls back to an already-provided credential for
the same provider before starting simulation at all.

Fixing this surfaced a second, previously silent gap: a simulated
container's child never inherited the simulator's credential, since it is
never part of any execution's requiredAuthorizations and so the ordinary
per-container authorization propagation loop never touched it. Every
simulated container subflow would have started failing the same upfront
check once it went in, so propagateSimulatorAuthorizationToChild copies the
already-validated credential down to the child before it starts simulating.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:54:26 +02:00
Lucio Lelii dec3295d6d Stop the assistant special-casing one provider by name for a credential
resolveProviderAuthorization used to name "InternalOllama" as the only
provider that could be used from the assistant without a credential, and
reject every other one with 409 - including a provider that plainly
declares requiresAuthorization() false, such as a credential-free remote
Ollama. The rule is now exactly that capability: !requiresAuthorization()
means no credential is asked for, whatever the provider is called.

The now-dead INTERNAL_PROVIDER_NAME constant goes with it - nothing else
referenced it.

AssistantSelectionResolverTest is new: this method had never been tested
in isolation, only indirectly through AssistantControllerTest, which
never exercised a credential-free provider under any name but the one
that used to be hardcoded.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:29:14 +02:00
Lucio Lelii 5bd2caf884 Migrate Gemini onto the shared HTTP transport base
Gemini is the provider that verifies AbstractHttpLLMProvider is really
transport only: it authenticates through a "key" query parameter, not a
header, so if the base assumed a bearer header this migration would have
had nowhere to go. It did not need to - the base has no opinion on
authentication mechanics at all, so query-string auth needed no hook,
just its own call site, same as before.

GeminiLLMProviderBodyTest was written first, against the
still-unmigrated provider, specifically to survive this move: it pins the
retry policy (ten attempts, 30s backoff), the two-minute timeout, the
role mapping (system and user both become "user", only assistant becomes
"model"), and the deliberate exclusion of seed from supportedParameters -
every one of which the interface's own defaults or the base's own
defaults could have silently replaced if an override were dropped by
accident. All thirteen assertions pass unchanged after the migration.

The retry policy and timeout are now the explicit overrides
AbstractHttpLLMProvider expects (retryPolicy(), timeout()) rather than
being built inline in the one method that used them - same values, same
pinned constants, just named as what the base already knows how to ask
for.

Gained for free, the same way Ollama did: 4xx responses now carry their
body, and every failure is logged - Gemini previously had no logging of
its own at all.

A live end-to-end HTTP test was attempted and dropped: Gemini's base URL
is a private constant with no way to redirect it to a local test server
without either changing production code or wrapping WebClient.Builder in
a test double fragile enough to break on the next Spring release. Given
that AbstractHttpLLMProviderTest already proves the shared HTTP mechanics
generically, repeating that proof through Gemini's specific, unreachable
URL would not have added real coverage.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:29:04 +02:00
Lucio Lelii 5e6636d3b9 Extract Ollama's protocol into a shared base; add a remote Ollama provider
InternalOllamaLLMProvider carried its own client, its own error mapping
and logging, and its own request bodies and parsing, all mixed together.
OllamaProtocolProvider pulls out everything that is genuinely the
protocol - nested options, think:true, top_k, the JSON-path defaults the
flow assistant depends on, response parsing - onto the transport base,
leaving InternalOllamaLLMProvider about fifty lines: its own URL, its own
key, and nothing else. getName() still returns "InternalOllama" - that
string is persisted in seventeen places in workflow-editor-init/flows.json,
in every existing flow, and in the vault's provider column, so it could
not change even in a refactor this size.

Not extended from OpenAIProtocolProvider, even though Ollama also exposes
an OpenAI-compatible endpoint: the native shape differs enough - nested
options, think, top_k, none of which OpenAI has - that a subclass would
override every method the parent provides, which is not a subclass, it is
a different implementation wearing one.

The base ended up with two hooks instead of the OpenAI family's one,
because there is a real asymmetry here that family does not have:
resolveApiKey() lets InternalOllamaLLMProvider ignore whatever credential
a caller passes and always use its own server-configured key, while
RemoteOllamaProvider - the new provider, for an Ollama instance other than
our own - requires the caller's. baseUrl() has the same shape as
OpenAICompatibleProvider's: a constant for the internal instance, read
from the credential (and validated through OutboundEndpointGuard) for the
remote one. RemoteOllamaProvider cannot list its models either, for the
same reason OpenAICompatibleProvider cannot: listing would run from the
editor, with no credential and therefore no endpoint to ask.

InternalOllamaLLMProviderBodyTest - the existing test pinning every
request body byte for byte - passes unchanged, which is what "extraction"
is supposed to mean here. InternalOllamaLLMProviderHttpTest is new: no
test before this one exercised the actual HTTP round trip, only the
bodies, so there was no way to confirm the 4xx-carries-its-body upgrade
(the whole reason for building the shared base) actually reached Ollama
until now.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:28:46 +02:00
Lucio Lelii a222c8a6d7 Add OpenAI, OpenRouter and OpenAI-compatible providers
OpenAIProtocolProvider is the shared request/response shape - flat
sampling parameters, native system/user/assistant roles, no top_k (OpenAI
has no equivalent, so it is excluded from supportedParameters rather than
silently ignored) - built on the transport base added earlier. Three
concrete providers sit on it:

  - OpenAIProvider and OpenRouterProvider have a constant endpoint, the
    way any client of either service does; requiresEndpoint() stays false
    and their baseUrl() ignores whatever the credential carries.
  - OpenAICompatibleProvider is the one whose endpoint the user supplies -
    a self-hosted vLLM, LM Studio, a company gateway - so
    requiresEndpoint() is true and baseUrl() reads the credential's
    endpoint, validated through OutboundEndpointGuard before every call.

Registering "OpenAI" as a real provider bean was checked against the
existing tests that used that exact name as a stand-in for an
*unregistered* provider (VaultCredentialGateTest, UserSecretControllerTest)
- none of them break, and the ones that specifically assert the
unregistered-name behaviour now exercise the real bean instead, which is
closer to what they were meant to prove.

The 4xx-carries-its-body improvement from the transport base applies here
from day one: with a free-typed model name, that body is often the only
thing that says whether the model does not exist, the key lacks access to
it, or the endpoint is wrong - OpenAI's own error responses say so
directly ("The model 'x' does not exist or you do not have access to it").

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:28:26 +02:00
Lucio Lelii 247de4485b Resolve a ProviderCredential everywhere a provider is called, not a string
LLMCredentialResolver.resolve now returns ProviderCredential instead of a
bare String, so the endpoint travels with the value from the vault all
the way to the provider that needs it. Every one of the eight call sites
had to change to compile - there was no way to touch only one - so all of
them now pass the resolved credential straight through instead of
unwrapping it first.

That turned out to be the right amount of change, not more than
necessary. Where an endpoint-aware provider is not actually reachable yet
(the interaction simulator and the bias judge choose their descriptor
after the execution already exists, and never had their authorization
requirement computed up front to begin with - a separate, pre-existing
gap this does not close), a missing credential fails exactly as it always
did: LLMCredentialResolver still throws "Missing saved credential" when
the authorizations map has no entry for the provider's key, whether the
caller then unwraps .value() or keeps the whole ProviderCredential makes
no difference to that failure. The only place behaviour actually changes
is the success case, and only for a provider that reads the endpoint at
all - every existing provider still only reads .value() through the
interface's own default unwrapping, so Gemini, InternalOllama and every
test stub keep behaving exactly as before.

LLMCredentialResolverTest is new: this resolver was previously exercised
only indirectly, through a full execution.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:28:12 +02:00
Lucio Lelii b32e5e7b12 Let a saved credential carry an endpoint, guarded against SSRF
A vault secret has always been one opaque value. A provider whose
endpoint is not known in advance - coming in the next commits - has to
get it from somewhere, and the credential the user already picks is the
one place that does not mean typing a URL on every run.

UserSecretEntity gains endpoint: nullable, unencrypted (it identifies
where to connect, not a secret, and needs to be readable to show in a
credential picker), populated only for a provider whose requiresEndpoint()
is true. UserSecretService enforces that at both create and update -
present when required, absent otherwise - and, when present, validates it
before it is ever stored.

That validation is OutboundEndpointGuard: http/https only, no credentials
embedded in the URL, and a rejection of loopback, link-local (including
169.254.169.254, the instance-metadata endpoint on every major cloud and
the single most valuable SSRF target there is), private and multicast
addresses, with an operator override for a legitimate private-network
endpoint. Deliberately not a reuse of HTTPServerCallService's existing
guard: that one resolves a hostname once and never again, which is
exactly the DNS-rebinding gap. This one is meant to be called immediately
before use as well as at save time, so the resolution it checks is the one
about to be connected to - though even then it does not pin the resolved
address for the connection that follows, so it is a baseline, not a
complete defence.

Saving is a courtesy: a comprehensible 400 instead of a mysterious
failure at execution time. The check that actually protects runs where
the endpoint is used, not here - that call site is not in this commit yet.

The provider catalog (LLMProviderMetadata, over the wire at
/llm/providers) grows the matching requiresEndpoint flag, which is what
will let the "Add credential" dialog show the field only where it applies.

Test literals worth a note: OutboundEndpointGuardTest resolves only
literal IP addresses and loopback names, never a real hostname, so it
needs no network access and cannot be flaky because of one. The new vault
tests use public IP literals (8.8.8.8 and 8.8.4.4) for the same reason -
subdomains of example.com mostly do not resolve at all, and a real DNS
name would make these tests depend on network access they should not need.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:27:43 +02:00
Lucio Lelii ddc792ff28 Let a provider declare it needs an endpoint, and accept one as a credential
Two additions to LLMProvider, both additive defaults so no existing
provider or test stub changes behaviour.

requiresEndpoint() names the one thing every provider until now has had
in common without anyone needing to say so: a base URL it already knows,
whether server-configured or a constant of the service it talks to. A
provider whose endpoint is not known until a credential names it -
coming next - is the first that needs to say otherwise.

ProviderCredential carries that endpoint alongside the secret value a
provider has always received. The three new generate/generateJson/chat
overloads that take one default to unwrapping .value() and calling the
String-authorization overload above them, so a provider that only
overrides the old ones - which today is every one of them, including
every anonymous test stub across the suite - keeps behaving exactly as it
did. Only a provider that overrides the new overloads directly gets to
read .endpoint() at all.

Adding an abstract method instead would have broken every one of those
stubs, since none of them implement anything beyond the three methods the
interface already requires.

One ambiguity fell out of this at the call site InternalOllamaLLMProvider
used to have: chat(model, messages, null, null) no longer resolves
unambiguously, since a bare null now fits both the String and the
ProviderCredential overload equally. Not visible in this diff - that call
site was rewritten away in the Ollama extraction - but worth naming since
it is the shape of thing this kind of overload addition can trigger
elsewhere too.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:27:19 +02:00
Lucio Lelii 68c3a1e4a5 Add a shared HTTP transport base for LLM providers
Every provider that speaks HTTP wrote its own version of the same handful
of things - a client to reuse, an error mapping, a blocking call with
logging - and each ended up with a different half of it. Gemini has a
retry policy and no logging at all; Ollama has logging but no retry; both
map only 5xx responses and drop the body of a 4xx, which is exactly the
detail that would say whether a model name is wrong, a key lacks access,
or an endpoint is misconfigured.

AbstractHttpLLMProvider consolidates that: a WebClient cache keyed by base
URL (needed once an endpoint can vary per call, which a user-supplied one
will), 4xx and 5xx both mapped through LLMProviderHttpException carrying
the response body, and failure logging that distinguishes an HTTP error
from a connection failure from a timeout.

It is deliberately transport only - no generate/chat/generateJson, no
request body, no parsing. Two hooks, timeout() and retryPolicy(), are
overridable rather than fixed, because Gemini authenticates through a
query parameter rather than a header and any future provider might too;
a base class that assumed otherwise would not be a base class Gemini could
actually sit on.

Not wired into any real provider yet - AbstractHttpLLMProviderTest proves
the plumbing generically, through a minimal test-only subclass and a real
JDK HttpServer, before anything depends on it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-17 12:27:00 +02:00
Lucio Lelii e7f3ee5474 Add the SPDX licence header to every source file
Mechanical: three comment lines above the package declaration of every
.java file under src, main and test alike, and nothing else. Its own
commit because it moves the blame line on 501 files and would otherwise
bury the licence change it belongs to.

The short SPDX form rather than the full GNU notice - machine-readable
under REUSE, sufficient to keep the licence notice intact, and it defers
the attribution term to LICENSE-ADDENDUM rather than repeating it five
hundred times.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 12:10:34 +02:00
Lucio Lelii 9b354b4286 Release under the AGPL with an attribution term
Same licence and the same additional term as the web repository: the two
halves are one product, and licensing them differently would leave the
question of what a derivative of the whole owes unanswerable.

The AGPL rather than the GPL because this service is meant to be hosted,
and section 13 is what obliges whoever hosts a modified copy to offer its
source to the people using it. The section 7(b) term in LICENSE-ADDENDUM
requires the attribution to be preserved, including in the Appropriate
Legal Notices a derivative displays.

The pom's licence, developer and scm blocks were the empty placeholders
Spring Initializr generates, so the published artifact declared no licence
at all - now they say what the LICENSE file says.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 12:10:26 +02:00
Lucio Lelii a13092dcea Log the whole completed operation, not just the field we read
Only result.result is modelled here, so anything else the bridge reports
is dropped at deserialization - which makes a bridge that returned its
answer under another name indistinguishable from one that produced
nothing. "No output generated" is exactly the case where that distinction
decides what to fix, and the bridge's schema is not in this codebase, so
the payload itself is the only way to tell.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 16:51:32 +02:00
Lucio Lelii 337b2939ac Stop reporting a client that went away as a server failure
The editor polls a running execution and drops the request when it
navigates, refreshes or supersedes it - routine, and more likely the
longer the response takes to write, which an execution view carrying a
large global input does. Each one was logged as "Unhandled request
failure" with a hundred-line stack trace, for something nobody can act on
and where the reply goes to a connection that is already gone.

Recognised through the cause chain, because Jackson wraps the broken pipe
several times over before it surfaces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 16:28:49 +02:00
Lucio Lelii 97f3ab0e66 Give a query with attachments the same step budget, within the documented ceiling
max_steps went only on the JSON query operation, so a query carrying an
attachment still ran on the bridge's default and could abort the same way.
The multipart operation takes it too.

The default drops from 200 to 100, which is the ceiling the bridge
documents: asking for more is at best ignored and at worst refused.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 15:10:56 +02:00
Lucio Lelii 4ce74e9863 Ask the bridge for more steps by the name it actually takes
The query kept aborting with "Recursion limit of 60 reached". That
message names LangGraph's own recursion_limit, which is what the bridge
sets internally - so that is the key that was sent, on the session and on
every query, and it changed nothing. The bridge's API takes the budget as
max_steps, on the query operation, which is worth more than any amount of
reasoning about its error text.

Sent there and nowhere else now: the budget belongs to a query, not to
the session that may run several.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 15:04:54 +02:00
Lucio Lelii 03db2dee5c Ask an upload row which source it is on, then show only that branch
The row offered both ways of giving a file at once and greyed out the one
not in use, which still leaves the reader working out which half is live.
It reads better as what it is: one choice, then the fields that choice
needs - an input name, or which global input holds the file.

Greying was all the schema could express, so this adds the annotation for
showing a field only in the state it belongs to. The distinction earns
the second annotation: a field that still tells the reader something
while unavailable should stay and grey, but the branch nobody picked is
not unavailable, it is irrelevant.

A row saved before the choice existed says which branch it is on by what
it carries, so it is stamped on read rather than left reading as the
default and demanding an input name it never had.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 14:46:01 +02:00
Lucio Lelii fed7141c06 Offer an upload row's two sources as the alternatives they are
Picking a global and nothing else still failed the block update: the kind
select posts "" when nobody chose from it, and an enum cannot be coerced
from that, so the request came back naming a field that is optional and
that the person had deliberately left alone. Blank now reads as unset.

The row also left both sources on offer at once, with nothing saying which
one wins. Filling either now greys out the other, and the input name is
required exactly when no global is chosen - which needed a way to say
"while this other field is empty", since a primitive boolean cannot tell
an unspecified present() from a false one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 14:37:39 +02:00
Lucio Lelii 665a05899d Ask an upload input only for what it alone can answer
Choosing a global input and nothing else broke the block update outright:
"multiple" was a primitive boolean, so a row posted before that box had
been touched failed to deserialize, and the editor showed a bad request
naming a field nobody had filled in. Absent now means single, which is
what a half-filled row means.

The other two fields were being asked for without earning it. A name is
the port's name, so an attachment taken from a global has none to give -
and declaring one grew a port that asked for the same document a second
time, once as a global and once as a step input. A kind is written before
any file exists, can be wrong by accident and can be wrong on purpose, so
what a file is now comes from its own bytes when the query is built; the
declaration only filters the picker, and only where there is one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 14:18:50 +02:00
Lucio Lelii 1d5d1eb78b Keep isFromGlobalInput out of the payload it is not part of
Jackson reads an is-prefixed no-arg boolean as a property, so the helper
added with the field put a "fromGlobalInput" into every block payload and
every persisted flow - a key the generated schema never declares, sitting
next to the one it is derived from. ModelParameters.isEmpty carries the
same guard for the same reason.

Also cover the two steps between this field and the code that reads it:
the generated schema has to carry it, and a configuration posted back by
the editor has to survive deserialization into a block whose ports reflect
the choice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 14:11:54 +02:00
Lucio Lelii f4693ac6ab Attach a global input's file, and name a file by its name in a prompt
Two halves of the same trap. An MCP agent handed ${{global.document}} got
this server's temp path - "/tmp/document1190411166785638981plan.pdf" -
and spent its turns trying to open a file it cannot reach, reporting it
could not access the plan. Meanwhile the only way to attach a file for
real was a port on the block, so a document needed by four agents had to
be uploaded four times.

An upload input can now name a global input to take its file from, chosen
in the editor from the flow's file-typed globals (the retriever pattern
sharedSessionRef already uses); named that way it grows no port, since a
global reaches a block by being named and a port for it would sit
unsatisfiable. And a file interpolated into a prompt now renders as its
name, which is also the name the bridge is told the attachment has - for
which the upload had to stop mangling it, so each one now lands in a temp
directory of its own under the name it arrived with.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:53:01 +02:00
Lucio Lelii 818f9c28e1 Stop re-deleting MCP sessions the bridge has already dropped
The bridge expires sessions on its own, so the idle sweep regularly asks
it to delete one that is already gone. That 404 was treated as a failure,
and the cost was not just the warning and stack trace every minute: the
removal from activeSessions sat after the call that threw, so the dead
session was never untracked and the sweep retried it forever - and each
attempt counted against the circuit breaker that guards real MCP calls,
where five of them open it.

A session the bridge no longer has is the outcome this method wants, so
404 now completes it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:34:42 +02:00
Lucio Lelii 4516105d74 Give a global input of a file kind somewhere to be uploaded to
The only endpoints for a global input took JSON, so an upload aimed at one
was refused before it reached any handler: "Content-Type
'multipart/form-data' is not supported". A flow whose global input is a
file - the plan document of the orchestrator flows, for one - could
therefore never be given its file at all.

Add the multipart pair the node inputs already had, named the same way
because it is the same operation on the other scope.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:31:43 +02:00
Lucio Lelii 01d739ee7d Stop a short input name from failing a file upload
The uploaded file's temp name used the input's own name as the prefix
handed to File.createTempFile, which rejects anything shorter than three
characters. Uploading to an input called "dc" therefore threw
IllegalArgumentException - past the IOException catch, so the client saw
only a 500 that reads as "Failed to upload file" in the UI, with nothing
naming the real cause. A filename carrying a path separator failed the
same way.

Sanitise and pad both parts: neither is the uploader's mistake to pay
for, and nothing downstream reads meaning out of the temp name beyond the
extension, which is preserved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 12:19:20 +02:00
Lucio Lelii cc820554ac Close shared MCP sessions when their execution is cancelled
ExecutionContext.cancel() cleared executionVariableDescriptors before
transitioning to CANCELLED, so ExecutionsService.cleanupManagedResourcesIfFinal
(triggered from the same state-change notification) always found an empty
map and closed nothing. The CLOSE_RESOURCE mechanism already worked
correctly on SUCCESS and ERROR - only cancel/stop silently leaked any
shared MCP bridge session still open at that point, until the bridge's
own idle timeout, eventually hitting its session cap ("Reached the
maximum limit of 5 sessions").

Defer clearing executionVariableDescriptors until after the state-change
notification runs, so the cleanup sees the still-registered CLOSE_RESOURCE
session descriptor and closes it before the map is scrubbed. Added a
regression test that reproduces the leak (confirmed red without this
change) and asserts closeSessionQuietly runs on cancel.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 11:39:01 +02:00
Lucio Lelii 44991f4a79 Raise the MCP bridge's agent-loop recursion limit for complex queries
A production run hit "Recursion limit of 60 reached without hitting a
stop condition" on initialize-persistent-orchestrator, which needs many
tool round-trips (inspect workspace, write/verify two files) in one
query - the bridge's own LangGraph agent loop aborted before finishing,
even though nothing on our side errored.

Send recursion_limit (configurable via app.mcp.bridge.recursion-limit,
default 200) on both the session-open request and each query operation,
same best-effort spirit as think: the bridge's request schema isn't in
this codebase, so this is sent on faith it's honoured somewhere - if it
isn't, nothing changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 10:55:54 +02:00
Lucio Lelii 9ad2394090 Keep a thinking model's reasoning trace out of block/query outputs
A reasoning model (e.g. qwen3) mixed its <think> preamble into the same
text this code treats as the final answer, since neither the direct Ollama
calls nor the MCP bridge session ever asked for reasoning to be kept
separate. That let a Conditional's SpEL condition (or any other consumer)
silently see reasoning prose instead of the expected value - in a
LoopContainer this meant looping through iterations without ever taking the
intended branch, with nothing logged to show why.

- Ask Ollama to think explicitly (think: true) on both the direct provider
  and the MCP bridge's llm_provider options, so reasoning is returned
  separately instead of folded into response/content - full reasoning
  quality kept, unlike think:false which would ask the model to reason
  less.
- Defensively strip a closed <think>/<thinking> block from an MCP query
  result, and fail loudly instead of returning garbage when the block is
  unterminated or leaves nothing behind - the bridge is an external service
  we don't control, so this is the fallback for whatever it sends anyway.
- Log every MCP query result's raw text to its own file
  (mcp-agent-responses.log), mirroring the existing assistant-responses
  log, since there was previously no way to see what a given model/bridge
  combination actually returns.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-12 13:27:12 +02:00
Lucio Lelii 3ccdeca637 Key coding-agent-mcp workspace by root execution id, not per-iteration id
A LoopContainer iteration spawns a fresh child ExecutionObject each time, so
context.executionId (used to key the coding-agent-mcp workspace subpath) was
never stable across iterations once MCPAgent nodes stopped sharing one MCP
session. Add context.rootExecutionId (the top-level execution an iteration
belongs to) and use it for the workspace subpath instead, so independently
sessioned MCPAgent nodes still land in the same workspace across iterations.

Also fixes a narrow, real race in JensenStructuredFlowsExecutionTest: a
step's in-memory status flip and its listener-triggered persistence happen
on the same background thread but aren't atomic, so evicting the execution
from cache and reloading it could occasionally observe a stale snapshot.
Retry the evict-and-reload instead of asserting on a single attempt.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-11 17:27:58 +02:00
Lucio Lelii 6f9a0d1929 Add MCP multipart file uploads 2026-09-11 14:50:16 +02:00
Lucio Lelii aa308fd70f Use asynchronous MCP query operations 2026-09-11 10:48:11 +02:00
Lucio Lelii 89126a2b9e Avoid retrying MCP session creation 2026-09-10 12:04:11 +02:00
Lucio Lelii 18f2ab2fa0 Let a secure retriever answer the questions the editor asks it
The editor asks every retriever-backed field whether its list is open,
and whether the field is required, on endpoints suffixed onto the
field's own URL. Only the unsecured retriever had them, so opening an
MCP agent's shared session picker asked /secure-retriever/.../open and
got a NoResourceFoundException stack in the log. The editor swallowed
the failure and fell back to a closed list, which is why nothing looked
wrong.

Both questions now exist on the secure side too, defaulting the same
way, so the answer comes from the retriever rather than from a 404.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 11:35:19 +02:00
Lucio Lelii 90c97f9aad Harden Maven downloads in Docker build 2026-09-10 11:34:39 +02:00
Lucio Lelii 760554c2d7 Fix shared MCP session ownership in subflows 2026-09-10 11:34:37 +02:00
Lucio Lelii a2743c6c66 Validate inherited subflow globals 2026-09-09 17:17:04 +02:00
Lucio Lelii ec009411e1 Fix shared MCP sessions in container subflows 2026-09-09 16:34:05 +02:00
Lucio Lelii 289bcfb1db Note how to run the service locally and build its image
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 12:24:51 +02:00
Lucio Lelii 44f722e102 Publish which of an enumerated field's values is the default
A field with a fixed set of values had no way to say which one it opens on, so the MCP server
dialog started with no source type chosen: nothing was selected, nothing said it had to be, and the
server could be saved that way. `@SchemaAllowedValues` now takes a `defaultValue`, emitted as the
schema's `default`, and both MCP configurations declare CATALOG - which is also the choice whose
own required fields the dialog can then enforce.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 12:23:42 +02:00
Lucio Lelii 8d6ddfecfb Replace Flyway schema setup with JPA mappings 2026-09-09 12:20:36 +02:00
Lucio Lelii e71ba9c290 Stop a second LLM assessment from erasing the one still running
Two assessments of the same report, started minutes apart, left only one:
observed in a real run as two REPORT_JUDGE jobs both COMPLETED (14:33-14:37
and 14:36-14:37) against a report that ended up holding a single judgement -
the one that finished last.

judgeReport was one @Transactional method that read the report, spent minutes
in one model call per compared pair, then wrote. The second assessment read
the report while the first was still calling models, saw no history, and
saved its own judgement as the only one there. Nothing warned, because from
each writer's side the write succeeded.

The model calls now happen outside any transaction - holding one open across
minutes also pins a connection for no reason - and the write moved to
BiasJudgementStore, a bean of its own so the transaction starts there and a
retry re-enters through the proxy rather than inside the transaction that just
failed. It re-reads the report as it is at that moment, puts the new verdicts
on those pairs and the summary on that history, and the report entity carries
a @Version so a writer working from a stale read is refused instead of
overwriting: the next attempt reads the assessment that landed meanwhile and
appends after it.

Optimistic rather than SELECT ... FOR UPDATE because the tests said so:
Hibernate renders PESSIMISTIC_WRITE as "for no key update", which H2 - what
the suite runs on - cannot parse. A fix that only holds on one dialect is not
one.

The regression test runs the two assessments on two threads, with a stub
provider that blocks inside the model call until both have reached it, and
asserts the report keeps both. Reinstating the old shape fails it.

This prevents further losses; it does not recover an assessment already lost.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 15:14:55 +02:00
Lucio Lelii 02b284d368 Keep every LLM assessment of a bias report, and fall back to text when JSON comes back empty
A report held one BiasJudgeSummary, so asking a second model overwrote the
first: there was no way to compare two opinions on the same comparison, and
a report survived only until someone reevaluated it. BiasImpactReport now
keeps a capped history (10) of assessments, newest first, each carrying its
own verdicts per compared pair rather than a single shared set - the point
of keeping several is being able to trust each one's own reasoning, not just
its headline. Recomputing an outdated report now carries the history forward
instead of discarding it.

schemaVersion had to become Integer rather than int in the same change:
Jackson reads an absent field as null, and a primitive rejected every report
persisted before today - a bug this history change would otherwise have
inherited silently, caught by a new compatibility test that reads an old
report's JSON.

Separately, the judge asked for its verdict in JSON mode only, and Ollama's
format=json makes some models - reasoning models especially - answer with an
empty body or a degenerate {}, since they have nowhere to put their thinking
under that flag. Every compared pair failed with "the model returned an
empty answer". The flow assistant has ripped this exact seam out before and
falls back to a text call; the judge now does the same, extracting the
verdict object from whatever prose it comes wrapped in.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-08 11:40:21 +02:00
Lucio Lelii 7d3f6b088c Resolve every configurable-as-input field through one rule
Three fields are declared @ConfigurableAsInput - the model of an LLM
descriptor and the two on the MCP blocks - and the editor offers the same
choices on all of them: leave it blank and feed the port, or write a template
such as ${{global.modelName}}. It was resolved three times and differently.
The MCP executors each carried an identical private copy that fell back to the
configured value raw, so a placeholder written there reached the MCP service
verbatim, and a global input could not decide the model of an MCP block at all.

The rule now lives in ConfigurableInputBinding: a value bound to the port
wins, otherwise the configured value is resolved as a template against the
same inputs and variables a prompt is. LLMDescriptorInputBinding keeps only
what is specific to it - rebuilding the record around the resolved model, and
resolving the model alone, since the provider decides the credential and the
sampling parameters and cannot arrive mid-execution.

The MCP model fields carry @AcceptsVariablePlaceholder to match, so the editor
declares what the runtime has always been asked to accept.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 10:52:52 +02:00
Lucio Lelii 8f9017891c Compare a bias variant subject by subject, and let an LLM assess it
A full-flow comparison measured one number: an edit distance between the
outputs of every activated node stringified into one map. On a flow that
evaluates one subject per iteration - the multi-CV case - that number was
dominated by the map's punctuation and could not say which subject moved,
by how much, or whether anything of substance had changed.

The comparison is now field by field, and list values are paired element by
element, so an iterated node reports per subject. Where both sides label a
number the same way, its movement is reported in the flow's own units
(Score 8 -> 5), which is the only figure here a reviewer can act on. Values
that only one side produced are marked rather than guessed at, and both the
element count and the edit distance are capped, the latter because it was
quadratic with no bound on exactly the long model outputs it runs on.

Iterator and loop containers are compared iteration by iteration by joining
the child executions on parentIterationIndex, which they already record.
Guard subflows are excluded: a loop creates one per main iteration, and
pairing a guard with a main run reported every loop as rewritten. This is
what turns "the accumulated list changed" into "iteration 3, on this inner
node".

None of that says whether the change is the intervention doing what its
probe described or the same model answering differently, so a report can now
be handed to a model: one call per aligned pair, answering on two separate
axes - how far the meaning moved, and whether the change carries the
intervention's fingerprint. Provider, model and sampling arrive as an
LLMDescriptor and are resolved the way the interaction simulator's are. The
level and the attribution shown are rolled up from the per-pair verdicts in
code; only the narrative comes from the model, so re-reading a report cannot
show a different headline than the pairs it is made of. A model that answers
with something unusable leaves an error on its pair and nothing else.

Finally, a rerun now carries the simulator of the run it repeats - the
descriptor only, never the enabled flag - and the report records what
answered the interactive steps on each side. Two runs answered by different
simulators differ for a reason the intervention had no part in, and until now
nothing said so.

Reports are persisted and served from storage, so they carry a schema version
and an older one is recomputed instead of serving its new sections empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 10:50:15 +02:00
Lucio Lelii d024f80239 Separate "takes a ${{...}} placeholder" from being drawn as a textarea
LongText.acceptVariableAsPlaceholder welded a fact about the value to a
rendering choice: only a field also drawn as a textarea could say it is
interpolated. That left LLMDescriptor.model unable to declare it - the
capability worked, since the executors resolve the field as a template,
but nothing in the schema said so and nothing in the editor showed it.

@AcceptsVariablePlaceholder is that fact on its own. LongText keeps its
flag as the shorthand for the many prompt fields that are both, and the
new metadata is applied after it so that LongText writing the same key as
false cannot win: either annotation saying yes is enough.

The model field also gains a description, so the editor can say what an
empty value and a placeholder each do rather than leaving both to be
guessed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 14:31:12 +02:00
Lucio Lelii 3c5b44e108 Let a node's LLM model come from an input instead of the configuration
In the four blocks that hold an LLMDescriptor - LLM, ChatInteraction,
Conditional and Switch - the model can now be decided at run time rather
than only at design time. Leaving the field blank turns it into a "model"
input port; writing ${{global.modelName}} into it resolves the value from
a global input, which is the only way one can reach a block since globals
travel through template interpolation and never feed a port.

The mechanism already existed: @ConfigurableAsInput, until now used only
on the flat model field of the MCP blocks. What was missing was reaching a
field one level down. configurableInputDescriptors now descends into the
objects a configuration holds, and the editor needed nothing at all - it
already reads the binding at a dotted path and exempts a bound field from
its required marker.

Three states, and only the middle one is new. No descriptor at all offers
no port: a freshly dropped node is a scaffold, and a port there would ask
for an input before a provider had even been chosen. A descriptor with the
model set needs no port. A descriptor whose model was left blank is asking
for one.

@NotBlank had to go from LLMDescriptor.model, because blank is the state
the binding is made of - the MCP equivalent does not carry it either. The
consequence is deliberate and changes a tested behaviour: a blank model
used to make the flow DRAFT, and now makes it EXECUTABLE. That is right,
because an unconnected port is simply one of the values the run asks for
before starting, exactly like the ${{...}} placeholders a prompt declares.
The field stays required in the published schema so the editor still marks
it.

The port is named "model", which a ${{model}} placeholder in the same
block's template could also claim. Deduplicating the two would be worse
than failing - the value wired for the prompt would silently become the
model - so that collision is refused with an error naming it.

LLMDescriptorInputBinding is the single place the rule lives: the port
wins over the template, a configured literal is returned unchanged rather
than copied, and an empty port never blanks out a configured model. The
provider is deliberately not bindable: it decides the credential, which
AuthorizationRequirementResolver has to resolve before the run starts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 14:24:01 +02:00
Lucio Lelii a1163adb46 Let hosted providers take a typed model, and cap temperature at 1.0
A provider whose catalogue cannot be listed without a credential now says
so, and the editor offers a free text field instead of a select. Gemini is
the first: its four hardcoded model names went stale as fast as Google
renamed them, and the whitelist guard rejected models that do exist.

canListModels() is a declared capability rather than an inference from an
empty list, because for our own Ollama an empty answer means it is
unreachable - not that anything goes. isOpen() answers it per field on a
sibling endpoint, the way isRequired() already does, so nothing needed a
new annotation or a new schema key.

Temperature now stops at 1.0, below what Gemini's API accepts: past that
the output is noise, and one range that holds for every provider beats a
per-provider ceiling nobody can remember.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 11:53:28 +02:00
Lucio Lelii 19f9fcdbe0 Mark a nested object as a group of optional settings
@UiOptionalGroup says what a field is - a group of settings that are all optional
- rather than how to draw it. The editor will render it as one control that opens
a dialog instead of unfolding five empty chips inline; the read-only execution
view will keep showing only what was actually set. Two renderings, one
declaration.

A field annotation, not a type one, and that is forced rather than preferred.
Shared definitions are hoisted into sharedDefinitions only when their JSON is
identical in every schema containing them, so a label that varies by owner would
un-share ModelParameters silently. The web merges a property's x-ui-* keys over
the definition it $refs, so the renderer sees the marker on the object node
anyway - the constraint costs nothing.

The test asserts the marker is on the property, that the $ref survives beside it,
that it is absent from the definition, and that ModelParameters is still shared.
Making the label vary by owner fails it.

Applied to LLMDescriptor.parameters and to both MCP configurations. Nothing reads
it yet; the editor comes next.

534 backend tests green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-04 12:24:34 +02:00