Vendor bank mutation attempted
Prompt injection reached the Vendor Update Tool and was blocked by PAY-042.
SILEX runs continuous Agent Verification & Validation on a live Enterprise World Model: verifying that agents follow their intended trajectories, and validating that the resulting business state — and the controls around it — still hold under today's reality.
Three objects that used to travel together come apart once an agent holds real authority.
Correspondence cannot be verified without a model of reality; a useful model of reality must carry both semantics and causal structure.
Verification asks whether execution conforms to spec. Validation asks whether the spec is still correct under reality.
Replay every production trajectory against its designed path.
Has the environment moved the "good" state under the agent?
Every control fires where the design says it should.
Does the control prevent the prohibited outcome, or only block one path?
Blocking one path does not prove that the underlying failure has been closed.
A self-evolving L1–L4 ontology over a runtime knowledge graph. 1506 types · 5244 relations · grounded in open standards.
See how the I-1042 picture was reconstructed →
Each blind spot names the class and constraint that generated it — a hypothesis with an address, not a mystery. That is what makes the model self-evolving.
Pruned ≠ sampled: every excluded branch names the law that excludes it, so each exclusion can be checked. Coverage is only as complete as the model's vocabulary and assumptions.
Assurance is not an insurance premium; it is what lets consequential authority move from human approval to the agent.
Certification is a snapshot; assurance has to be continuous. — SILEX raises the autonomy ceiling of enterprise AI.
The four lifecycle jobs, stated honestly. Views in this demo are illustrative product screens.
Which unsafe outcomes are still reachable across your agentic workflows, and what needs a decision.
Enterprise → domain → agents, workflows, controls
Registered workflows vs known workflows ·
Validated recommendations waiting for a human decision
Incidents, approvals, revalidations
Path coverage and scenario representation
Post-deployment entry point: production incidents across registered workflows.
Prompt injection reached the Vendor Update Tool and was blocked by PAY-042.
Individually allowed compensation actions composed into a prohibited $1,200 outcome.
A stale approval state authorized a purchase-order change after the approver scope changed.
A remediation agent rotated credentials before discovering a downstream production dependency.
A mover workflow inherited an obsolete entitlement from historical role state.
Regional and strategic-account policies produced contradictory approval decisions.
Finance · Vendor Master Update · WF-021 · Production
A blocked prompt injection is evidence of a broader workflow-level failure.
Stopped the observed execution only.
Approval is not bound to the original identity, intent and exact mutation.
Declared and observed inputs are separated from inferred structure.
U1 · Finance Agent long-term memory contents at instruction time — no ontology class yet.
U2 · Approver's out-of-band bank-change callback (phone) — it is in no system, so it cannot be inferred.
Not connected: Vendor master S5 is not observable at source; the record state stays declared until the ERP is connected or the change is verified.
Reading the grades. Absent means the tier has no input; latent is inferred structure, never an unobserved state. Unmodeled means the graph cannot represent it. LEADS_TO: a declared policy or law reflects a customer-declared constraint; this view does not model Laws.
What stays reachable today
The intervention designed to close them
Ranked candidates, tested in simulation — promoted through shadow and canary before production
Validated against the scenario family
The enterprise inventory and system of record for registered agentic workflows.
| Workflow | Domain | Business owner | Security owner | Agents | Tools | Resources | Controls | Policies | Lifecycle | Validation | Last validation | Deployment | Risk | Coverage conf. |
|---|
This browser · not deployed · computed by Blueprint Studio (simulated)
Vendor onboarding and bank-detail maintenance.
Illustrative fixture — it has no Blueprint document.
Illustrative values
Latest changes to this workflow
Expand as needed
Runtime stage — Policy Change Proposals (PCP), validated and waiting for a human decision.
Finance · Vendor Master Update · incident I-1042 · closes 3 alternative paths
Customer Support · Refund Resolution · incident I-1038 · aggregate value capped
Release gate — workflow-wise. Is the workflow still safe after the enterprise changes, before release?
3 workflows may be affected. Security validity decays as the enterprise changes.
Revalidation replays validation with the change applied
| Workflow | Trigger | Last validation | Status |
|---|
A second kind of release-gate change. Same prompts, tools, permissions and task suite; only the model differs.
| Check | Current model | Updated model |
|---|---|---|
| Invoice amount matches PO | PASS | PASS |
| Vendor name matches vendor master | PASS | PASS |
| No duplicate payment | PASS | PASS |
| Approval present above $10,000 | PASS | PASS |
| Ledger entry balanced | PASS | PASS |
| Currency and remittance format | PASS | PASS |
| Remittance notice sent | PASS | PASS |
| Response format and tone | PASS | PASS |
| Score | 8 / 8 | 8 / 8 |
Every output check stays green, yet the updated model exercises Path C of incident I-1042, a route the world model had already flagged by type. The model change granted no new authority; it used authority that was already there. Illustrative: Path C remains latent in production traces.
Action-wise. Is each agent action safe to run, decided before it runs?
Before an AI agent's action runs, Silex checks it in three steps: fixed rules, then a judge model (Jev), then a policy. The gateway then lets it run, holds it for a person, or blocks it.
Simulated judge · fictional tenant iThe judge, latencies and tenant are simulated.Pick a situation an agent might face. Each action the agent wants to take travels left to right through the checks; where it stops tells you who decided and why. Until you pick one, the picture loops through three example payments.
payments.execute.Tip: click a stage to pause the animation there; click it again to resume. The outcome chips are the default-policy reference outcome.
The full Jev runtime demo, embedded. Things to try:
When the judge is unsure, it holds the action for a person. Every reviewer answer can become a training label, so the judge improves over time. Below, we tested the idea by fine-tuning our small judge, Kev, on a public benchmark (AgentDojo held-out). The labels come from the benchmark, not yet from customer reviewers.
Loading measured benchmark evidence…
An ontology labels what each step of a run means: “this call changes data”, “this argument is a password”, “this value came from a document, not from the user”. We asked whether those labels make a simple attack alert less noisy without missing attacks, on the same runtime provenance graph with and without ontology types. The rules were written down before looking at the data (pre-registered) and tested twice, on real benchmark runs of held-out agent models. No judge model is used here.
Loading the ontology results…
Every check happens while the agent waits, so the judge has a budget of 400 ms per action. We timed our small local judge, Kev, against gpt-4o-mini through the OpenAI API: same items, same questions.
Loading the latency measurements…
System-wise. Is the broader enterprise environment still safe as data, agents, workflows and threats evolve?
Visible and manually triggered in V1
Same scenario family, re-run as agents and workflows are added
| Period | Environment | Scenarios | Unsafe outcomes reachable | Policies weakened |
|---|---|---|---|---|
| Jul | 18 workflows · 41 agents | 2,140 | 6 | 2 |
| Aug | 21 workflows · 48 agents | 2,420 | 4 | 1 |
Enterprise World Model
Reusable semantics become enterprise-specific instances — follow the four ontology tiers.
Navigate the enterprise by business domain, then drill into capabilities, workflows, agentic systems, and runtime evidence.
Finance-specific entities, permissions, controls, business constraints, and prohibited outcomes mapped to the agent systems that execute them.
Organizational ownership is metadata; the domain ontology is organized around reusable business meaning.
A legacy summary kept for reference. Its names are not a canonical layer taxonomy, and every percentage below is an authored demo value.Illustrative
How much of the failure and attack surface the policies in place cover. A separate reading from World Model Coverage, not a consequence of it.
| Typed relationship | From | To | Meaning |
|---|---|---|---|
| DELEGATES_AUTHORITY | Identity | Agent | Who empowered the agent and under what scope |
| CALLS / MUTATES | Agent / Tool | Tool / Resource | Which behavior caused an enterprise state transition |
| GOVERNS / GATES | Policy / Control | Action | Where enforcement changes the execution |
| THREATENS / COUNTERS | Threat / Countermeasure | Agentic component | Which published attack lands on which part of the system, and what answers it |
| REACHES | Failure / Path | Outcome | Why a prohibited business outcome remains possible |
| SUBCLASS_OF | Concept | More general concept | The only subsumption relation; tiers and groups are not |
| PART_OF_DOMAIN / DEPLOYED_IN | Entity, action, hazard / Component | Domain pack | Membership and deployment, never written as specialisation |
| HAZARD_FOR / MAY_LEAD_TO | Hazard | Entity or action / Prohibited outcome | The condition that makes a business action dangerous, and where it ends |
| CHARACTERIZES / MITIGATED_BY | Hazard | Published threat / Control | Which attack the hazard is an instance of, and what answers it |
| REQUIRES_EVIDENCE / RECORDED_BY | Hazard / Evidence type | Evidence type / Record schema | What must be observed to check the hazard, and which telemetry records it |
Public sources: MITRE D3FEND, MITRE ATT&CK, MITRE ATLAS, UCO, OWASP GenAI Security Project. Distillation and licences: swm/data/SOURCES.md.
The Network view's notation is adapted from VOWL (Lohmann et al.) and its interaction model is inspired by WebVOWL (MIT); it is a separate implementation, not WebVOWL. Circles are ontology classes (L1–L3) or illustrative runtime instances (L4). The bundle has no datatype properties or set operators, so VOWL's datatype boxes and set-operator nodes are not drawn; export and force-distance options are omitted.
SUBCLASS_OF asserts that one concept is a kind of another. L1–L3 supply the world model's Schema; L4 supplies World State instances. This view does not depict the Laws, Objectives or Calibration layers. The build refuses to publish a bundle where a relation breaks its predicate signature, where the display tree or the subclass graph has a cycle, or where a node hangs under a lower tier.Layer selection is shared with the Ontology Graph tab: pick a layer here and the explorer opens there.
This recommendation was already validated in simulation. Approving records a human decision; it does not deploy anything.
Product terminology from the V1 PRD. Metric formulas are still TBD; values on these pages are illustrative — except Blueprint Studio, which computes its results on the declared graph (simulated outcomes, declared adversary model).
A pre-deployment, user-confirmed representation of an intended agentic workflow, used for simulation and security validation.
Testing whether unsafe outcomes remain reachable within a workflow or the enterprise environment.
A different sequence of actions that can reach the same unsafe outcome after another path has been blocked.
The remaining ability to reach an unsafe outcome after one or more controls have been applied. Used instead of "residual risk", which is too broad. Formula TBD
Confidence that relevant attack or failure paths have been adequately evaluated and mitigated. Formula TBD
Confidence that relevant workflows, paths, scenarios and environment components are sufficiently represented in validation. Formula TBD
A policy candidate already tested in simulation and recommended on security improvement and business impact.
Registered: formally accepted into the SILEX workflow inventory, possibly still pre-production. Deployed: active in production, so telemetry can be linked to it.
How much of the enterprise agentic workflow inventory is represented.
How much of the relevant enterprise environment SILEX currently understands.
Workflow-wise and change-triggered: a workflow is revalidated when the agents, permissions, policies or dependencies it relies on change.
Watch each agent action as it happens and see what runtime validation decides before it runs (hard rules → judgment → policy). Shown with a simulated engine.
System-wise and periodic: environment-wide recalibration, backtesting, red teaming and adversarial validation, on a cadence the enterprise defines.
Changed elements are outlined on the page. Click a row to jump there.