SILEX / Assurance
Simulated agents · illustrative data
YY
Continuous Production Assurance

Proof that your agents still do what they should — as everything keeps changing.

SILEX runs continuous Agent Verification & Validation on a live Enterprise World Model: verifying that agents follow their intended trajectories, and validating that the resulting business state — and the controls around it — still hold under today's reality.

82%
Weighted coverage
▲ +4 pts vs last calibration
29.4K
Entities understood
962 ontology types, typed & linked
94%
World-model fit
simulation vs observed agreement
8
Known blind spots
▼ 2 critical · named & addressable
01

The Correspondence Problem

Three objects that used to travel together come apart once an agent holds real authority.

Designed Intentprompts · blueprints · policy · permissions
↔execution drifts
from design
Executed Behaviormodel reasoning · memory · tool responses
↔outcome drifts
from intent
Enterprise Realitypermissions · other agents · business state

Correspondence cannot be verified without a model of reality; a useful model of reality must carry both semantics and causal structure.

02

Agent Verification & Validation

Verification asks whether execution conforms to spec. Validation asks whether the spec is still correct under reality.

Verification — conforms to spec?
Validation — spec still correct?
Outcome side

Actual state = expected state

Replay every production trajectory against its designed path.

1,240 trajectories replayed · 3 deviations flagged

Expected state still correct?

Has the environment moved the "good" state under the agent?

2 workflows' expected outcome invalidated by env change
Control side

Policy enforced as specified

Every control fires where the design says it should.

61 control types checked · enforced as specified

Failure Closure →

Does the control prevent the prohibited outcome, or only block one path?

5 prohibited outcomes · 4 latent paths still open to "Unverified Bank Change"

Blocking one path does not prove that the underlying failure has been closed.

03

Enterprise World Model — the evidence base

A self-evolving L1–L4 ontology over a runtime knowledge graph. 1506 types · 5244 relations · grounded in open standards.

L1 · GeneralReusable agent-security semantics
547 · 74%
L2 · Domain packsBusiness meaning per vertical
204 · 78%
L3 · Agentic-systemPlanner · memory · tools · MCP
187 · 60%
L4 · Runtime graphDeployed & observed now
24 · 81%
L4 nodes: 24 illustrative + 544 benchmark (public benchmark runs of named models in a research environment; not this enterprise's runtime, not in coverage)
D3FEND 213ATLAS 131ATT&CK 101UCO 72OWASP 25
Path grades →Observed in the graphLatent possible by typeDeclared from config
KNOWN BLIND SPOTS · 8 OPEN
CriticalProcurement vendor data not connected — bank changes inferred from tool calls, not observed at source.
CriticalLong-term memory unmodelled — 31 agents persist memory across runs; contents not in the graph.
SeriousRefund path below target — WF-021 at 68%; goodwill-credit branch untested.
SeriousHR workflows partly represented — 2 of 19 workflows registered · 68%.
WarningSub-agent delegation partly observed — inherited authority across hand-off is inferred.

Each blind spot names the class and constraint that generated it — a hypothesis with an address, not a mystery. That is what makes the model self-evolving.

How candidate routes are evaluated
observe → reconstruct → explore → compare fixes · incident I-1042
Illustrative counts
2,304candidate branches in the modelled set (input × agent × action × resource × control state)
→
2,117excluded by named laws, each exclusion checkable
→
187simulated executions of what remains
→
4 routesexamined to "Bank destination changed" · 3 remain open
Why branches were excluded (2,117)
Action not in any agent's permission scope1,204L1 Permission ⊄ Agent.scope
Resource type cannot lead to the prohibited outcome571L2 path law: Resource ↛ Payment
Violates a grant constraint (amount, expiry, scope)238L3 permission-scope constraints
An existing control stops it on every edge104L4 controls, e.g. PAY-042

Pruned ≠ sampled: every excluded branch names the law that excludes it, so each exclusion can be checked. Coverage is only as complete as the model's vocabulary and assumptions.

04

The autonomy flywheel — the effect of assurance

Assurance is not an insurance premium; it is what lets consequential authority move from human approval to the agent.

Assurance→Authority→Production actions→Runtime evidence→Better world model→Stronger validation↺
Before — human-gated
Vendor bank change → human approval on every request
→
After — failure closure validated
Agent handles changes up to $2,000; re-auth bound to identity + intent + amount

Certification is a snapshot; assurance has to be continuous. — SILEX raises the autonomy ceiling of enterprise AI.

05

Where this is today

The four lifecycle jobs, stated honestly. Views in this demo are illustrative product screens.

Enterprise Agentic Security

Which unsafe outcomes are still reachable across your agentic workflows, and what needs a decision.

Enterprise Security PostureIllustrative values
Critical open risks Critical
3
Unsafe outcomes still reachable ·
Defense confidence Formula TBD
78%
Across validated workflows ·
Workflows at risk
5
of 6 registered ·
Needing revalidation Changed
3
After environment changes ·

Domain Suites

Enterprise → domain → agents, workflows, controls

Workflow Coverage

Registered workflows vs known workflows ·

24
registered
9
known in the environment
18
deployed
67% workflow coverage ·

Policy Decisions

Validated recommendations waiting for a human decision

2 pending
More detailCollapsed by default

Recent material activity

Incidents, approvals, revalidations

›
Incident I-1042 observedVendor bank mutation blocked by PAY-042; alternative paths remain
Recommendation validatedBind approval to vendor + intent + mutation · ready for approval
Environment change detectedNew vendor tool connected · 3 workflows may be affected

Coverage confidence detail

Path coverage and scenario representation

Formula TBD›
Path coverage92%
Workflow coverage77%
Revalidation status21 / 24 valid
Opening Blueprint Studio…

Incident Queue

Post-deployment entry point: production incidents across registered workflows.

Open incidents
12
Last 30 days
Critical incidents
3
Unsafe outcome still reachable
Affected workflows
7
of 18 deployed
Residual failures
5
Blocked once, alternative paths open
5 shown
I-1042 · 18m agoCritical

Vendor bank mutation attempted

Prompt injection reached the Vendor Update Tool and was blocked by PAY-042.

Finance · Vendor MasterValidating
I-1038 · 2h agoHigh

Aggregate refund limit exceeded

Individually allowed compensation actions composed into a prohibited $1,200 outcome.

Support · Refund ResolutionInvestigating
I-1031 · 6h agoHigh

Delegated approval reused

A stale approval state authorized a purchase-order change after the approver scope changed.

Procurement · PO ApprovalValidation queued
I-1019 · 1d agoMedium

Rotation bypassed dependency check

A remediation agent rotated credentials before discovering a downstream production dependency.

IT · Account RemediationMechanism inferred
I-1007 · 3d agoMedium

Dormant entitlement restored

A mover workflow inherited an obsolete entitlement from historical role state.

Identity · Access ProvisioningMonitoring
I-0994 · 5d agoClosed

Discount policy conflict

Regional and strategic-account policies produced contradictory approval decisions.

Sales · Quote-to-ContractRe-validated
/ I-1042

Vendor bank mutation attempted

Finance · Vendor Master Update · WF-021 · Production

Critical
1Evidence
→
2Alternative Paths
→
3Candidates & Recommendation
→
4Decision
Incident evidence · what happenedI-1042 · Blocked tool mutation

A blocked prompt injection is evidence of a broader workflow-level failure.

Existing controls · what blocked itPAY-042

Stopped the observed execution only.

Failure mechanismIntent–approval binding failure

Approval is not bound to the original identity, intent and exact mutation.

Known attack / failure path
External emailObserved input
→
Finance AgentAgent
→
Vendor Update ToolAction
→
Bank destination changedBlocked by PAY-042

Blocked once. Is the unsafe outcome still reachable another way?

Technical detail · causal workflow graph (for security engineers)
Layered workflow graph
Incident evidence is an overlay on enterprise structure—not the first cause in the chain.
External artifactExternal emailUntrusted instruction sourceObserved · I-1042
AgentFinance AgentIdentity: svc-finance-agentObserved
Identity / authorityFinance ApproverDelegates bank-update authority
Causal mechanismIntent provenance lostApproval state becomes reusableInferred · 86%
Current controlPAY-042Pattern-based tool guard
ToolVendor Update ToolWRITE vendor.bank_accountObserved
Resource / stateVendor bank recordCritical financial resource
Alternative agentProcurement AgentInherited vendor authority
Business outcomeBank destination changedProhibited · $2.4M exposure
Candidate controlBind identity + intent + mutationdo(candidate control)
Observed evidence6 nodes · 5 transitions confirmed
Inferred mechanismApproval loses binding to verified intent
Counterfactual result3 of 187 executions remain reachable
Unsafe outcome: Bank destination changed
Observed exercised in your runtime graph · Latent asserted possible by type over your real agents, no trace yet · Declared from configuration, before deploy (appears in Blueprint Studio)
Route list (text)

Residual reachability. The same outcome is still reachable through 3 paths.

Residual paths

What stays reachable today

Recommended control

The intervention designed to close them

✓Bind approval to vendor + intent + mutationValidated in simulation · closes every open path aboveCloses all

Policy Change Proposal

Ranked candidates, tested in simulation — promoted through shadow and canary before production

✓ Validated
Current state
Defense confidence62%
Residual reachability21%
Critical paths open3
→
Recommended interventionBind approval to vendor + intent + mutation

Validated against the scenario family

→
Validated state
Defense confidence93%
Residual reachability2%
Critical paths open0
97%legitimate workflows still complete
+420 msadded latency
2 stepsrequire re-authentication
Why, in type terms: the block covered one route, but the failure is authorization reuse through lost intent provenance — an approval state that stays valid across executions. The latent paths reach the same prohibited outcome via a different tool, a delegated agent, or a poisoned document. Binding approval to vendor + intent + mutation closes the class.
Validation stageSimulation→Shadow→Canary→ProductionNo policy reaches production from simulation alone. Rollback preserved at every stage.

Recommendation ready. A human decides.

RecommendationBind approval to vendor + intent + mutation
Validated stateDefense 93% · Reachability 2% · 0 critical paths
On approval, SILEX opens a pull request against your enforcement surface (IAM / gateway / policy-as-code) — never inline, never holding the block decision. Rollback covers agent-side state; business effects get a compensating-action recommendation run through your systems.

Approve accepts the validated recommendation. Modify changes it and re-runs validation. Reject returns to the other candidates.

Workflow Library

The enterprise inventory and system of record for registered agentic workflows.

Registered
6
shown in this demo
Deployed
5
active in production
Registered · not deployed
1
pre-production
Needs revalidation
2
after environment changes
Lifecycle: Draft→Confirmed→Validating→Ready for Approval→Approved→Registered→Deployed→Needs Revalidation
WorkflowDomainBusiness ownerSecurity ownerAgentsToolsResourcesControlsPoliciesLifecycleValidationLast validationDeploymentRiskCoverage conf.

Registered from Blueprint Studio

This browser · not deployed · computed by Blueprint Studio (simulated)

/ WF-021

Finance · Vendor Master Update

Vendor onboarding and bank-detail maintenance.

Deployed
DomainFinance
StatusDeployed
Business ownerFinance Operations
Security ownerM. Chen
Last validationToday 11:12
Defense confidence78%

Workflow graph

Illustrative fixture — it has no Blueprint document.

Security posture

Illustrative values

Open risks3
Residual reachability21%
Coverage confidence86%
Approved policies4

Activity

Latest changes to this workflow

Detail

Expand as needed

PCP · Policy Review

Runtime stage — Policy Change Proposals (PCP), validated and waiting for a human decision.

Human approval required
Used by security operators. Each recommendation stays connected to its incident or blueprint, workflow, failure path, control and validation result. Decisions are made in that context.
Finance
Ready for approval

Bind approval to vendor + intent + mutation

Finance · Vendor Master Update · incident I-1042 · closes 3 alternative paths

62→93%DEFENSE
21→2%REACHABILITY
+420msLATENCY
Customer Service
Ready for approval

Bind control to cumulative outcome

Customer Support · Refund Resolution · incident I-1038 · aggregate value capped

58→90%DEFENSE
26→4%REACHABILITY
+310msLATENCY

Pre-release · Workflow Validation ?

Release gate — workflow-wise. Is the workflow still safe after the enterprise changes, before release?

Environment change detected

3 workflows may be affected. Security validity decays as the enterprise changes.

3 changes

Affected workflows

Revalidation replays validation with the change applied

WorkflowTriggerLast validationStatus

Model change: Finance Agent model update

A second kind of release-gate change. Same prompts, tools, permissions and task suite; only the model differs.

Illustrative model comparison, not recorded runs
OUTPUT CHECKS · AP output check suite v3 · 8 checks
CheckCurrent modelUpdated model
Invoice amount matches POPASSPASS
Vendor name matches vendor masterPASSPASS
No duplicate paymentPASSPASS
Approval present above $10,000PASSPASS
Ledger entry balancedPASSPASS
Currency and remittance formatPASSPASS
Remittance notice sentPASSPASS
Response format and tonePASSPASS
Score8 / 88 / 8
GRAPH DIFF · execution routes exercised by the suite
Finance → Vendor Tool → approved mutationboth models
Finance → Procurement → PO lookupboth models
Latent Finance → Support → vendor merge → inherited authorityupdated model only

Every output check stays green, yet the updated model exercises Path C of incident I-1042, a route the world model had already flagged by type. The model change granted no new authority; it used authority that was already there. Illustrative: Path C remains latent in production traces.

Runtime Observation ?

Action-wise. Is each agent action safe to run, decided before it runs?

Before an AI agent's action runs, Silex checks it in three steps: fixed rules, then a judge model (Jev), then a policy. The gateway then lets it run, holds it for a person, or blocks it.

Simulated judge · fictional tenant
Scripted scenarios
—
AP payments + SOC triage
Actions checked
—
Tool calls, before they run
Stopped before running
—
Held or blocked at the gateway
Decided by a hard rule
—
No model can override these

Watch one agent action get checked

Pick a situation an agent might face. Each action the agent wants to take travels left to right through the checks; where it stops tells you who decided and why. Until you pick one, the picture loops through three example payments.

Simulated engine
Ready
Choose one of the 12 scripted situations; the chip shows its expected outcome under the default policy.

How to read the picture

Agent actionThe tool call the agent wants to make, such as payments.execute.
Hard rulesFixed checks, such as an amount limit or a domain allowlist. Any one can stop the action, and no model can override it.
Judge (Jev)A small model answers risk questions: is this off the user's goal? Is the payee suspicious?
PolicyTurns the judge's answers into allow, hold or block.
EvidenceA record of every decision, kept for audit. In this preview nothing leaves the browser.

Tip: click a stage to pause the animation there; click it again to resume. The outcome chips are the default-policy reference outcome.

System Validation ?

System-wise. Is the broader enterprise environment still safe as data, agents, workflows and threats evolve?

Last system validation
Aug 31
Monthly cadence
Workflows in scope
24
All registered workflows
Environment changes since
+3 / +5
Workflows / agents
Revalidation status
21 / 24
Workflows still valid

What system validation does

Visible and manually triggered in V1

Not started
1 · Recalibrating environment
Enterprise World Model
2 · Updating workflow dependencies
24 workflows
3 · Generating adversarial scenarios
incl. historical incident replay
4 · Running attack mutations
red teaming · backtesting
5 · Testing policy resilience
approved policies
6 · Updating residual risk
residual reachability
System validation complete. 2,700 scenarios · 3 unsafe outcomes still reachable · 1 policy weakened by the new vendor tool.

Backtest history

Same scenario family, re-run as agents and workflows are added

PeriodEnvironmentScenariosUnsafe outcomes reachablePolicies weakened
Jul18 workflows · 41 agents2,14062
Aug21 workflows · 48 agents2,42041

Enterprise World Model

Make the invisible legible.

Reusable semantics become enterprise-specific instances — follow the four ontology tiers.

Existing data · illustrative figures · not a live system

Enterprise Domain Suites

Navigate the enterprise by business domain, then drill into capabilities, workflows, agentic systems, and runtime evidence.

5 domain packs · 53 agents · 179 workflows
Enterprise›Finance›Business capabilities

Finance Domain Suite

Finance-specific entities, permissions, controls, business constraints, and prohibited outcomes mapped to the agent systems that execute them.

Accounts PayableInvoice Processing · Vendor Management · Payment Approval3 open incidents
TreasuryCash Management · Liquidity Forecasting · Bank Operations1 open incident
ControllershipJournal Entry · Close · Reconciliation0 open incidents

Selected Domain Context

Organizational ownership is metadata; the domain ontology is organized around reusable business meaning.

Domain ontologyFinance Security Pack v2.6
Primary ownerCFO · Finance Operations
Agentic systemsAP Agent · Treasury Agent
Runtime scope6.8K entities · 21.4K edges
Each suite is a navigational lens over the same federated ontology and runtime graph—not a separate security data silo.
Loading the coverage model…
About World Model Coverage · sources and caveats
World Model Coverage: how much of the relevant enterprise environment SILEX currently understands. Drill from enterprise into domain, capability and workflow — the six-dimension view re-reads at whichever level you are standing on. Formula TBD; every figure is an authored demo value, not a computed metric.
Illustrative model-component summary · authored figures · not the ontology tiers or the world model's layers

A legacy summary kept for reference. Its names are not a canonical layer taxonomy, and every percentage below is an authored demo value.Illustrative

O
Federated Ontology GraphGeneral semantics plus domain and agent-system overlays
142 types
G
Runtime Security Knowledge GraphPopulated identities, agents, tools, resources, controls, incidents, and outcomes
84%
C
Causal FoundationMechanisms, evidence strength, counterfactuals, and interventions
79%
H
Business HarnessObjectives, legitimate completion, prohibited outcomes, cost, and constraints
82%
S
Simulation & CalibrationBehavior agreement, outcome agreement, drift, and reality validation
94%

Coverage Gaps

How much of the failure and attack surface the policies in place cover. A separate reading from World Model Coverage, not a consequence of it.

8 open gaps2 critical4 serious2 warningEach names the class + constraint that generated it — a hypothesis with an address.
CriticalProcurement vendor data not connectedVendor master changes are inferred from tool calls, not observed at source · enterprise / procurement
CriticalLong-term memory contents unmodelled31 agents persist memory across runs; what is stored there is not in the graph
SeriousRefund path coverage below targetWF-021 · Customer Refund (SWM fixture) at 68% · goodwill-credit branch and partial-failure retries not yet exercised
SeriousPayroll adjustment mostly unobservedWF-043 at 60% · compensation records read through an unmonitored channel
Instrument
SeriousHR workflows only partly represented6 agents known, 2 of 19 workflows registered · coverage 68%
Serious3 known workflows not yet registeredFound in the environment but missing from the Workflow Library
WarningSub-agent delegation partly observed19 delegations seen; authority inherited across the hand-off is inferred, not confirmed
Later
WarningHorizontal agents unassigned to a domainFile-reading and analytics agents used by every domain · taxonomy for later versions
Later
Loading the ontology graph…
About the Ontology Graph · sources, typed relationships and caveats
The published L1 classes are distilled from public ontologies — MITRE D3FEND, ATT&CK, ATLAS, UCO and the OWASP GenAI lists — and keep their real identifiers. L1 also holds Silex-authored core concepts, and those, like everything else Silex adds, carry a review grade in the inspector: curated (core concepts, domain packs, components, hazards and mappings), heuristic (keyword-mapped threat placements) or illustrative (registered workflows, the L4 runtime graph and all coverage figures). CRM and Legal are candidate packs, outside the coverage figures. Final naming is still pending technical alignment.
Typed relationshipFromToMeaning
DELEGATES_AUTHORITYIdentityAgentWho empowered the agent and under what scope
CALLS / MUTATESAgent / ToolTool / ResourceWhich behavior caused an enterprise state transition
GOVERNS / GATESPolicy / ControlActionWhere enforcement changes the execution
THREATENS / COUNTERSThreat / CountermeasureAgentic componentWhich published attack lands on which part of the system, and what answers it
REACHESFailure / PathOutcomeWhy a prohibited business outcome remains possible
SUBCLASS_OFConceptMore general conceptThe only subsumption relation; tiers and groups are not
PART_OF_DOMAIN / DEPLOYED_INEntity, action, hazard / ComponentDomain packMembership and deployment, never written as specialisation
HAZARD_FOR / MAY_LEAD_TOHazardEntity or action / Prohibited outcomeThe condition that makes a business action dangerous, and where it ends
CHARACTERIZES / MITIGATED_BYHazardPublished threat / ControlWhich attack the hazard is an instance of, and what answers it
REQUIRES_EVIDENCE / RECORDED_BYHazard / Evidence typeEvidence type / Record schemaWhat must be observed to check the hazard, and which telemetry records it

Public sources: MITRE D3FEND, MITRE ATT&CK, MITRE ATLAS, UCO, OWASP GenAI Security Project. Distillation and licences: swm/data/SOURCES.md.

The Network view's notation is adapted from VOWL (Lohmann et al.) and its interaction model is inspired by WebVOWL (MIT); it is a separate implementation, not WebVOWL. Circles are ontology classes (L1–L3) or illustrative runtime instances (L4). The bundle has no datatype properties or set operators, so VOWL's datatype boxes and set-operator nodes are not drawn; export and force-distance options are omitted.

Loading the ontology chain…
About Ontology Layers · the tiers, the principle and caveats
Final naming is still pending technical alignment. The four tiers, L1 → L2 → L3 → L4, are presentation groups, not taxonomic ranks: only SUBCLASS_OF asserts that one concept is a kind of another. L1–L3 supply the world model's Schema; L4 supplies World State instances. This view does not depict the Laws, Objectives or Calibration layers. The build refuses to publish a bundle where a relation breaks its predicate signature, where the display tree or the subclass graph has a cycle, or where a node hangs under a lower tier.
Architecture principle: one federated ontology carries every domain pack; a domain pack is a root of its own whose entities and actions are kinds of general L1 concepts; an agentic component is a kind of a general L1 concept, deployed in the domains where it runs; and the runtime graph is the one place where those semantics meet observed behaviour. Membership, deployment and grouping are navigation, not subsumption. Organisational ownership stays a navigational attribute.

Layer selection is shared with the Ontology Graph tab: pick a layer here and the explorer opens there.

Planned for V2 and hidden in this demo: Runtime Knowledge Graph, Cross-Domain Risk, Business Harness, Model Health, rich ontology exploration.
Ready