Trust & evidence

Trust is not asserted. It is earned through evidence.

OperatorWorks is designed to prove five things before a workflow is treated as reviewer-ready: correctness, precision, determinism, reproducibility, and claims safety.

Assurance doctrineNo computation without assumptions. No result without evidence. No release without replay. No claim without proof.
Fail-closed outcomes

A trustworthy system does not always say yes.

When a claim cannot be established within the implemented method and stated assumptions, the safe result is visible uncertainty—not manufactured confidence.

PASS

Established within scope

The claim is supported under the recorded assumptions, method, version, and support boundary.

FAIL

Contradicted within scope

The claim does not hold under the recorded assumptions and supported method.

Inconclusive

The available method or evidence does not establish either pass or fail.

Unsupported

The requested case falls outside the current implemented and evidenced boundary.

Assurance layers

Validation moves from static integrity to real-user proof.

The testing program is layered so that individual rules, system interactions, notebooks, installers, evidence artifacts, benchmarks, and external claims are evaluated separately and together.

L0–L1

Static and unit integrity

Imports, schemas, dependencies, claims lint, individual algebraic rules, canonical zero behavior, and typed error paths.

L2–L3

Property and regression tests

Linearity, symmetry, antisymmetry, idempotence, invariant preservation, golden identities, and prior release behavior.

L4

Cross-product integration

Derive → verify → certify → bundle → replay across Engine, Workbench, Verify, DataGen, and Evidence Bundles.

L5–L6

Notebook and installer replay

Fresh-kernel notebook execution, clean install, app launch, Jupyter startup, kernel registration, repair, and uninstall paths.

L7–L8

Evidence and benchmark validation

Manifest completeness, hashes, replay instructions, environment capture, fair comparisons, limitations, and reproducibility.

L9

Qualified user validation

Reviewer comprehension, scientific utility, repeat use, design-partner evidence, and commercial pull before expansion claims.

Release gates

Reviewer-ready means every gate has evidence.

Correctness, replayability, evidence completeness, and claims safety are not waivable polish items. They are the product.

Requirements gate

Each sanctioned requirement has an owner, acceptance test, evidence artifact, dependency, and release status.

Traceable
Correctness gate

Core identities, invariants, canonicalization, rewrite convergence, and negative cases pass deterministic checks.

Blocking
Fail-closed gate

Ambiguous and unsupported cases return explicit outcomes rather than silent success.

Blocking
Notebook gate

Shipped notebooks execute start-to-finish from a clean kernel without hidden local assumptions or manual repair.

Blocking
Installer gate

Application, disk image, managed runtime, Jupyter path, repair, and clean-install behavior are validated independently.

Blocking
Evidence gate

The build emits manifests, hashes, environment context, assumptions, warnings, logs, certificates, and replay instructions.

Required
Claims gate

Website, documentation, demos, release notes, and commercial materials remain bounded by what the build proves.

Zero drift
Reviewer gate

A technically qualified reviewer can reproduce the core artifact without undocumented private knowledge.

Final gate
Validation evidence pack

Every serious build should leave a reviewable record.

The evidence pack connects requirements, tests, installers, notebooks, benchmarks, claims, and the release decision.

01

Requirements traceability matrix — requirement → acceptance test → evidence → status.

02

Test and notebook reports — passes, failures, skips, replay status, and exact failure points.

03

Installer and environment reports — supported paths, dependency locks, kernels, clean launch, repair, and package inventory.

04

Correctness, Verify, DataGen, and benchmark reports — product-specific quality and limitations.

05

Claims lint and release decision — what can be said externally and the evidence-backed ship/no-ship rationale.

Local data boundary

  • Supported Workbench workflows remain on the reviewer-controlled machine.
  • Public forms should receive workflow categories, not confidential equations or datasets.
  • Selected artifacts can be redacted and intentionally exported for review.
  • Any future hosted service requires a separate security, privacy, governance, and disclosure model.

Claims governance

  • No production, cloud, API, customer, revenue, or superiority claim without direct evidence.
  • Current, controlled-review, sample-supported, and roadmap states remain visibly distinct.
  • Benchmarks disclose method, environment, comparison basis, and limitations.
  • Uncertainty is preserved rather than edited out for marketing convenience.
Technical evaluation

Judge the artifact, not the adjective.

Open the review kit
Current posture: evidence-gated review of local and artifact-producing workflows where supported. This page describes the assurance doctrine and required evidence standard; it does not imply that every roadmap product, domain, hosted path, or external benchmark claim is currently implemented or validated.