1. Executive Summary & Epistemic Scope#
This research brief establishes the empirical evaluation framework for Composable Intelligence Instruments (CII) and articulates the formal benchmark methodology for testing governed instrument composition against constituent instruments and naive tool chaining.
Under WinMedia's evidentiary taxonomy, this document is classified as a $C_1$ Conceptual Specification & $C_4$ Benchmark Research Program. It specifies the exact experimental protocols, boundary variables, and baseline comparison models required before empirical outcomes can be ratified as a $C_2$ Demonstrated Experimental Result.
Evidence classification: C1 (Conceptual Protocol) / C4 (Benchmark Program)
Scope: Formal specification of treatment conditions, evaluation metrics, and invariant verification protocols
Preregistration / Protocol Reference: WMR-PR-CII-001 (Planned)
Generalization beyond specified protocols: not established
2. Experimental Architecture & Benchmark Protocol#
2.1 Candidate Instrument Interface#
The benchmark protocol evaluates the composition of two distinct instrument classes across an explicit typed interface:
- Structural Problem Explorer (Design Space Calculus): Operates over typed problem-space decompositions using formal constraint propagation.
- Invariant Verification & Discovery Evaluator (Novum): Operates over candidate state transitions, evaluating structural compliance, boundary conditions, and invariant preservation.
2.2 Task Family: Multi-Constraint Structural Synthesis#
The reference benchmark family focuses on multi-constraint structural synthesis tasks where solutions require both exploratory state-space search and strict invariant satisfaction across hierarchical boundaries.
2.3 Required Experimental Conditions & Baselines#
Every empirical evaluation under this program must benchmark across six explicit treatment arms:
C0Uninstrumented Baseline: Prompted end-to-end task generation without externalized instrument representations.C1Constituent A Alone: Single-instrument formal structural calculus.C2Constituent B Alone: Single-instrument epistemic discovery medium.C3Naive Chaining Baseline (A → B): Output of unconstrained generation passed sequentially without typed contracts.C4Naive Chaining Baseline (B → A): Output of unconstrained generation passed sequentially without typed contracts.C5Governed CII Composition: Heterogeneous composition connected through an explicit typed handoff contract, preserving invariant boundaries across transitions.
2.4 Evaluated Primary Variables#
Empirical execution records must directly report:
- Invariant Violation Rate: Percentage of candidate or final outputs violating declared domain constraints.
- Search Efficiency: Mean candidate evaluations or search iterations required to reach a verified valid configuration.
- Attributable Failure Rate: Percentage of failure states whose cause is formally localized to specific boundary contract mismatches or operator exhaustion.
3. Provenance & Ratification Rules#
To prevent ungrounded claims, WinMedia applies the following ratification rules to empirical CII reporting:
- No Synthetic Numbers: Numerical performance claims require direct reference to an immutable experimental run, source dataset, and reproducible scoring harness.
- Separation of Feasibility and Advantage: Demonstrating a valid composition ($A \circ B$) establishes composition feasibility, while demonstrating that $A \circ B$ outperforms both $A$, $B$, and naive chaining establishes composition advantage.
- Claim Boundary: Benchmark results are published under $C_2$ only when accompanied by complete, reproducible execution artifacts; unexecuted protocol designs remain governed under $C_1/C_4$.