The explorer in Appendix B opens the whole model to navigation, and the
notebooks in Appendix C show what the model can prove by computation. This
appendix holds the other half of the account: the subjective choices,
judgments and interpretations the model is grounded in. Every term whose
definition a citation could not settle alone, every step whose canon match
was a reading, every seam added on a judgment call, was raised as a concern
and answered by a named adjudicator, dated. The rulings below are the
adjudicator’s decisions in the register’s own words; the message as sent is
kept in the graph as ogc:verbatim and printed by ogc ruling. Those answers are
implicit in every objective result the computational proofs produce. The
front page’s account of science asks that the assumptions held fixed while
a test runs be written down; these are ours, held fixed while the model was
built, and the claim that the method is scientific holds with them as its
auxiliary assumptions. Appendix D is the specification holding itself to its
own definition.
The table is rendered from rulings/adjudications.ttl and checked by
shapes/rulings.shapes.ttl; a concern marked ruled without a ruling fails
the gate. Open concerns await a ruling.
| # | Concern | Severity | Status | Ruling | Adjudicator | Date | Change |
|---|---|---|---|---|---|---|---|
| 1 | C-01: test coverage: abstract tests vs concrete applications | H | ruled | SEVOCAB test coverage (ISO/IEC/IEEE 29119-2:2021 3.28) is adopted. The distinction between the abstract tests that cover requirements and the concrete applications of those tests during a case, which produce the evidence, is recorded as a scope note rather than as a term. | Z (Michael Zargham), adjudicator | 2026-09-05 | Glossary cites ISO/IEC/IEEE 29119-2:2021 3.28 via SEVOCAB; scope note: acceptance criterion is the coverage item, probe is the test case, evidence is the concrete application. Paper asked to align. |
| 2 | C-02: operating environment: which canonical anchor | M | ruled | The headword is operational environment, adopted from SEVOCAB (IEEE 982-2024 3.1). Operating environment remains an alternative label; the automotive Operational Design Domain sources (SAE J3016, SAE J3259, ISO 34503) leave the canon. | Z (Michael Zargham), adjudicator | 2026-09-05 | prefLabel operational environment (IEEE 982-2024 via SEVOCAB); altLabel operating environment; ODD and SAE dropped from the canon. |
| 3 | C-03: operational envelope: no canonical entry | M | ruled | Operational envelope is kept as a house label, defined by composition and anchored to ISO 9000:2026 3.5.1 requirement as a specialization, with requirement set as an alternative label. It is not counted as a coinage. | Z (Michael Zargham), adjudicator | 2026-09-05 | Class refined; anchored to ISO 9000:2026 3.5.1 requirement with relation specializes; not counted as coinage. SysML item def RequirementSet. |
| 4 | C-04: redistribution of source snapshots | M | ruled | SEBoK v2.14 and the ISO 9000:2026 screenshots are both held locally and not committed. The repository commits their hashes and digests only, and the checks verify the quotes against the digests. | Z (Michael Zargham), adjudicator | 2026-09-05 | sources.ttl posture heldLocally for SEBoK v2.14 and the ISO 9000:2026 screenshots; digests with quotes and locators committed; CI verifies quotes against digests with pinned counts. |
| 5 | C-05: deployment sensitivity: refined term or coinage | L | ruled | Deployment sensitivity is a refined term that specializes SEVOCAB risk-based testing (ISO/IEC/IEEE 29119-2:2021 3.16). It is not a coinage. | Z (Michael Zargham), adjudicator | 2026-09-05 | Anchor SEVOCAB risk-based testing (29119-2 3.16), relation specializes, altLabel risk weight. Coinage stays at three. |
| 6 | C-06: IRI base: w3id now or Pages URL | L | ruled | IRIs are authored under https:// | Z (Michael Zargham), adjudicator | 2026-09-05 | All IRIs under https:// |
| 7 | C-07: scope: requirements and interfaces, or architecture too | M | ruled | The scope is requirements and interfaces; architecture is excluded. Part and port definitions stay at the interface level, and the infrastructure drawing is cited as a realization sketch with nothing allocated to it. | Z (Michael Zargham), adjudicator | 2026-09-05 | SysML carries part, port, interface, action and requirement defs only; no allocation, no physical structure; the infra drawing is cited as a realization sketch. |
| 8 | C-08: adequacy vs appropriateness | H | ruled | The pairing is appropriateness and sufficiency. The earlier label adequacy was a misquotation of the primary source; the functional use was correct and only the label was wrong. The primary source is to be checked and cited correctly. | Z (Michael Zargham), adjudicator | 2026-09-05 | Every artifact uses appropriateness (of the declared context) and sufficiency (of the evidence); citations to Hawkins et al. 2011 3.2/3.3 with ISA 500 5(b)/5(f)/A8 as seeAlso; paper and the authors’ term contract draft asked to align. |
| 9 | C-09: the coinage set: which terms do we own | H | ruled | Coinage is confined to the project’s novel contributions. The two primary coinages are Domain-Specific Ontology and Evaluation Process Ontology; the method is named either CAIE or Ontology-Grounded CAIE, abbreviated OG-CAIE. | Z (Michael Zargham), adjudicator | 2026-09-05 | Coinage is exactly three: Domain-Specific Ontology, Evaluation Process Ontology, OG-CAIE (the method name). Probe and deployment sensitivity are refined terms (R-05); operational envelope is a house label (R-03). Julie Hollek and Z agreed the DSO and EPO expansions in the paper comments (a co-author’s note, 2026-09-05). |
| 10 | C-11: technical expert: one kind or several on the evaluation team | M | ruled | The world model has more than one kind of technical expert. The domain expert is a technical expert in the domain-specific ontology: a person who can attest to the appropriateness of the Domain-Specific Ontology for the particular case being evaluated, that is, whether this system under test passes these requirements in this operating environment. The red-teamers are also technical experts, though not in the auditing sense: their expertise is in AI evaluation; they perform the tests the EPO specifies, conditioned on the DSO the domain expert supplies, and their job is to assemble and interpret evidence. Domain experts may be called on to consult on the interpretation of the resulting evidence. Both domain experts and AI experts belong on the evaluation team. | Z (Michael Zargham), adjudicator | 2026-09-06 | Glossary: technical expert kept as the genus (ISO 9000:2026 3.12.9), narrative rewritten to name the two kinds on the evaluation team and what each is expert in; bound to both DomainExpert and Evaluator. Model: doc comments of DomainExpert and Evaluator state their expertise and the case (system under test, requirement set, operational environment) the appropriateness judgment is about. |
| 11 | C-12: assessing the assessors | L | ruled | Technical experts may in principle have mirrored counterparts who assess the assessment, but this document does not model assessing the assessors. The one plausible place for the idea is a closing call to action, offering the reader OG-CAIE or an audit of the reader’s own AI evaluation practice against this specification. | Z (Michael Zargham), adjudicator | 2026-09-06 | No assessor-of-assessors parts or shapes. The conclusion closes with a call to action: use OG-CAIE, or have an AI evaluation practice audited against this specification. |
| 12 | C-13: the chain from requirement to attestation: is the test plan explicit | H | ruled | Traceability from the requirement, as a hypothesis that may be falsified, through an epistemically valid process that collects and interprets evidence to render the judgement, is critical for requirements to be useful concepts. The existing definitions are not to be overwritten. The record must carry a chain of evidence connecting the test plans to the evidence they would produce, the implementation of those plans that produces the evidence, the interpretation of that evidence, and the justified attestation that a requirement is or is not met, bound to that evidence and to the test plan that produced it. This is the level of abstraction at which the abstract EPO is encoded. | Z (Michael Zargham), adjudicator | 2026-09-06 | Added, never overwriting: glossary terms test plan and expected results (SEVOCAB, 29119); seeAlso test execution on test, test result on attestation, test scenario (29119-11) on scenario. EPO: TestPlan (objectives, means), expectedResult on the criterion, underPlan from attestation to probe. Record: the plan, expected results on a1-a3, both attestations bound to probe-1, recommendation derived from the plan. Shapes: S2 criterion requires an expected result; S3-TestPlan; S3 probe must be a plan’s means; S6 chain rule (evidence used comes from a run of the bound probe which exercises the attested criterion); S8 recommendation names the plan. Counterexample attestation-off-plan. Model: TestPlan item, expectedResultStated, probeId on the attestation, SCI-04 and SCI-06 extended. Traceback query returns plan, probe and expected result. |
| 13 | C-14: probe versus run: trajectories, sessions and state feedback policies | H | ruled | The eight-term package is adopted as proposed: session and dialogue (NIST); trajectory, with the output-feedback anchor for the policy (IEC 60050-351); test strategy (ISO/IEC/IEEE 29119-2); non-deterministic system (29119-11); repeatability, reproducibility and measurement uncertainty (VIM, NIST TN 1297). Session and Turn enter the record; a Strategy is a plan’s means; the chain rule is extended accordingly. | Z (Michael Zargham), adjudicator | 2026-09-06 | Sources: IEC 60050-351 (Electropedia, cite-only, quotes pending), JCGM 200:2012 (held locally), NIST TN 1297 (committed). Glossary: eight terms; probe and test narratives tightened to one turn and to the session. EPO: Session, Turn, Trajectory, Strategy (a prov:Plan), inSession, turnIndex, policy, groundedIn; a TestPlan’s means is a Probe or a Strategy. Record: session-1 with turn-1 and trajectory-1 replace run-1. Shapes: S4-Session, S4-Turn, S3-Strategy; S5 evidence from a turn; S6 chain rule through the turn; S3 probe belongs to a plan directly or through a strategy. Counterexample attestation-off-turn (two-turn session under a strategy). Model: Session item, sessionId and turnIndex on Evidence, SCI-05 extended. Traceback query returns session and turn. |
| 14 | C-15: operational envelope: what it adds to a set of requirements | M | ruled | The operational envelope refines the concept of requirements in its specificity: it conditions explicitly on an operating environment and on the system under test, stating that this system in this environment must meet these requirements. It is closely related to a set of requirements, but Contextual AI Evaluation expects different requirements for different systems under test in different operating environments. The distinction is deliberate: it separates the method from benchmarks and aligns it with the practice for safety-critical systems. | Z (Michael Zargham), adjudicator | 2026-09-06 | Anchor unchanged (ISO 9000:2026 3.5.1 requirement, specializes). Narrative and scope note rewritten. EPO: RequirementSet names its system under test (epo:systemUnderTest, an earl:TestSubject) as well as its environment; shape S2 requires both; SysML RequirementSet gains sutNamed and SCI-02 checks it; the measles record names chatbot v1. |
| 15 | C-16: requirements traceability and metrological traceability | M | ruled | Requirements traceability and metrological traceability are kept distinct. When both are done computationally with knowledge graphs they can be chained, because they are in principle the same notion at different levels of abstraction. A note records that computational instruments are used to measure computational systems, so the concept of metrological traceability still applies. | Z (Michael Zargham), adjudicator | 2026-09-06 | VIM 2.41 metrological traceability added as a seeAlso on requirements traceability with a scope note stating the relation; canonical anchor unchanged (SEVOCAB, 29148). |
| 16 | C-17: conformity, conformance, and what SHACL calls validation | H | ruled | Since ISO deprecates conformance, the word is reclaimed for the specific requirement that the record be correctly constructed according to the SHACL shapes, and conformity is used for fulfilment of a requirement in the ISO sense. The reclamation is to be noted explicitly. One reason for it is to break the problem that SHACL “validation” is, by this specification’s standard, verification; using conformance specifically for the correct-construction requirements applied to the record keeping the EPO demands makes it a special case of conformity that covers requirements which might otherwise create confusion in V&V terms. | Z (Michael Zargham), adjudicator | 2026-09-06 | Two glossary terms. conformity: adopted, ISO 9000:2026 3.5.9, fulfilment of a requirement. conformance: refined, specializes conformity, anchored to the same clause’s Note 1 (the deprecated synonym, reclaimed on purpose), technical binding W3C SHACL section 3.5 Conformance Checking; the machine-checked, correctly constructed record that satisfies the EPO’s record-keeping requirements. Prose: conformance for the SHACL check, conformity for fulfilling a requirement; SHACL’s word validation is not used for it. Model: ConformityChecker renamed ConformanceChecker. The docs test no longer bans the word. |
| 17 | C-18: evaluation record: what the EPO binds it to provide | M | ruled | The evaluation record in this context is bound by the EPO, which determines what must be present for formalized traceability in both directions and for coverage. The EPO makes all three directly checkable by SHACL over the evaluation record; a conformant evaluation record therefore provides both traceability directions and coverage. | Z (Michael Zargham), adjudicator | 2026-09-06 | Narrative of evaluation record rewritten to say so; scope note names the two directions (downward flow-down from requirement to criteria to plan to probes, upward derivation from attestation to evidence to probe to criterion to requirement). New shape S2-Requirement: every requirement belongs to the requirement set and has at least one acceptance criterion derived from it, so the downward direction is checked; the upward direction was already checked by S6 and S8. |
| 18 | C-19: evidence: at what level, and what it maps to | H | ruled | Evidence can exist at the probe level, at the session level, and rolled up to a multi-session scenario such as a sensitivity test or a robustness battery. It is therefore not necessarily bound to a probe, though it may be. Evidence must be related to the test plan and in particular maps to an expected result: something that can be ruled on directly as a Boolean, whether it did or did not happen. | Z (Michael Zargham), adjudicator | 2026-09-06 | EPO: Response (what the system produced at a turn) separated from Evidence; Evidence answers exactly one acceptance criterion’s expected result with a Boolean observed value (a determination, ISO 9000:2026 3.11.1), derives from one or more responses or trajectories, is attributed to whoever determined it, and is bound to the test plan; TestSuite (SEVOCAB, altLabel battery) groups sessions for the multi-session level. Attestation binds to evidence; the chain rule now requires every evidence an attestation uses to answer the attested criterion under a plan that has it as an objective. Record: response-1, evidence-1a (a1, observed false), evidence-1b (a2, observed false). Shapes S5-Response, S5-Evidence, S4-TestSuite; S6 chain rule rewritten. Glossary: test suite added; objective evidence narrative rewritten. Model: Response and two Evidence items, level enum, SCI-05 and SCI-06 rewritten. |
| 19 | C-10: system under test: weak canonical entry | L | ruled | The recommendation on C-10 is accepted, option (b): the headword is test item (SEVOCAB, ISO/IEC/IEEE 29119-2:2021, work product to be tested), with system under test, SUT, test object, application and system of interest as alternative labels and the related entries as seeAlso citations. Prose and the model may continue to say system under test. | Z (Michael Zargham), adjudicator | 2026-09-06 | The recommendation accepted: headword test item (SEVOCAB, ISO/IEC/IEEE 29119-2:2021, work product to be tested) with system under test, SUT, test object, application and system of interest as alternative labels; seeAlso system-of-interest (15288), NIST AI 700-2 application, AI-based system (29119-11), object of conformity assessment (ISO/IEC 17000 4.2), and the 14756 system under test entry kept visible as the only standard definition of the phrase. Prose and the model keep saying system under test. |
| 20 | C-20: evidence and determination: domain and codomain | H | ruled | The R-18 implementation is wrong: evidence is the domain and the determination is the codomain. Evidence is collected and serves as the basis for a determination of whether the expected outcome has been met. The determination is not strictly Boolean: per EARL, cantTell is admitted. | Z (Michael Zargham), adjudicator | 2026-09-06 | Evidence keeps only what it is: collected material (a response, a trajectory, or several) bearing on one criterion’s expected result, under the test plan, attributed to its collector. Determination is a separate node, an earl:Assertion and prov:Activity: it uses one or more evidence items, tests the criterion, and records the EARL outcome passed, failed or cantTell, its mode and its assertor; the paper’s word for it is judgment. The attestation aggregates the determinations for a criterion. Glossary: determination becomes a term (ISO 9000:2026 3.11.1, altLabel judgment); objective evidence’s narrative corrected. Shapes: S5-Evidence without observed; S6-Determination; the closure rule and chain rule run attestation to determination to evidence. Model: Determination item def replaces the observed attribute. |
| 21 | C-21: parties to the evaluation are implicit | H | ruled | Explicit stakeholders are introduced. There are affected stakeholders; the organization responsible for the system under test (the test item); the organization performing the testing; and possibly an organization sponsoring the testing that is not the organization operating the test item. Affected stakeholders are populations of actors, whereas the sponsor organization, the testing organization and the organization accountable for the test item are concrete single entities. The requirements are established and agreed between the sponsor organization and the testing organization; the arrangement is viewed as a contract to perform OG-CAIE as a service. The organization that owns the system under test may or may not be the sponsor, and members of the affected stakeholder classes may or may not be interviewed as part of the requirements development process. Most often the interests of affected stakeholders are represented in the requirements on the basis of input from the domain expert, who is employed in the selection of the domain-specific ontology and in the appropriateness assessment of the requirements; the domain expert may also participate in developing or approving test plans, in the interpretation of the evidence collected, and in attestations to the conformity, or not, of the system under test with regard to those requirements under those test plans. This domain expert role is distinct from the expert role responsible for administering the tests and gathering the evidence per the test plan. Part definitions and port definitions specify this as an abstract wiring diagram before it is tested by concrete walkthroughs that exercise the specification on a simple demonstration case. | Z (Michael Zargham), adjudicator | 2026-09-06 | Parties become part defs: SponsorOrganization (ISO 9000:2026 customer), TestingOrganization (provider), AccountableOrganization (ISO/IEC 17000 first party), AffectedPopulation with multiplicity 0..*, EvaluationTeam (refined from audit team). A service agreement (ISO 9000:2026 contract) between sponsor and testing organization precedes the requirement set. Populations are interviewed (a StakeholderInput) or represented by a domain expert, and the record says which. The EPO becomes seven steps (superseded by R-32: six evaluation steps nested in seven contracting steps): agree, scope, declare requirements, plan, execute, determine and attest, report. Glossary: organization, customer, provider, first-, second- and third-party conformity assessment activity, contract, evaluation team added (this slice); model, shapes and record follow in the next slices. |
| 22 | C-22: which artifact is canonical: the SysML source or its RDF rendering | H | ruled | SysML is used primarily for part definitions, port definitions and wiring rules: structure only, with value checks living in SHACL. The rendered RDF version of the models is the canonical record, and the SysML source is treated as a view, even where it is the view used for authoring. OpenSysML renders SysML models as RDF triples because it is designed to work with Flexo MMS; the conversion is deterministic and uses the OMG SysML vocabulary with stable IRIs. A parsimony heuristic prunes the TBox not required by the model, following the pattern applied and described in the authors’ earlier work. | Z (Michael Zargham), adjudicator | 2026-09-06 | SysML holds structure only: part defs, port defs, interface defs, the action def and the assembly; requirement defs and -satisfy leave the model. The pruned RDF rendering of the model (model/og-caie.model.ttl, committed, byte-identical in the gate) is the canonical structure and the SysML source is the authoring view. Wiring rules are SHACL shapes over the model graph; value and provenance rules stay SHACL shapes over the record. Parsimony follows the authors’ earlier lifecycle model builds: a term map of the OMG terms kept, a manifest, and a triple budget with a rationale. |
| 23 | C-23: actor categories on the testing organization’s side, and named persons in the demo | M | ruled | A domain expert and an evaluation operator are abstract actor categories that populate distinct slots in the EPO. There is also an account executive on the side of the evaluation organization, who is the signatory on the contract and the legal representative of the organization performing the evaluation. The worked example names the Humane Intelligence executive (Mala), the domain expert (Annie) and the evaluation operator (Theo). These are references to real people, but the case is synthetic: the people named will recognise their roles, and they are not asked to make attestations within the demonstration case, which is a simplified synthetic walkthrough of the process. DomainExpert, EvaluationOperator and AccountExecutive are all roles held by members of the company performing the evaluation. Each role must be filled, so there is an at-least-one rule. The account executive is exactly one, being the counterparty who represents the organization in the contract; the domain expert and evaluation operator roles are one or more. | Z (Michael Zargham), adjudicator | 2026-09-06 | Three abstract actor categories, each a distinct part def and EPO slot, all held by the testing organization (accountExecutive[1]; team.domainExpert[1..]; team.operator[1..]; shape M5-Cardinality): DomainExpert, EvaluationOperator (headword evaluation operator, refined from the SEVOCAB operator; AI evaluation expert becomes its altLabel) and AccountExecutive (signatory and legal representative of the testing organization; no evidence, determination or attestation port). The worked example names Mala (account executive), Annie (domain expert) and Theo (evaluation operator); the record and the record page state that the case is synthetic, that the names recognise the roles of real people, and that no attestation in the record was made by them. |
| 24 | C-27: port reuse across actor categories: activities confused with actors | M | ruled | Port sharing is bad practice: it confuses activities with actors. Actors may perform multiple activities, but each activity needs precise inputs and outputs. | Z (Michael Zargham), adjudicator | 2026-09-06 | Every step of the EPO action def now declares its inputs as well as its outputs. Delivery is an item kind of its own (the report and the recommendation handed to the sponsor), the output of the account executive’s delivery activity, carried on DeliveryWrite through a delivery seam to the sponsor and a second seam into the record; the sponsor’s report and recommendation ports and the executive’s reuse of RecommendationWrite and ReportWrite are gone. Shape M5-PortsBelongToRoles now restricts the account executive to AgreementWrite and DeliveryWrite. The one activity two actor categories may both perform is determining on evidence (R-21), so DeterminationWrite is the only supplier port definition carried by two roles, and the wiring test says exactly that. |
| 25 | C-28: one wire per port, or shared output wires | M | ruled | Input wires must be unique. Output wires may be shared, provided that what flows over them is information whose use is nondestructive. | Z (Michael Zargham), adjudicator | 2026-09-06 | Shape M2-InputsUniqueOutputsShared replaces M2-PortConnectedOnce: every port is wired; an input (conjugated) port is the end of exactly one seam; an output port may feed several. The recorder has one record output read by six, the probe deriver one probe output read by the operator and the recorder, the account executive one delivery output read by the sponsor and the recorder: 47 ports, 27 seams. Everything that flows is information. |
| 26 | C-29: what the wiring rules are rules over, and what the EPO is as a process | M | ruled | Every wiring must be checkable and explainable locally. At the part level, all of a part’s input ports and output ports are examined to ensure that every input is present and every output goes somewhere. At the wire level, each wire is examined for where it comes from and where it goes: an output port on a part to an input port on a part. If all rules are defined over kinds of parts and kinds of ports, then enforcing the rules for all parts and wires yields a well-constructed wiring diagram; ports are how wires are addressed. This is essentially the typed block diagram once directionality and the cardinality rules are enforced: one input wire per port, and output wires that may split as long as they are informational. In all likelihood the EPO is a process DAG. Loops are allowed if they are needed to guarantee that the final output record has full requirements traceability in both directions and test coverage, computable by SHACL constraints. The ability to ensure a full record with executable checks on traceability and coverage is precisely what allows the final result to be trusted, and it supports drilling in and deterministically answering the “why” questions related to that result. | Z (Michael Zargham), adjudicator | 2026-09-06 | Shapes M2-Wire (per wire: from an output port on a part to an input port on a part, one item kind, ends typed by the seam) and M2-Part (per kind of part: every input present on exactly one wire, every output going somewhere) replace the port-targeted shape; M4-ProcessDag checks the succession graph has no cycle (loops stay admissible by ruling if the end state needs them; none does). queries/wiring.rq explains every port locally and renders as a table on the assemblage page. SCI-11 restated. |
| 27 | C-24: contract: no capture of ISO 9000:2026 clause 3.3.13 | L | ruled | The ISO 9000:2026 clause for contract is 3.3.13. The page carrying clauses 3.3.9 to 3.3.13 is captured and the quote, binding agreement, is verified; sheet 03 row 7 is ticked. | Z (Michael Zargham), adjudicator | 2026-09-06 | Z captured the OBP page for clauses 3.3.9 to 3.3.13 (screenshot 38, hash in sources/sources.ttl): the 2026 wording of contract is the 2015 wording, binding agreement. The quote is verified by Z on 2026-09-06; sheet 03 row 7 ticked. |
| 28 | C-31: evaluation team anchored to audit team: the word audit reads as legal | M | ruled | Evaluation team is anchored to the SEVOCAB evaluator. The other two candidates, ISO 9000:2026 3.12.7 audit team and SEVOCAB assessment team (ISO/IEC 33001:2015 3.2.10), are kept as seeAlso citations. | Z (Michael Zargham), adjudicator | 2026-09-06 | Evaluation team is now refined from the SEVOCAB evaluator (ISO/IEC 25000:2014, 4.18: individual or organization that performs an evaluation), the same SQuaRE family as the adopted term evaluation; ISO 9000:2026 3.12.7 audit team (pending Z’s tick, sheet 03 row 8) and SEVOCAB assessment team (ISO/IEC 33001:2015, 3.2.10) are seeAlso. The label evaluator leaves the evaluation operator’s altLabels. Technical expert stays under R-10 with its scope note. |
| 29 | C-32: the front matter does not prepare the audience; the glossary is rendered whole; Popper is everywhere | H | ruled | The Popper concepts connect to the main OG-CAIE ones through a simple crosswalk table from the Popperian conception of science to the project’s terms. The entire glossary is not rendered; the glossary navigation approach from the authors’ earlier work is adopted; key terms carry mouse-over definitions in MyST. This plan defers the work through the whole executable EPO specification, with parts, ports and wirings that demonstrably guarantee a scientific record, and the worked health example. The flow for contracting and mapping stakeholders is separated from the flow for performing the evaluation into distinct chapters. All of that remains required, but the glossary is operationalized first and the prose of the first part tightened, so that the audience is prepared to enter the specification for the contracting workflow and then for the evaluation workflow. Each chapter is carefully scoped and revised one at a time as it is reached, working forward. This breaks the modelling into bounded definitions and checks, lowers the modeller’s cognitive overhead, and helps present the model clearly in the MyST site. Shorthands are provided: DSO for Domain-Specific Ontology, EPO for Evaluation Process Ontology, CAIE for Contextual AI Evaluation, and OG-CAIE for CAIE performed using the EPO and DSO method this document puts forth. That is the full extent of the coinage; everything else is grounded in cited standards and literature. The Popper-to-standards crosswalk exists to set up what would make the evaluation scientific, and the conclusion must be that OG-CAIE encodes exactly what qualifies as scientific. Popper’s definitions occur only in the set-up and in the final conclusion: a coordinate transform from the low-dimensional Popperian conception of science to the high-dimensional engineering standards and back, so that the main narrative arc is accessible to less technical readers. The top-level page is rescoped from why this counts as science onward. The bridge to formal engineering standards for evaluation is the endpoint of the first chapter, which sets up a more rigorous, though not exhaustive, glossary; the glossary in turn supplies the terms needed to express the executable specifications and then to demonstrate their properties. | Z (Michael Zargham), adjudicator | 2026-09-06 | Slice A of the site restructuring, one chapter at a time, working forward. Coinage becomes exactly four: Contextual AI Evaluation (CAIE) added as a coined term, OG-CAIE redefined as CAIE performed with the EPO and DSO method, DSO and EPO unchanged. The front page runs from why this counts as science to the bridge into the engineering standards and ends at the forward Popper crosswalk (vocabulary/crosswalk.ttl, six rows citing the BoF deck, rendered to generated/popper.md); the conclusion page reads the crosswalk backwards and states that the essentials, shapes and record encode exactly what qualifies as scientific; a test confines the Popperian words to those two pages. The glossary page carries the site map, the discipline, and a hover glossary of exactly the terms the prose references ({term} roles rendered to a {glossary} directive by scripts/render.py); the full table is generated for the paper only. Deferred to later slices, each scoped when reached: the contracting chapter, the evaluation chapter (absorbing the assemblage and record pages), the ogc navigation CLI ported from the authors’ earlier glossary work; C-30 stays open. |
| 30 | C-33: the bookends test banned ordinary words | L | ruled | Care is needed with the Popperian terms because they are common English words. The restriction is not on using the words but on using them in Popper’s sense with Popper’s definitions. The discussion of what is science, grounded in Popper, stays in the bookends, since it is the paper’s core argument: “CAIE would be science if it covers these things”, ending with “OG-CAIE demonstrably covers these things”. The inner sections achieve this through engineering standards. The introduction and conclusion cover why and what; the inner chapters cover what and how in enough detail, and executably, to show both in theory through the executable specification and in practice through the worked example that the bar for science is met. | Z (Michael Zargham), adjudicator | 2026-09-06 | The test now bans only the name (Popper, Popperian) and falsifiability’s forms from the inner pages; hypothesis, prediction, auxiliary and evidence are free words there, and the rule about their sense is stated in CLAUDE.md and read, not grepped. The two scrubbed sentences on the record page stay scrubbed: they used the words in Popper’s sense. |
| 31 | C-34: the seven steps were not bound to a citation | M | ruled | The first page, OG-CAIE as an executable specification, is in good shape for now. One thing stands out: the seven steps are not bound to a citation. Given the INCOSE canon and quality management, a process model for evaluation is almost certainly documented already; it is to be cited and the steps matched to its steps, to reinforce that nothing here is invented, only operationalized through executable code. | Z (Michael Zargham), adjudicator | 2026-09-06 | Each EPO step in vocabulary/epo.ttl now carries a canonical citation and seeAlso citations: agree, scope and declare requirements match the ISO/IEC/IEEE 15288 acquisition, stakeholder needs definition and system requirements definition processes as SEBoK v2.14 describes them (pp. 353, 549, 557; machine-located); plan, execute and report match the ISO/IEC/IEEE 29119-2 test strategy and planning (3.51), test execution (3.40) and test completion (3.25) processes as SEVOCAB defines them (machine-located); determine and attest matches ISO/IEC 17000 7.2 decision with 7.1 review, and Annex A’s functions (selection, determination, review, decision and attestation) frame the whole, cited by heading. The EPO term cites 29119-2’s test process, SEBoK’s four 15288 process groups and 17000 4.1 Note 3. The seven ISO/IEC 17000 quotes are pending on sheet 04. Rendered as a table on the glossary page; ogc steps prints it; the citation tests and ogc verify cover the steps. |
| 32 | C-35: the contracting lifecycle was absent from the model; the two layers of assumptions were not distinguished | H | ruled | The contracting lifecycle model receives the same treatment as the evaluation steps. Both cycles are in the paper. The contracting cycle connects to the stakeholders and counterparties and, in this work, pins the first layer of auxiliary assumptions. There must be an explicit separation between the layer of work pinned at the point of the contract and what is pinned and recorded within the scope of an evaluation being performed. Although the evaluation is the main focus, the contract provides axiomatic boundary conditions that bound the scope and also supply assumptions which must be documented for the evaluation to be scientific. The stakeholder terms already introduced are the actors within that lifecycle, which runs from contracting through delivery on the contract. The lifecycle is not to be invented; it is taken from the relevant engineering standards. The contracting lifecycle exists because work has been performed, but it was absent from the model. It is made explicit because it bears on the auxiliary assumptions: the record must distinguish what was pinned by the contract from what was decided in the course of the contract’s fulfilment. | Z (Michael Zargham), adjudicator | 2026-09-06 | The contracting lifecycle is now explicit: six steps C1 need, C2 propose, C3 agree, C4 access, C5 deliver, C6 acceptDelivery (epo:ContractingStep; accept is a SysML keyword), each citing the canon step it matches: SEBoK’s business or mission analysis (p. 543), supplier and acquisition process (pp. 352, 353); ISO/IEC 17000 access (4.10), scheme (4.9) and acceptance (9.6); ISO 9000:2026 contract (3.3.13) and requirement (3.5.1); SEVOCAB test environment (29119-2 3.34), test completion report (3.26) and acceptance (24748-5). The evaluation keeps six steps, scope to report; agree left it. New item kinds Need, Proposal, Acceptance. Every item class declares ogc:pinnedAt epo:contract or epo:evaluation; shape S0-Layers checks that the first layer is closed before the second opens; S0-Need, S0-Proposal and S9-Acceptance check the new items; the record’s measles run carries a need, a proposal and an acceptance. The Popper crosswalk’s auxiliary-assumption row splits into the two layers; SCI-13 states the separation. The model follows in the same slice: a ContractingProcess action def wrapping the EvaluationProcess. |
| 33 | C-36: two cycles as two action defs, or one nested model | M | ruled | The two-layer model is to be one model exercising the nesting power of SysML, provided OpenSysML supports it: the outer model treats fulfilment as a black box with well-defined inputs and outputs, and the drill-down specifies that black box in more detail as a specification of its own, which necessarily conforms to the inputs and outputs, and carries the assumptions, of the layer above. This is committed to the plan and to memory as the way the two layers are modelled. | Z (Michael Zargham), adjudicator | 2026-09-06 | Possible, and done. Probed on OpenSysML v0.4.3: action fulfil : EvaluationProcess; inside the contracting action def validates strictly and converts to a sysml:ActionUsage typed by the inner def; flow a.x to b.y between steps converts to sysml:FlowUsage with two resolvable ends; a flow from an action def’s own parameter is refused (both ends need dot notation), so bind step.param = param; binds it (sysml:BindingConnectorAsUsage). The contracting lifecycle now owns seven steps, need, propose, agree, access, fulfil, deliver, acceptDelivery; fulfil is the evaluation process as a black box, fed the agreement and the access by flows and feeding the report and the recommendation to deliver; the evaluation process binds its own parameters to its steps and flows its items between them; the assembly performs one action. Shape M4-Nesting checks that fulfil is typed by the evaluation process and that every input of the inner process is fed by a flow from the outer and every output used by one. The prune script resolves flow ends (an inline action resolves against its own parameters) and records them as ogm:flowSource and ogm:flowTarget; FlowUsage, BindingConnectorAsUsage and sysml:references join the term map; the budget rises to 5200 from measurement. Recorded in the plan and in memory as the way the layers are modeled: type a step by an action def, bind its parameters, never restate its interface. |
| 34 | C-37: the site had no stated relation to the model; pages were not paired specification and example | H | ruled | The site is replanned. The MyST site is a presentation layer on top of a rich model that documents itself in RDF. The glossary records the terms, their definitions, their lineages, their cross-relationships and the judgement calls that led to the choice of those definitions. The SysML model covers the actual processes, which are bound to the standards and literature they implement; the processes encoded become executable, so their properties become verifiable. The worked examples demonstrably conform to those processes and produce data that meets the specifications for the wires in the model. The formal properties can be checked (correct wiring, and any rule encoded as SHACL shape conformance), and the examples walk through to make it concrete and intuitive for a human reader. Each page of the MyST site is therefore a curated view of the underlying model. The updated outline delivers a story arc that uses the Popper bookends to make this intuitive for a general audience, then inside uses the formal engineering standards for contracting and for quality evaluation processes as a nested SysML model, rendered as RDF and enriched with the glossary content to ensure citation traceability; ultimately guaranteeing an auditable record of the AI evaluation done using OG-CAIE, by encoding and enforcing a DSO-conditioned EPO, and concluding that this record meets Popper’s bar for being scientific. The task builds out what the model (representation) and the site (presentation) each must do. The separation of representation and presentation is one of the main value propositions of doing this with GitHub, MyST, Jupyter, Python, OpenSysML, rdflib, pySHACL and other executable systems engineering tools, which allow text-based standards to be codified in a formal record with machine-provable traceability in both directions and coverage. Eight pages, as listed; each page includes both the formal abstract specification and the relevant concrete example, the two running in parallel throughout, with a recognisable pattern so that the reader can find them. There is always more: highlighting the separation principle makes clear that the presentation layer (the site) is a view rather than the model itself (the repository). The presentation is calibrated for a human reader, so the texture and the volume of the content are tuned to human sensibilities. The representation layer is made further accessible to AI through CLI tools and skills, as was done for the glossary. In practice the views are tailored to carry the most important information for the human reader, and the concrete examples are made didactic relative to the abstract concepts; the two always appear as a pair. | Z (Michael Zargham), adjudicator | 2026-09-06 | The plan of record (2026-09-06) states what the representation must hold and prove and what each of eight pages shows and is rendered from: Why this counts as science, The vocabulary, Contracting, The evaluation, The nested model, What the record proves (executed live in the build), Rulings, Conclusion. Every chapter from Contracting on follows five titled blocks in order (What the standards say; The specification; The walkthrough; Checked; There is more in the model), enforced by a test, with a per-page prose budget. Slice B lands the Contracting chapter: renderer slices by item kind and by step, page tags on the essentials, and rulings sheet 05, generated from the model graph, for Z’s block-by-block and wire-by-wire validation (C-30). |
| 35 | C-38: what happens after acceptance: narration, not a step | L | ruled | The walkthrough may describe two further steps that are explicitly not part of the contract lifecycle and are identified instead as observed in practice: the county health office tells the chatbot provider to implement changes that result in demonstrable behaviour change per the findings of Humane Intelligence, and the chatbot provider implements the changes, at which point the county health office may require new testing to certify that the flagged issues were addressed. This is not claimed to be in the contracting loop of the standard contracting process; it is an opportunity to show how the evaluation results in demonstrable system behaviour change. For safety-critical deployments these evaluations would be expected to be paired with a release. Since the sponsor organization and the chatbot provider are both already in the example, this is primarily narration, to help the reader understand what the process is like in practice and why it is not an academic exercise. | Z (Michael Zargham), adjudicator | 2026-09-06 | Narration only: a paragraph in the Contracting chapter’s walkthrough block, after the record rows, marked as observed in practice and deliberately outside the lifecycle; no step, item or shape added. The standards’ word for the repeat, surveillance (ISO/IEC 17000 8.1), is quoted and cited as a seeAlso on the contracting lifecycle class, pending Z’s tick on sheet 04 row 12. |
| 36 | C-39: surveillance as a headword raises eyebrows | L | ruled | The headword surveillance would raise eyebrows and is replaced. The citation to ISO/IEC 17000 8.1 surveillance stays mapped to the concept, since it is the formal doctrine on the subject, but an alternative headword is chosen, monitoring or oversight. | Z (Michael Zargham), adjudicator | 2026-09-06 | Headword monitoring, adopted from ISO 9000:2026 3.11.3 (determining the status of a system, a process or an activity; Note 3: carried out at different stages or at different times), which outranks 17000 by the precedence rule; surveillance (ISO/IEC 17000 8.1) is the alternative label with its clause and a seeAlso citation, so the doctrine stays mapped. Oversight is not a term of either standard. The walkthrough prose says monitoring; both quotes pending on sheet 04 (rows 12, 13). |
| 37 | C-40: no seam carried the sponsor’s obligations towards the affected populations | M | ruled | An extra seam is added to account for any data covering the obligations, duties or expressed mission the sponsoring organization has with regard to the affected populations. This is context that informs what goes into the contract. | Z (Michael Zargham), adjudicator | 2026-09-06 | Item kind Mission (ISO 9000:2026 3.4.11 mission, with 3.4.5 policy as neighbour; glossary term mission, pending Z’s tick), pinned at the contract, produced at the need step before the need, regarding one or more affected populations. Seams missionSeam (sponsor to recorder) and missionToExecutiveSeam (sponsor to account executive), the sponsor’s one output shared; flows carry it into propose and agree; shape S0-Mission checks it is the sponsor’s, names the populations, and that the need is stated under it. The measles record opens with the county office’s mission; SCI-10 says it; the crosswalk’s first layer names it; the Contracting chapter says it in the standards block, the specification and the walkthrough. |
| 38 | C-41: the mission’s relation to the affected populations was in the record but not in the model or the view | M | ruled | Some connectivity should show how the mission of the sponsor organization relates to the affected population; it may be covered in the data but not in the view. Diagrams tend to be overloaded, so in practice they must stay simple, even if some wires are braided into bundles, since the information flows together in practice. Diagrams are important views for humans, so their density must be balanced against the point they make. As long as the code that produces a diagram reads from the model itself and the perspective the view encodes is documented (what it brings into focus and what it leaves out), any view that is compelling narratively is permitted. The view code is hardened by building reusable views and committing the methods with a CLI and skills for their use, which makes other views possible even though the presentation layer curates. | Z (Michael Zargham), adjudicator | 2026-09-06 | Model: connection def Obligation (sponsor end, population end) and the assembly’s connection obligation from sponsor to affected; the Mission item owns ref part regards : AffectedPopulation[1..*], which the record’s epo:regards binds to; derived ogm:relatesFrom and ogm:relatesTo; shape M1-Obligation with the counterexample no-obligation; budget 5500. Views: registry ogc/views.py (layers, assemblage, contracting, evaluation), each with a title, what it brings into focus and what it leaves out; wires between the same two parts braid into one bundle labelled by the item kinds in the order the process produces them; relations dotted; the site’s figures render from the registry with the perspective as caption; ogc views and ogc view ; skill updated; sheet 05 gains a Relations table. |
| 39 | C-42: every chapter says there is more in the model, but a human reader has only the files and the CLI to find it | L | ruled | A knowledge graph explorer is added as an appendix, modelled after the authors’ earlier d3 work in several repositories. Since the text keeps saying there is more, it is easier to cut the text and link to the explorer. Given the extent of the OWL, SHACL and SPARQL work and the existing machine navigation tools, the human-facing graph exploration interface belongs in the appendix; it must back onto the RDF model (Oxigraph may be used for speed). The work may proceed in parallel with the main body. The appendix is an alternative user interface, not new content. | Z (Michael Zargham), adjudicator | 2026-09-06 | An appendix page with a d3 knowledge graph explorer rendered deterministically from the RDF (vocabulary, sources, rulings, essentials, shapes, crosswalk, process, model graph, measles record), each view with its perspective, every node naming the ogc command that prints it; served as static files next to the site; block 5 of every chapter links to it. Built on the branch explorer-appendix, merged after the gate. |
| 40 | C-46: the sponsor’s decision to interview or represent each affected population had no item of its own | M | ruled | In the example, one of the boundary items between the contract and the fulfilment of the evaluation contract covers the sponsor’s decision that the commuters needed to be interviewed while the residents could simply be represented. This is the kind of judgement call that affects the scope of work: whether expert representation is accepted or interviews and documented stakeholder-needs input are required. The example case splits across two affected stakeholder groups to show that both options are valid and traceable under the EPO. In practice representation is the most common, with stakeholder interviews reserved for extenuating circumstances because of the increased level of effort, generally for under-represented stakeholders or under-documented stakeholder needs. | Z (Michael Zargham), adjudicator | 2026-09-06 | Item kind StatementOfWork (SEVOCAB statement of work, p. 406; glossary term statement of work with altLabels SOW and scope of work), pinned at the contract, produced at the agree step by the sponsor, carried by two seams (to the recorder and to the account executive) and by a flow into the evaluation process’s scope step; one engagement decision per affected population, interview or representation (enum Engagement; epo:decides, epo:population, epo:engagement). Shape S0-StatementOfWork; S0-Population now requires the decision and that the record realizes it. The measles record: commuters interviewed, residents represented, by the county’s decision on 31 July. Executor: the agree template emits it; mutation engagement-mismatch is caught by S0-Population; counterexample engagement-mismatch.ttl. Contracting chapter says it in the specification and the walkthrough. |
| 41 | C-47: the rulings sat among the chapters as a page most readers will not want | L | ruled | The rulings move to an appendix. The concept is important, but most readers will not need it. The rulings are already part of the knowledge graph explorer; they fit best as Appendix C. The knowledge graph explorer allows navigation of the whole model and the computational proofs demonstrate the guarantees the model can prove computationally; Appendix C is the acknowledgement of the subjective choices, judgements and interpretations the model is grounded in. These subjectivities are implicit in the objective results the computational proofs produce, and the claim that this is scientific holds with these judgements as auxiliary assumptions. Appendix C is framed as the specification holding itself to its own definition of science. | Z (Michael Zargham), adjudicator | 2026-09-06 | docs/rulings.md becomes Appendix C, last in the toc after Appendix B, framed as Z says: the subjective choices the model is grounded in, implicit in the computational proofs, held as the auxiliary assumptions of the claim to science. The main path is seven pages; the vocabulary page’s map and the README say so. |
| 42 | C-48: the outline’s page names did not tell a skimming reader what the site contains | L | ruled | The page names are tightened: Glossary becomes Standards and Definitions; Contracting becomes Stakeholders and Contracting; The evaluation becomes Context and Evaluation; What the record proves becomes Records and Reporting; Conclusion stays. This gives a reader who skims the outline a much clearer picture of what the site contains. The appendices provide the first rung of backup and the repository as a whole an even richer one, so the reader need not take the authors on trust: the reader can verify. | Z (Michael Zargham), adjudicator | 2026-09-06 | Page titles renamed as Z lists them, in the site’s sentence case; file names and links unchanged. The nested model keeps its name. |
| 43 | C-49: the toolchain that makes the specification executable was nowhere reviewed for the reader | M | ruled | An Appendix D, Toolchain and Reproducibility, is added: a detailed review of the toolchain from virtual environments to the packages used, modelled after the portion of the SciPy proceedings supplementary material that covers the same subject. The reader should see that this is a computational paradigm, not just another document. | Z (Michael Zargham), adjudicator | 2026-09-06 | Appendix D, toolchain and reproducibility, after Appendix C: rendered from the lock file, the pinned converter digests, the gate script and the CI workflow, never typed by hand; a reviewer recipe modeled on the SciPy proceedings supplementary material (uv sync, the gate, ogc doctor, the site). |
| 44 | C-50: the nested lifecycle page was named for the model and did not say that the record rests on both cycles | L | ruled | The page title The nested model becomes A Nested Lifecycle: model is a throwaway word for most readers and should not trip anyone up. The nested lifecycle page must also address the fact that the records and reporting rely on information produced during both the contracting lifecycle and the evaluation lifecycle. | Z (Michael Zargham), adjudicator | 2026-09-06 | Page title A nested lifecycle (sentence case as the site’s titles are); its specification block now says which items each cycle contributes to the one record, that coverage and the traceback read both sides in one query, and that S0-Layers and M4-Nesting guarantee it; the walkthrough’s layer table is introduced as that count. |
| 45 | C-51: the toolchain review did not say why OpenSysML is outside the runtime environment, nor which ontologies the graphs are written in | L | ruled | OpenSysML is absent from the runtime environment because it was used only for modelling, through the features that render the models to RDF. The toolchain review therefore needs a section on ontologies, addressing PROV-O, EARL, SysML and the others. | Z (Michael Zargham), adjudicator | 2026-09-06 | Appendix D says OpenSysML is the modelling tool, used at authoring time to validate and to render the model to RDF, fetched by digest outside the lockfile, with everything downstream reading only the RDF; and gains an ontologies section, a register of every namespace the committed graphs use (W3C RDF, RDFS, OWL, XSD, SKOS, PROV-O, EARL, SHACL; OMG SysML v2 as OpenSysML renders it and its tool facts; the specification’s own namespaces) with the classes, properties and subjects each contributes, counted from the files; a namespace in use but not registered fails the test. |
| 46 | C-52: the layers figure on the nested lifecycle page was illegible | M | ruled | The layers diagram is illegible to a human and needs the treatment given to the earlier diagrams. Views are designed to communicate clearly as views, not to be exhaustive; this may require several diagrams, or it may be as simple as black-boxing subsystems more effectively. Wiring two systems together often means focusing on how they interface with each other, even when the interfaces span scales. | Z (Michael Zargham), adjudicator | 2026-09-06 | The layers poster is retired from the registry. The view nesting draws the outer chain with fulfil as a black box, the inner chain it opens into, and only the flows that cross the boundary, bundled by item kind and read from the model graph (the outer flows into the process’s parameters and the binds that hand them to inner steps); the page says what it leaves out. The rule joins the views doctrine: a view communicates one thing; black-box the rest; when wiring two systems, draw the interface. |
| 47 | C-53: round one of the simulated user test: fifteen findings that needed a judgment (rulings sheet 06) | H | ruled | Round one of the simulated user test is adjudicated item by item (sheet 06). The front page opens in three sentences and names the measles case early. The bridge cites the public Birds-of-a-Feather session at SciPy 2026 as a fact on the record rather than a private artefact; the forthcoming blog post may be added later. The chapters are reorganized: prose with headings that shape the outline, information blocks used sparingly, and code blocks that demonstrably pull content from the model, in the manner of the vendor fraud example. The rulings are rewritten in a formal register with their intent preserved and the message as sent kept in the graph. Terms are defined at first use, as the SciPy proceedings reviewers approved. The canonical source of outcome is SEVOCAB’s test result; NIST TN 1297 is a reserve source; fragment quotes say so; Popper 1959 stands alongside every crosswalk row. The four weak step matches are re-cited from process canon, since the posture is implementation of standards and not their re-litigation, the contribution being the application to contextual AI evaluation and the executable specification. The coined terms cite nothing and are owned. ogc reads the record and exposes the executor’s parameters. The thirteen pending quotes are ticked. Appendix D is aligned with the code it describes. | Z (Michael Zargham), adjudicator | 2026-09-06 | 06-14: the thirteen quotes ticked (human, Z, 2026-09-06); none pending. 06-06: outcome’s canonical is SEVOCAB test result, EARL a seeAlso and the binding. 06-07: NIST TN 1297 rank reserve. 06-08: six fragment quotes say so in their locators. 06-10: the four coined terms cite nothing; they carry ogc:coinedBy; the internal sources leave the register. 06-02: the crosswalk cites the public SciPy 2026 session (src:scipy-2026-bof, digest) instead of the deck; 06-11: Popper 1959 stands alongside every row (ogc:also). 06-01, 06-03, 06-05: the front page and the chapters restructured (prose and headings, judicious info blocks, code blocks that pull content, terms defined at first use). 06-04: the rulings rewritten in a formal register with the verbatim messages kept in the graph as ogc:verbatim. 06-09: process canon researched for C2, C4, C6 and step 2 and re-cited. 06-12: ogc reads the record. 06-13: ogc execute exposes the executor’s parameters. 06-15: Appendix D’s descriptions rendered from the commands’ own output. |
| 48 | C-54: round two of the simulated user test: nine findings needing a judgment, and three broader directions (rulings sheet 07) | H | ruled | Round two of the simulated user test is adjudicated (sheet 07). C-30 is narrowed to its human half, the block-by-block and wire-by-wire validation, the executor chapter being the computational half. The bridge’s second-layer row carries a definition in the register’s clinical words; the slogan that leaked in from the informal record belongs to the prose. Record is the abstract term and evaluation record its specialization, by subclassing; specializations are made explicit throughout the vocabulary, and hover definitions are used for clarity only, never for every term. The nesting view keeps its vertical alignment and shows that access to the test item comes from a different actor. The bridge table’s quotations move to a paragraph below it. The accountable organization’s duty to provide access is made explicit; whether accountable organization is the right headword is reconsidered. Citation traceability for the terms grounded outside the standards is reviewed and tightened with NIST and academic sources. References to the unpublished paper are purged: the site is the novel content and the paper will cite it. The open blocks and wires are walked through one at a time; sheet 05 is regenerated first, the model having been refined since. Status tags such as human and machine leave the primary tables for the explorer and the tool, on the principle that the graph’s detail is shared only where it helps the reader. The ontology’s terminological box is audited and cleaned: SKOS matches, subclassing and disjointness where they belong, disconnected content joined. The explorer gains a focus mode on a node’s neighbourhood. A BibTeX works-cited is kept: citations close each chapter and a new appendix aggregates them for the whole project. | Z (Michael Zargham), adjudicator | 2026-09-06 | 07-01: C-30 narrowed. 07-02: the crosswalk’s second-layer quote is a definition in the register’s words. 07-03: record and evaluation record by subclassing (tbox audit). 07-04: the nesting view labels each boundary wire with the part that supplies it. 07-05: the bridge’s quotations move below the table. 07-06: C4 carries ISO/IEC 17000 4.10 access as a neighbour and a scope note on the accountable organization’s duty; the headword goes to sheet 08. 07-07: citation traceability research. 07-08: references to the paper purged. 07-09: the block-and-wire walkthrough. Tags out of the primary tables. TBox audit, explorer focus mode, bibliography: separate slices of the plan. |
| 49 | C-55: the block-and-wire walkthrough (the human half of C-30): what each block’s ports and wires should be | H | ruled | The block-and-wire walkthrough of the model is adjudicated block by block (sheet 09). The sponsor organization, the domain expert, the evaluation operator and the test item are validated as wired. The account executive’s ports and wires are validated and its headword becomes authorized representative: the person is the organization’s interface with its contractual counterparties and is not responsible for the determinations in the report. The accountable organization’s headword becomes test item provider; it must provide access to the test item, and that access reaches the evaluation operator as well as the recorder; no further involvement of the provider is assumed, and the monitoring narrated in prose stays out of the wiring. Each affected population carries its engagement, interviewed or represented, and a responsible party, a person or an organization; the domain expert’s approval of the DSO release is preconditioned on every population’s input or representation being available. The report step is opened: a report assembler assembles the report with coverage and performance recomputed; the conformance checker records its verdict that the record and the report are correctly constructed, traceable and covered; a domain expert then approves the report’s contents, since correct construction alone risks correctly constructed nonsense; delivery requires that approval; coverage is also checked by a SHACL constraint on the report. The machine that administers the test plan is the test driver. The record is wired to the three machines only; people read it through views tailored to their roles, the record feeding forward through the report to the sponsor and, in prose, potentially to the test item provider. | Z (Michael Zargham), adjudicator | 2026-09-06 | Ticks in rulings/validated.ttl, rendered on sheet 05. Model: accessToOperatorSeam; TestDriver, ReportAssembler; ConformanceVerdict and ReportApproval items, ports and seams; the report step emits both and deliver consumes the approval; the three human recordIn ports and seams removed. Shapes: S7-ConformanceVerdict, S7-ReportApproval, S8-Delivery requires the approval, S7-Report recomputes coverage; the party and population model of the ontology audit (sheet 08). Headwords: authorized representative (closes C-26), test item provider, test driver; identifiers unchanged. The population block returns for validation with its new slots. |
| 50 | C-56: round three of the simulated user test: the six personas’ findings and the report step read one row at a time | H | ruled | Round three (sheet 10) is adjudicated item by item, and sheet 05 read one row at a time. Contract: the sponsor approves the requirement set (10-01); the report is the record’s delivered export; the record is auditable and reproducible but delivered only as a special case (10-02); acceptance is the sponsor’s act on receipt, recognizing completion, for which a correctly constructed record is necessary and not sufficient (10-03); the access grant carries the version identity, the period and the instrument (10-04); the agreement accepts the proposal and the statement of work, the access and the delivery are under the agreement (10-05); the sponsor has a signatory person (10-06); independence and user interest are declaration items (10-07); 17000’s own reading of party holds, the sponsor second-party, the independent testing organization third-party, the co-incidence cases named (10-08); the test item provider is accountable for the test item as its provider, the predicate is providesTestItem, and the sponsor’s obligation to the affected populations is not transferred (10-09). Evaluation: minPassRate is dropped for this version and weights carry a rationale and a timestamp (10-10); the report approval covers the recommendation, which rests on every attestation the coverage used or names its exclusions (10-11); appropriateness covers the assumptions and the test planning, and sufficiency relates the evidence to the claim: insufficient evidence implies cannot tell, sufficient evidence gives pass or fail (10-12); a person holds at most one actor role and any team member may represent a population (10-13); no cherry-picking: a determination is justified against all relevant evidence in the record (10-14); S5 mirrors M5 and the record states its independence level (10-15); deviations from the plan are items (10-16); the test item is bound across the record and records are linked by revision (10-17); the verdict and the coverage computation carry digests of the shapes, the ontology and the query (10-18); the ordering rules, the recomputed rates and the computed verdict are accepted as built (10-19); test driver stays, human and machine work complementary, separated in the test plan a human deems appropriate (10-20). Standards: 17000 4.10 leaves C4 for C2, with Annex A A.2 on C4 (10-21); the outer cycle names its canon per step and is performed under an agreement (10-22); attestation is refined, attesting passed, failed or cannot tell on specific evidence with supersession by later attestations (10-23); determination is refined as 17000’s decision (10-24); technical expert is refined: not all technical experts attest, every attestation comes from a relevant role (10-25); the evaluation operator is anchored to tester (10-26); authorized representatives exist per organization and no canonical definition is read reversed (10-27); ISO 9000:2026 stays a preview, checked 2026-09-07 (10-28); the SEVOCAB statement is recorded as printed (10-29); mission’s obligations towards the affected populations are marked as the specification’s facets of the organization’s purpose (10-30). Graphs: the record becomes a PROV bundle now (10-31); the register and the model vocabularies are declared with headers (10-32); the term map is realized in the model graph and the step derived from it now (10-33); the PROV subproperty on actsOnBehalfOf is dropped (10-34); the w3id namespace is assumed to merge (10-35); report and recommendation become terms and the machines and roles one paragraph (10-36); the retired word stays paraphrased (10-37); the site carries its version and commit (10-38); the gate prints a failing step’s cause (10-39); sources carry BibTeX keys and shapes their counterexamples (10-40); the conformance check is on the record and precedes the final report, a draft report may carry flagged gaps and a final report cannot, and the report assembler is rebuilt so (10-41). Sheet 05: B3, B9, B11, B12 and the wires W22, W30, W32, W34, W35, W36, W39 are validated; B8 awaits its rebuild, and concern C-30 narrows to it. The works cited name Hollek and Zargham for the SciPy 2026 session, Hollek first. | Z (Michael Zargham), adjudicator | 2026-09-07 | Sheet 10’s ruling column (10-01 to 10-41; the follow-on rulings of the same walkthrough are R-51); ticks in rulings/validated.ttl; the register’s ISO 9000 status and SEVOCAB retrieval note; the works cited author order. The model, ontology, shape, record and chapter changes the rulings call for follow in slices. |
| 51 | C-56: round three of the simulated user test: the six personas’ findings and the report step read one row at a time | H | ruled | Four rulings given during the sheet 10 walkthrough, after its items. The worked example is the measles evaluation, not the measles run: neither word is a glossary term, run reads at the execution scale (SEVOCAB’s test, concern C-14) and case at the scale of SEVOCAB’s test case, the anchor of probe, while the whole is one Contextual AI Evaluation and ISO/IEC 17000’s conformity assessment; the record’s namespace and file are renamed and the prose keeps walkthrough for the narrative (10-42). Synthetic content is tagged as such in the graph, on the measles evaluation’s bundle and every member, and the explorer can include or exclude synthetic data (10-43). A sample report for the synthetic measles evaluation becomes the first appendix: a simple dashboard rendered from the record alone, built with d3, that answers the sponsor’s question and does not overwhelm; the appendices are reordered and Records and reporting links to it as the sample report for the synthetic case (10-44). The SEVOCAB absence check behind every NIST and reserve canonical is recorded with its date: tester, session, measurement uncertainty, trajectory, appropriateness, sufficiency, attestation, red teaming, operational envelope and coverage are absent; ontology, repeatability, reproducibility and dialog are present in the specification’s sense and take SEVOCAB as canonical by the precedence rule, with Gruber, JCGM 200 and NIST AI 700-2 kept as neighbour cites; guardrail and probe are present in another sense and keep their canonicals (10-45). The source ranks are ordinal, 1 to 8, with no reserve or internal class: ISO 9000, SEVOCAB, NIST, W3C, the ISO-family vocabularies, other bodies, academic works, the authors (10-46). Each chapter keeps one references block, the generated sources cited with locators, and MyST’s compiled bibliography renders only in the works cited appendix; the prose sheds the source names the block carries (10-47). The sample report is redesigned first, positive and terse, then the synthetic case is redesigned to make the demo, with every criterion planned and attested; a final report requires full coverage, and only a flagged draft may carry gaps; the report’s primary reader is an executive who knows the domain and little about AI, the non-profit director or the head of an administration, and the page carries no technical jargon (10-48). The identifiers follow the headword authorized representative: the part def, the role individual and the term IRI are renamed, the alternative label account executive kept (10-49). | Z (Michael Zargham), adjudicator | 2026-09-07 | The record renamed; ogc:synthetic; the sample report appendix; the four canonicals moved and the SEVOCAB digest’s check rows. |
Open concerns, awaiting a ruling: 5.
| Concern | Severity | Surfaced | The question |
|---|---|---|---|
| C-25: first, second and third party: three adopted terms or one refined term | L | 2026-09-06 | ISO/IEC 17000 defines first-, second- and third-party conformity assessment activity as three headwords. The glossary adopts all three verbatim (no coinage); the alternative is one refined term, conformity assessment party, with the three as altLabels, which would be a fourth coinage. Z to confirm the three-term reading or rule otherwise. |
| C-30: the model is a draft until Z validates it block by block and wire by wire (the human half); the computational half, the end-state demonstration, is done | H | 2026-09-06 | Narrowed again by R-50 (sheet 10, 2026-09-07): every block and wire of sheet 05 is validated except B8, the report assembler, which is rebuilt so that the conformance check on the record precedes the final report; C-30 closes when B8 is ticked. Earlier: narrowed by R-48 (sheet 07-01): the executor chapter is the computational demonstration that the process yields a complete, traceable, coverable record for the runs it generates and that eight departures are each caught by a named check. What remains open is the human half: Z validates the model block by block and wire by wire on rulings sheet 05, regenerated from the model graph at every commit; the original wording, from the v0.2.0 tag, is kept in the register as history. |
| C-43: the S-shapes judge the items that exist; no shape requires a step’s output to exist | M | 2026-09-06 | The executor’s mutations show that a record without a plan approval, or without an access grant, conforms to S0 to S9: the shapes constrain a plan approval or an access grant once present but do not require one. Completeness against the model catches both, and the missing access also empties the traceback. Question for Z: should the S-shapes require the existence of each step’s outputs (a test plan approved by a domain expert before any session; an access grant behind every session), or is completeness against the model the right place for existence, keeping the shapes local to the items? |
| C-45: the model counterexample no-obligation also fires M5-Cardinality | L | 2026-09-06 | counterexamples/model/no-obligation.sysml is built to fail M1-Obligation and does, but its bare testing organization declares no account executive, so M5-Cardinality fires too; the test only asks that the named shape is among those fired. Either the counterexample gains an account executive so that it fails M1-Obligation alone, or the claim is read as sufficiency. |
| C-57: a final report requires full coverage (sheet 10-48): S7 does not yet require coverage one on a final report | M | 2026-09-07 | Z ruled that results without coverage do not reach a final report. The redesigned measles evaluation has coverage one and a draft report before its final one, and S7-Report recomputes a draft’s numbers over the attestations of its day, but no shape yet refuses a final report whose coverage is below one: the executor’s default run (two of three criteria planned) still emits a final report at two thirds. Closing it needs the executor to emit a draft, with its gaps flagged, when coverage is below one, and S7 to require coverage one and epo:draft false together; a change to the executor’s templates and its mutation expectations. |