The front page set out, in Popper’s terms as the authors stated them at their SciPy 2026 birds-of-a-feather session, what an evaluation would have to record to count as science: a hypothesis that could be shown false, the auxiliary assumptions held fixed, a prediction, evidence, at least one excluded outcome, and the context that says which assumption failed. The chapters between translated each of those into the terms of the engineering standards, stated the specification in those terms, first the contracting of an evaluation and then the evaluation itself as one nested model, and executed the process to show what the record proves. This page reads the bridge backwards: the first column names the item kinds the record holds and the shapes, the machine-checked rules, that check them.
| What the record holds and the shapes check | In the standard terms | Popper’s element |
|---|---|---|
AcceptanceCriterion, S2-AcceptanceCriterion, S2-Requirement: The expected result is written down before testing, so a determination can rule it not met. | acceptance criteria, expected results, requirement | hypothesis |
Mission, Population, S0-Access, S0-Layers, S0-Parties, S0-Population, S0-StatementOfWork, ServiceAgreement, StatementOfWork, TestItemAccess: Pinned at the contract and timestamped before the requirement set; the evaluation takes them as given, documents them, and cannot change them. | contract, customer, mission, provider, stakeholder, statement of work, test item | auxiliary assumption, first layer (pinned at the contract) |
DsoRelease, RequirementSet, S1-DsoRelease, S2-RequirementSet: Declared and timestamped before testing; every attestation says whether that declared context was appropriate. | Domain-Specific Ontology, operational envelope, operational environment | auxiliary assumption, second layer (pinned within the evaluation) |
Probe, S3-Probe, S3-TestPlan, TestPlan, Turn: The plan names its objectives and means before any session; every probe names the criteria it exercises. | expected results, probe, session, test plan | prediction |
Determination, Evidence, S5-Evidence, S6-Determination: Evidence derives from a recorded response at a recorded turn; a determination rules only on evidence that bears on the criterion it tests. | determination, objective evidence, trajectory | evidence |
Attestation, Report, S6-Attestation, S7-Report: A failed outcome is a recorded falsifier; coverage is recomputed from the record, and a criterion with no attestation counts zero. | attestation, outcome, test coverage | falsifiability |
S6-Attestation, S8-Recommendation: A failed prediction is attributed: the record says which named person judged the context inappropriate or the evidence insufficient, and why. | appropriateness, requirements traceability, sufficiency | context |
Read in this direction, the claim is exact. The thirteen essentials (SCI-01 to SCI-13, the statements the specification must keep), the shapes that check the model and the record, and the record itself encode what qualifies as scientific about an evaluation, and nothing else. An evaluation that follows the process leaves a record, and the executor shows it for the runs it generates: each conforms, is complete, has a coverage anyone can recompute and traces fully, and each of twelve ways of departing from the wiring is caught by a named check. That every possible run must do so is the open concern C-30, stated in Appendix D, not a claim made here. A record that conforms to the shapes shows the hypothesis was stated before the test, the assumptions were declared and judged, the prediction was made and observed, the evidence was ruled on by a named person, the excluded outcomes were counted, and the context was attributed. Nothing in the record says whether the experts were right; that is theirs, and it is recorded with their names.
1Where this goes¶
Two offers close this specification. Use OG-CAIE: the process, the vocabulary and the record format are open, and the measles example shows the whole chain across two chapters. Or have your own AI evaluation practice audited against it: every requirement here is checkable, so an existing practice can be walked through the thirteen essentials and shown where its record would and would not conform.
2Sources cited¶
ISO, ISO 9000:2026 Quality management — Fundamentals and vocabulary (3.1.4; 3.1.9; 3.11.1; 3.11.2 review; 3.3.13; 3.4.11; 3.4.5 policy; 3.5.1; 3.5.1 requirement; 3.5.11 traceability; 3.8.6; 3.9.1) International Organization for Standardization, 2026
IEEE Computer Society, Software and Systems Engineering Vocabulary (SEVOCAB), PDF export created 2026-09-02 (481 pp.) (acceptance criteria, p. 4 (ISO/IEC 33202:2024, 3.1); artificial intelligence-based system, p. 28 (ISO/IEC TR 29119-11:2020); expected results, p. 160 (ISO/IEC/IEEE 29119-4:2021, 3.32); objective evidence, p. 277 (ISO/IEC 33001:2015, 3.2.13); operational environment, p. 282 (IEEE 982-2024, 3.1); requirements traceability, p. 352 (ISO/IEC/IEEE 29148:2018); stakeholder, p. 403 (ISO/IEC/IEEE 12207:2026, 3.1.59; 15288:2023, 3.44); statement of work, p. 406 (ISO/IEC 33202:2024, 3.25); system under test (SUT), p. 423 (ISO/IEC 14756:1999); system-of-interest (SOI), p. 423 (ISO/IEC/IEEE 15288:2023); test case, p. 432 (fragment of the entry); test coverage, p. 433 (ISO/IEC/IEEE 29119-2:2021, 3.28); test item, p. 435 (ISO/IEC/IEEE 29119-2:2021); test plan, p. 437 (ISO/IEC/IEEE 29119-2:2021, 3.50); test result, p. 438 (ISO/IEC/IEEE 29119-2:2021, 3.56)) IEEE Computer Society and ISO/IEC JTC 1/SC 7, 2026
NIST, Assessing Risks and Impacts of AI (ARIA): ARIA 0.1 Pilot Evaluation Report (Appendix A, Application, p. 16 (fragment of the entry); Appendix A, Session, p. 17 (first sentence of the entry); Section 3.1 Validity Risk Assessment, p. 5 (one of the four annotation values, a fragment)) Amironesei et al., 2025
W3C, Evaluation and Report Language (EARL) 1.0 Schema (Section 2.7 OutcomeValue Class) W3C, 2017
IEC, IEC 60050-351:2013 International Electrotechnical Vocabulary, Part 351: Control technology (351-41-08 state variable, Note 1; 351-41-10 trajectory) IEC, 2013
ISO/IEC, ISO/IEC 17000:2020 Conformity assessment — Vocabulary and general principles (4.2 object of conformity assessment; 7.3) ISO/IEC, 2020
JCGM, JCGM 200:2012 International vocabulary of metrology, basic and general concepts and associated terms (VIM), 3rd edition (2.41 metrological traceability, p. 45) JCGM, 2012
INCOSE, Guide to Writing Requirements v4, Summary Sheet (C7 Verifiable, p. 2) INCOSE, 2023
IAASB, International Standard on Auditing 500: Audit Evidence (paragraph 5(b) (fragment of the paragraph); paragraph 5(f) (first sentence of the paragraph)) International Auditing and Assurance Standards Board, 2009
Hawkins, Kelly, Knight and Graydon, A New Approach to Creating Clear Safety Arguments (Section 3.2 Asserted context, p. 7; Section 3.3 Asserted solution, p. 10) Hawkins et al., 2011
Popper, The Logic of Scientific Discovery (as the source of those definitions) Popper, 1959
Hollek and Zargham, Building Scientific Approaches to Generative AI (the authors’ definition for this specification (sheet 07-02); the slide read ‘Context Matters!’; the definitions presented at the session) Hollek & Zargham, 2026
- International Organization for Standardization. (2026). ISO 9000:2026 Quality management — Fundamentals and vocabulary. International Organization for Standardization. https://www.iso.org/obp/ui/#iso:std:iso:9000:ed-5:v1:en
- IEEE Computer Society and ISO/IEC JTC 1/SC 7. (2026). IEEE Computer Society, Software and Systems Engineering Vocabulary (SEVOCAB), PDF export created 2026-09-02 (481 pp.). IEEE Computer Society and ISO/IEC JTC 1/SC 7. https://www.computer.org/sevocab
- Amironesei, R., Godil, A., Greenberg, C., Greene, K., Hall, P., Jensen, T., Fiscus, J., & Schulman, N. (2025). Assessing Risks and Impacts of AI (ARIA): ARIA 0.1 Pilot Evaluation Report (Techreport NIST AI 700-2). National Institute of Standards and Technology. 10.6028/NIST.AI.700-2
- W3C. (2017). Evaluation and Report Language (EARL) 1.0 Schema. W3C. https://www.w3.org/TR/2017/NOTE-EARL10-Schema-20170202/
- IEC. (2013). IEC 60050-351:2013 International Electrotechnical Vocabulary, Part 351: Control technology. IEC. https://www.electropedia.org/iev/iev.nsf/index?openform=&part=351
- ISO/IEC. (2020). ISO/IEC 17000:2020 Conformity assessment — Vocabulary and general principles. ISO/IEC. https://www.iso.org/obp/ui/en/#iso:std:73029:en
- JCGM. (2012). JCGM 200:2012 International vocabulary of metrology, basic and general concepts and associated terms (VIM), 3rd edition. BIPM. 10.59161/jcgm200-2012
- INCOSE. (2023). Guide to Writing Requirements v4, Summary Sheet. INCOSE.
- International Auditing and Assurance Standards Board. (2009). International Standard on Auditing 500: Audit Evidence. International Auditing and Assurance Standards Board.
- Hawkins, R., Kelly, T., Knight, J., & Graydon, P. (2011). A New Approach to Creating Clear Safety Arguments. In C. Dale & T. Anderson (Eds.), Advances in Systems Safety: Proceedings of the Nineteenth Safety-Critical Systems Symposium (pp. 3–23). Springer. 10.1007/978-0-85729-133-2_1
- Popper, K. R. (1959). The Logic of Scientific Discovery. Hutchinson.
- Hollek, J., & Zargham, M. (2026). Building Scientific Approaches to Generative AI. Birds-of-a-Feather session, SciPy 2026.