Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Conclusion

Authors
Affiliations
Dynamical Systems Group
Humane Intelligence

The front page set out, in Popper’s terms as the authors stated them at their SciPy 2026 birds-of-a-feather session, what an evaluation would have to record to count as science: a hypothesis that could be shown false, the auxiliary assumptions held fixed, a prediction, evidence, at least one excluded outcome, and the context that says which assumption failed. The chapters between translated each of those into the terms of the engineering standards, stated the specification in those terms, first the contracting of an evaluation and then the evaluation itself as one nested model, and executed the process to show what the record proves. This page reads the bridge backwards: the first column names the item kinds the record holds and the shapes, the machine-checked rules, that check them.

What the record holds and the shapes checkIn the standard termsPopper’s element
AcceptanceCriterion, S2-AcceptanceCriterion, S2-Requirement: The expected result is written down before testing, so a determination can rule it not met.acceptance criteria, expected results, requirementhypothesis
Mission, Population, S0-Access, S0-Layers, S0-Parties, S0-Population, S0-StatementOfWork, ServiceAgreement, StatementOfWork, TestItemAccess: Pinned at the contract and timestamped before the requirement set; the evaluation takes them as given, documents them, and cannot change them.contract, customer, mission, provider, stakeholder, statement of work, test itemauxiliary assumption, first layer (pinned at the contract)
DsoRelease, RequirementSet, S1-DsoRelease, S2-RequirementSet: Declared and timestamped before testing; every attestation says whether that declared context was appropriate.Domain-Specific Ontology, operational envelope, operational environmentauxiliary assumption, second layer (pinned within the evaluation)
Probe, S3-Probe, S3-TestPlan, TestPlan, Turn: The plan names its objectives and means before any session; every probe names the criteria it exercises.expected results, probe, session, test planprediction
Determination, Evidence, S5-Evidence, S6-Determination: Evidence derives from a recorded response at a recorded turn; a determination rules only on evidence that bears on the criterion it tests.determination, objective evidence, trajectoryevidence
Attestation, Report, S6-Attestation, S7-Report: A failed outcome is a recorded falsifier; coverage is recomputed from the record, and a criterion with no attestation counts zero.attestation, outcome, test coveragefalsifiability
S6-Attestation, S8-Recommendation: A failed prediction is attributed: the record says which named person judged the context inappropriate or the evidence insufficient, and why.appropriateness, requirements traceability, sufficiencycontext

Read in this direction, the claim is exact. The thirteen essentials (SCI-01 to SCI-13, the statements the specification must keep), the shapes that check the model and the record, and the record itself encode what qualifies as scientific about an evaluation, and nothing else. An evaluation that follows the process leaves a record, and the executor shows it for the runs it generates: each conforms, is complete, has a coverage anyone can recompute and traces fully, and each of twelve ways of departing from the wiring is caught by a named check. That every possible run must do so is the open concern C-30, stated in Appendix D, not a claim made here. A record that conforms to the shapes shows the hypothesis was stated before the test, the assumptions were declared and judged, the prediction was made and observed, the evidence was ruled on by a named person, the excluded outcomes were counted, and the context was attributed. Nothing in the record says whether the experts were right; that is theirs, and it is recorded with their names.

1Where this goes

Two offers close this specification. Use OG-CAIE: the process, the vocabulary and the record format are open, and the measles example shows the whole chain across two chapters. Or have your own AI evaluation practice audited against it: every requirement here is checkable, so an existing practice can be walked through the thirteen essentials and shown where its record would and would not conform.

2Sources cited

References
  1. International Organization for Standardization. (2026). ISO 9000:2026 Quality management — Fundamentals and vocabulary. International Organization for Standardization. https://www.iso.org/obp/ui/#iso:std:iso:9000:ed-5:v1:en
  2. IEEE Computer Society and ISO/IEC JTC 1/SC 7. (2026). IEEE Computer Society, Software and Systems Engineering Vocabulary (SEVOCAB), PDF export created 2026-09-02 (481 pp.). IEEE Computer Society and ISO/IEC JTC 1/SC 7. https://www.computer.org/sevocab
  3. Amironesei, R., Godil, A., Greenberg, C., Greene, K., Hall, P., Jensen, T., Fiscus, J., & Schulman, N. (2025). Assessing Risks and Impacts of AI (ARIA): ARIA 0.1 Pilot Evaluation Report (Techreport NIST AI 700-2). National Institute of Standards and Technology. 10.6028/NIST.AI.700-2
  4. W3C. (2017). Evaluation and Report Language (EARL) 1.0 Schema. W3C. https://www.w3.org/TR/2017/NOTE-EARL10-Schema-20170202/
  5. IEC. (2013). IEC 60050-351:2013 International Electrotechnical Vocabulary, Part 351: Control technology. IEC. https://www.electropedia.org/iev/iev.nsf/index?openform=&part=351
  6. ISO/IEC. (2020). ISO/IEC 17000:2020 Conformity assessment — Vocabulary and general principles. ISO/IEC. https://www.iso.org/obp/ui/en/#iso:std:73029:en
  7. JCGM. (2012). JCGM 200:2012 International vocabulary of metrology, basic and general concepts and associated terms (VIM), 3rd edition. BIPM. 10.59161/jcgm200-2012
  8. INCOSE. (2023). Guide to Writing Requirements v4, Summary Sheet. INCOSE.
  9. International Auditing and Assurance Standards Board. (2009). International Standard on Auditing 500: Audit Evidence. International Auditing and Assurance Standards Board.
  10. Hawkins, R., Kelly, T., Knight, J., & Graydon, P. (2011). A New Approach to Creating Clear Safety Arguments. In C. Dale & T. Anderson (Eds.), Advances in Systems Safety: Proceedings of the Nineteenth Safety-Critical Systems Symposium (pp. 3–23). Springer. 10.1007/978-0-85729-133-2_1
  11. Popper, K. R. (1959). The Logic of Scientific Discovery. Hutchinson.
  12. Hollek, J., & Zargham, M. (2026). Building Scientific Approaches to Generative AI. Birds-of-a-Feather session, SciPy 2026.