Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

OG-CAIE as an executable specification

Authors
Affiliations
Dynamical Systems Group
Humane Intelligence
Humane Intelligence and Dynamical Systems Group

A collaboration between Humane Intelligence and Dynamical Systems Group.

Contextual AI Evaluation (CAIE) is the evaluation of a deployed AI system against the needs of one domain, judged by people who know that domain. OG-CAIE is CAIE performed with the method this site sets out: an Evaluation Process Ontology (EPO), the same six steps for every domain, performed inside a contracting lifecycle drawn from the engineering standards, and a Domain-Specific Ontology (DSO), the vocabulary of one domain supplied or approved by its experts. The method is stated here in a form a machine can check and a person can read, and it is walked through on one synthetic case, a county public-health chatbot asked about measles during an outbreak.

This site is a collaboration between Dynamical Systems Group and Humane Intelligence, co-authored by Michael Zargham and Julie Hollek.

1Why this counts as science

Karl Popper’s account of empirical science is short. A hypothesis is a claim that could be shown false. It is tested under auxiliary assumptions, the background held fixed while the test runs. Together they yield a prediction: what should be observed if both are true. Evidence is an accepted result of observation. A claim is scientific to the extent that it excludes at least one possible outcome; and context matters, because a failed prediction does not by itself say which assumption failed.

An evaluation of an AI system counts as science on the same terms. It must write down, before any test, what would count as failure and under what assumptions; it must collect what the system actually did; it must record who ruled on that evidence, and whether the assumptions still held; and it must leave a record from which anyone can recompute what was covered and what was not. That is the whole standard. Everything on this site exists to meet it.

2The bridge to the engineering standards for evaluation

None of this needs new words. Each element of that account already has a settled name in the standards that govern testing, conformity assessment and quality: ISO 9000 for requirement, objective evidence and determination; ISO/IEC 17000 for attestation; the IEEE Software and Systems Engineering Vocabulary for acceptance criteria, expected results, test plan, test coverage and the test case a probe refines; NIST’s evaluation reports for session. Using those names rather than coining our own lets someone else check the record against the same definitions. At SciPy 2026 the authors led a birds-of-a-feather session, Building Scientific Approaches to Generative AI, and put Popper’s elements to the room in the words the table’s first column keeps. The table is the bridge: each row takes one of those elements to the standard terms it lands on, to the place it occupies in an evaluation record, and to the check that makes it more than a promise.

Popper’s element, as presented at the SciPy 2026 sessionThe standard terms it lands onWhere it lives in the recordWhat makes it checkable
hypothesisacceptance criteria, expected results, requirementAn acceptance criterion derived from a requirement, stating the expected result a test could observe.The expected result is written down before testing, so a determination can rule it not met.
auxiliary assumption, first layer (pinned at the contract)contract, customer, mission, provider, stakeholder, statement of work, test itemThe sponsor’s mission towards the affected populations, the service agreement with its statement of work, and the access to the test item: the counterparties, the object, the requirements’ frame and the method, fixed before the evaluation begins.Pinned at the contract and timestamped before the requirement set; the evaluation takes them as given, documents them, and cannot change them.
auxiliary assumption, second layer (pinned within the evaluation)Domain-Specific Ontology, operational envelope, operational environmentThe requirement set declaring this test item in this environment, and the DSO release, both declared and approved before any session.Declared and timestamped before testing; every attestation says whether that declared context was appropriate.
predictionexpected results, probe, session, test planA probe applied at a turn of a session under the test plan, against the criterion’s expected result.The plan names its objectives and means before any session; every probe names the criteria it exercises.
evidencedetermination, objective evidence, trajectoryThe response or trajectory collected as evidence bearing on one criterion; the determination is the accepted result: passed, failed or cannot tell.Evidence derives from a recorded response at a recorded turn; a determination rules only on evidence that bears on the criterion it tests.
falsifiabilityattestation, outcome, test coverageThe attestation’s outcome for each criterion, and the report’s coverage over all of them.A failed outcome is a recorded falsifier; coverage is recomputed from the record, and a criterion with no attestation counts zero.
contextappropriateness, requirements traceability, sufficiencyThe appropriateness and sufficiency judgments on every attestation, and the trace from each recommendation back to the requirement it answers.A failed prediction is attributed: the record says which named person judged the context inappropriate or the evidence insufficient, and why.

The elements in the first column, as the authors defined them at their SciPy 2026 birds-of-a-feather session, after Popper (1959), except the second-layer row, defined for this specification (sheet 07-02):

The next chapter gives those terms their canonical definitions, verbatim and cited. The chapters after it use them to state the specification, first the contracting of an evaluation and then the evaluation itself, and to show its properties. The conclusion returns to Popper’s terms and says what has been encoded.

Version unreleased, rendered at commit 6e53610: the specification’s version is its git tag and its commit (sheet 10-38).

3Sources cited

References
  1. International Organization for Standardization. (2026). ISO 9000:2026 Quality management — Fundamentals and vocabulary. International Organization for Standardization. https://www.iso.org/obp/ui/#iso:std:iso:9000:ed-5:v1:en
  2. IEEE Computer Society and ISO/IEC JTC 1/SC 7. (2026). IEEE Computer Society, Software and Systems Engineering Vocabulary (SEVOCAB), PDF export created 2026-09-02 (481 pp.). IEEE Computer Society and ISO/IEC JTC 1/SC 7. https://www.computer.org/sevocab
  3. Artificial Intelligence Risk Management Framework (AI RMF 1.0) (Techreport NIST AI 100-1). (2023). NIST. 10.6028/NIST.AI.100-1
  4. Amironesei, R., Godil, A., Greenberg, C., Greene, K., Hall, P., Jensen, T., Fiscus, J., & Schulman, N. (2025). Assessing Risks and Impacts of AI (ARIA): ARIA 0.1 Pilot Evaluation Report (Techreport NIST AI 700-2). National Institute of Standards and Technology. 10.6028/NIST.AI.700-2
  5. W3C. (2017). Evaluation and Report Language (EARL) 1.0 Schema. W3C. https://www.w3.org/TR/2017/NOTE-EARL10-Schema-20170202/
  6. IEC. (2013). IEC 60050-351:2013 International Electrotechnical Vocabulary, Part 351: Control technology. IEC. https://www.electropedia.org/iev/iev.nsf/index?openform=&part=351
  7. ISO/IEC. (2020). ISO/IEC 17000:2020 Conformity assessment — Vocabulary and general principles. ISO/IEC. https://www.iso.org/obp/ui/en/#iso:std:73029:en
  8. JCGM. (2012). JCGM 200:2012 International vocabulary of metrology, basic and general concepts and associated terms (VIM), 3rd edition. BIPM. 10.59161/jcgm200-2012
  9. INCOSE. (2023). Guide to Writing Requirements v4, Summary Sheet. INCOSE.
  10. International Auditing and Assurance Standards Board. (2009). International Standard on Auditing 500: Audit Evidence. International Auditing and Assurance Standards Board.
  11. SEBoK Editorial Board. (2026). Guide to the Systems Engineering Body of Knowledge (SEBoK) v2.14, released 18 May 2026 (N. Hutchison, Ed.). The Trustees of the Stevens Institute of Technology. https://sebokwiki.org/
  12. Hawkins, R., Kelly, T., Knight, J., & Graydon, P. (2011). A New Approach to Creating Clear Safety Arguments. In C. Dale & T. Anderson (Eds.), Advances in Systems Safety: Proceedings of the Nineteenth Safety-Critical Systems Symposium (pp. 3–23). Springer. 10.1007/978-0-85729-133-2_1
  13. Popper, K. R. (1959). The Logic of Scientific Discovery. Hutchinson.
  14. Hollek, J., & Zargham, M. (2026). Building Scientific Approaches to Generative AI. Birds-of-a-Feather session, SciPy 2026.