Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Standards and definitions

Authors
Affiliations
Dynamical Systems Group
Humane Intelligence

The front page ends at the bridge from a plain account of science to the engineering standards for evaluation. This chapter gives the terms on the far side of that bridge their definitions: rigorous, cited verbatim, and not exhaustive: exactly the terms this site uses; the full register is in the repository and the explorer.

1The map of the site

Seven pages, each a view of the model in the repository, and six appendices: A is the sample report, B opens the same model as a knowledge graph, C runs each chapter’s checks, D is the Rulings the model is grounded in, E reviews the toolchain and the reproduction, and F the works cited. The front page sets the bar for science in plain terms and crosses the bridge into the standards; this page gives the terms and the two cycles; Stakeholders and contracting is the outer cycle, from need to acceptance; Context and evaluation is the inner cycle, scope to report; A nested lifecycle shows why the two are one model; Records and reporting runs the checks and shows what the record proves; the Conclusion returns to the front page’s terms. Each chapter from the contracting on keeps one outline: what the standards say, the specification, the walkthrough, checked, and there is more in the model, with at least one block on each page showing an ogc command and what it prints.

2Vocabulary discipline

Terms are used, not owned. Each term cites exactly one canonical definition, chosen from the highest-ranked source that defines it in the sense used: ISO 9000:2026, Quality management: fundamentals and vocabulary, for every quality and process term; SEVOCAB, the IEEE Computer Society’s Software and Systems Engineering Vocabulary, for the systems-engineering and testing terms ISO 9000 lacks; NIST AI 700-2, the report of the ARIA pilot (Assessing Risks and Impacts of AI), for the AI-evaluation terms; and the W3C specifications for the technical binding only, never for a narrative definition: PROV-O for who did what and when, EARL for what was asserted and with what outcome, SHACL (the Shapes Constraint Language, in which every machine check on this site is written) and SKOS for the glossary itself. The other sources follow in ordinal ranks, which say whose definition wins when several define a term.

Exactly four terms are coined, with their shorthands: Contextual AI Evaluation (CAIE), OG-CAIE, Evaluation Process Ontology (EPO) and Domain-Specific Ontology (DSO). Everything else is grounded in cited standards and literature. Adopted terms are used exactly as the source defines them; refined terms carry a typed anchor to a standard term and keep the source’s word as an alternative label. One word is used against the grain of its source, and says so: conformance, deprecated by ISO 9000 as a synonym of conformity, is reclaimed for the machine-checked, correctly constructed evaluation record.

Every quote on this site carries a tag. Machine: the tests locate the quote in a content-hashed snapshot of the source or, where the source is held locally and not redistributed, in its committed digest. Human: a named person verified it against the source on a date, usually an ISO screenshot held locally. Pending: transcribed and awaiting that person’s tick, listed on a rulings sheet; a pending quote is cited but not yet verified. Authors: the authors’ own words on the public record, as the bridge’s definitions presented at the SciPy 2026 session. Cite-only: a locator that carries no quote, a neighbour cited for where it stands, never for its words. At this commit: 94 machine, 56 human, 0 pending, 7 authors, 6 cite-only.

3How to navigate

Hover any highlighted term on any page to see its definition. The entries below are exactly those terms: the narrative definition, then the canonical source with its locator and verbatim quote, then the alternative labels and the ruling the term rests on. Terms the prose does not reach, such as authoritative reference, stay in the full register.

The vocabulary graph answers directly from the command line: uv run -q ogc term probe gives one entry with its citations, rulings and the essentials it is stated in (the thirteen statements the specification must keep, SCI-01 to SCI-13); ogc define, ogc quote and ogc verify give the definition, the verbatim quotes and where each was found; ogc find searches labels and quotes; ogc check-word says whether a word is a headword, an alternative label or retired, and what to write; ogc sci, ogc rulings, ogc concerns and ogc sources read the rest of the record, ogc crosswalk prints the anchor table or, with a flag, the bridge rows of the front page, and ogc sparql takes a read-only query. Every answer starts with the command and the commit it was read at, shown on this site as <sha>. The tool follows the one the authors built for their earlier glossaries (ruling R-29). One answer, for the word this site’s prose is allowed to use for the test item:

check-word-system-under-test.md
$ uv run -q ogc check-word 'system under test'
# ogc check-word 'system under test' @ <sha>
system under test: registered as alt label 'system under test' of 'test item' (test-item); class
    adopted
  sense: The deployed AI system being evaluated, taken as a black box: only what it is given and
      what it produces are recorded.
  canonical: sevocab test item, p. 435 (ISO/IEC/IEEE 29119-2:2021) [machine]
  concern: C-10 (ruled) system under test: weak canonical entry
  concern: C-15 (ruled) operational envelope: what it adds to a set of requirements
  concern: C-21 (ruled) parties to the evaluation are implicit
  -> an alternate label; the headword is 'test item': write {term}`system under test <test
      item>` in prose
(exit 0)
acceptance criteria
A specific, checkable condition a response must meet for a requirement to count as satisfied. The unit that coverage is measured over: each criterion either traces to an attested outcome or it does not. Source: IEEE Computer Society, acceptance criteria, p. 4 (ISO/IEC 33202:2024, 3.1): “criteria that a system or component is required to satisfy to be accepted by a user, customer, or other authorized entity”. Also: acceptance criterion.
appropriateness
The expert judgment that the declared context (operational environment, requirement set and Domain-Specific Ontology) was the right frame for the judgment being made. Not whether the system passed, and not whether there was enough evidence. Source: Hawkins, Section 3.2 Asserted context, p. 7: “it is being asserted that the context is appropriate for the argument elements to which it applies”. Ruling R-08.
attestation
The recorded judgment of a named person that links one acceptance criterion to the evidence reviewed and records an outcome, together with the judgments of appropriateness of the declared context and sufficiency of the evidence. Source: ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 7.3: “issue of a statement, based on a decision, that fulfilment of specified requirements has been demonstrated”. Also: annotation. Ruling R-50.
conformance
The record is correctly constructed: it satisfies every rule the Evaluation Process Ontology’s shapes state about how an evaluation must be recorded. A mechanical yes or no, produced by machine. If the record conforms, the process was followed, its data is shaped, its required fields are filled, and coverage can be computed. A special case of conformity: fulfilment of the record-keeping requirements. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.5.9 conformity, Note 1: “The term “conformance” is synonymous but deprecated.”. Also: correct construction, correctness by construction. Ruling R-16.
conformity
Fulfilment of a requirement, in the ISO sense: the system under test conforms to an acceptance criterion when its response fulfils it, as a named person attests. For the machine-checked shape of the record itself, see conformance. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.5.9: “fulfilment of a requirement”. Ruling R-16.
Contextual AI Evaluation
The evaluation of a deployed AI system against the needs of a specific domain and deployment context, judged by people who know that domain. The class of evaluation this specification is for; OG-CAIE is the way it is performed here. Coined by the authors for this specification; it cites no source. Also: CAIE. Ruling R-29.
contract
The binding agreement between the sponsor organization and the testing organization to perform OG-CAIE as a service. It is signed for the sponsor by its signatory and for the testing organization by its authorized representative, precedes the requirement set and follows the mission, the need and the proposal in the record. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.3.13: “binding agreement”. Also: agreement, service agreement. Ruling R-21.
customer
A person or organization that receives a product or a service. Two products change hands around an OG-CAIE evaluation, so there are two customers: the evaluation customer, the sponsor, receives the evaluation as a service; the test item customer receives the AI system under test, as the sponsor sometimes does and as the affected populations do when they use it. The two may be the same organization, and the record says whether they are. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.9.1: “person or organization that can or does receive a product or a service that is intended for or required by this person or organization”. Ruling R-21.
determination
The ruling, made on one or more evidence items, that a criterion’s expected result was met, was not met, or cannot be told. It is the ruling made on evidence, recorded with an EARL outcome of passed, failed or cannot tell, made by machine or by a named person, and it is what an attestation aggregates for a criterion. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.11.1: “activity to find out one or more characteristics and their characteristic values”. Also: judgment. Ruling R-20, R-50.
dialogue
The ordered prompts to, and responses by, the system under test within one session. Its recorded form is the trajectory. Source: IEEE Computer Society, dialog, p. 129 (ISO TR 25060:2023, 2.4): “interaction between a user and an interactive system as a sequence of user actions (inputs) and system responses (outputs) in order to achieve a goal”. Also: conversation, dialog. Ruling R-13, R-51.
Domain-Specific Ontology
The vocabulary and relations of the problem domain, imported from an existing ontology or supplied by domain experts, that equip one evaluation with the concepts it needs. A domain expert supplies or approves the release used. It changes when the domain changes. Coined by the authors for this specification; it cites no source. Also: DSO. Ruling R-09.
evaluation
Systematically determining how far something meets its specified criteria. An OG-CAIE evaluation determines, for a declared requirement set, which criteria the system under test met, and reports coverage and performance separately. Source: IEEE Computer Society, evaluation, p. 155 (ISO/IEC 25001:2014, 4.1): “systematic determination of the extent to which an entity meets its specified criteria”. Also: AI evaluation.
evaluation operator
The technical expert who operates the evaluation: declares the requirement set with the sponsor, writes the test plan, administers the probes to the test item, collects the responses and the evidence per the plan, records any departure from the plan with its reason, and writes the recommendation. The operator makes determinations but does not attest; attestation belongs to the domain expert. An abstract actor category, filled by a named person in each evaluation. Source: NIST AI 700-2, Appendix A, Tester, p. 17 (first sentence of the entry): “Individual who interacts with an application within the ARIA test environment.”. Also: AI evaluation expert, operator, red teamer, tester. Ruling R-23, R-50.
Evaluation Process Ontology
The reusable, domain-independent description of how a rigorous evaluation is run: the standard operating procedure whose six steps, performed inside the contracting lifecycle, make a coverage claim checkable when followed. It assumes a Domain-Specific Ontology is in place and does not change across domains. Coined by the authors for this specification; it cites no source. Also: EPO. Ruling R-09, R-31.
evaluation record
The record of one evaluation, bound by the Evaluation Process Ontology: the DSO release, the requirement set, the test plan, the probes, the sessions and their trajectories, the evidence, the attestations, the report and the recommendations, with who did what and when. The EPO determines what must be present for requirements traceability in both directions and for coverage, and makes all three checkable by machine; a conformant record therefore provides both traceability directions and coverage. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.8.12 record: “document stating results achieved or providing evidence of activities performed”. Also: record (ISO 9000:2026 3.8.12, of one evaluation), test log. Ruling R-17.
expected results
What the system under test should observably do if it meets an acceptance criterion under the probe’s conditions. Stated on the criterion before testing, so the criterion is defined in terms of evidence a test can collect, and the attestation is a judgment that the actual result did or did not correspond. Source: IEEE Computer Society, expected results, p. 160 (ISO/IEC/IEEE 29119-4:2021, 3.32): “observable predicted behavior of the test item under specified conditions based on its specification or another source”. Also: expected result. Ruling R-12.
first-party conformity assessment activity
Evaluation performed by the organization that provides or is accountable for the test item: the test item provider. In OG-CAIE the test item provider is a party to every evaluation, since it grants access to the test item, but it does not perform the evaluation unless it is also the sponsor and the testing organization, and the record says so. Source: ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 4.3 first-party conformity assessment activity: “conformity assessment activity that is performed by the person or organization that provides or that is the object of conformity assessment”. Also: first party. Ruling R-21.
guardrail
A requirement on the system stating what may be shared with a user and what must be withheld. Recommendations at the end of an evaluation typically propose changes to guardrails. Source: NIST AI 700-2, Appendix A, Guardrail, p. 17 (first sentence of the entry): “An application requirement specifying both 1) permitted information that can be shared with a user, and 2) prohibited information that should be withheld from a user.”.
measurement uncertainty
A non-negative number that says how dispersed the values attributed to a measured quantity are. A probe pass rate estimated from replicate sessions carries one; coverage, counted exactly over the record, does not. Source: JCGM 200:2012 International vocabulary of metrology, 2.26 measurement uncertainty, p. 41: “non-negative parameter characterizing the dispersion of the quantity values being attributed to a measurand, based on the information used”. Also: uncertainty. Ruling R-13.
mission
The sponsor organization’s purpose for existing, as its top management expresses it, and with it the obligations and duties it holds towards the affected populations. Recorded before the need is stated, it is the context the contract answers to: what the sponsor owes the people the test item will serve. Pinned at the contract. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.4.11: “organization’s purpose for existing as expressed by top management”. Also: mandate, obligations to the affected populations. Ruling R-37, R-50.
monitoring
Determining the status of a system at different stages or times. After an evaluation, a sponsor may require the test item provider to change the system per the findings and then require new testing to certify that the flagged issues were addressed; in conformity assessment that repeat is called surveillance. Observed in practice; not a step of the contracting lifecycle. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.11.3: “determining the status of a system, a process or an activity”. Also: re-evaluation, surveillance (ISO/IEC 17000 8.1). Ruling R-36.
non-deterministic system
A system that, given the same inputs and starting state, will not always produce the same outputs. The system under test is one, which is why a single probe is weak evidence and why pass rates are estimated over replicate sessions. Source: IEEE Computer Society, non-deterministic system, p. 272 (ISO/IEC TR 29119-11:2020, testing of AI-based systems): “system which, given a particular set of inputs and starting state, will not always produce the same set of outputs and final state”. Also: nondeterministic system. Ruling R-13.
objective evidence
Data that supports the existence or truth of something. Here, what is collected as the basis for a determination: a response at one turn, a session’s trajectory, or several of them rolled up over a test suite, as in a robustness battery. Each evidence item bears on one acceptance criterion’s expected result and is bound to the test plan it was collected under. Evidence is the domain; the determination made on it, passed, failed or cannot tell, is the codomain. On its own evidence verifies nothing; a determination rules on it and an attestation judges it. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.8.6: “data supporting the existence or verity of something”. Also: evidence. Ruling R-18, R-20.
OG-CAIE
Contextual AI Evaluation performed with the method this specification sets out: the Evaluation Process Ontology, the same six steps for every domain inside a contracting lifecycle, applied with a Domain-Specific Ontology supplied or approved by the domain’s experts, leaving a record anyone can check for conformance, traceability and coverage. Coined by the authors for this specification; it cites no source. Also: Ontology-Grounded Contextual AI Evaluation. Ruling R-09, R-29.
operational envelope
The requirements this system under test must meet in this operational environment, with their acceptance criteria and weights, declared before any probe is run. It refines the idea of a requirement by adding specificity: the same system in a different environment, or a different system in the same environment, gets a different envelope. That is what makes the evaluation contextual, and what separates it from a benchmark; the nearest practice is that of safety-critical systems. Coverage is measured over the envelope and only over it. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.5.1 requirement: “need or expectation that is stated, generally implied or obligatory”. Also: requirement set. Ruling R-03, R-14.
operational environment
The conditions and assumptions about the deployment setting under which acceptable behavior is defined: who uses the system, for what, and what is out of scope. Source: IEEE Computer Society, operational environment, p. 282 (IEEE 982-2024, 3.1): “physical context, setting, and circumstances used to support the in-service operation of a system”. Also: operating environment. Ruling R-02.
outcome
The result recorded for a judgment: passed, failed or cannot tell, the three EARL values the record uses (EARL also allows inapplicable and untested). Not binary on purpose. Source: IEEE Computer Society, test result, p. 438 (ISO/IEC/IEEE 29119-2:2021, 3.56): “indication of whether a specific test case has passed or failed, i.e. if the actual results correspond to the expected results or if deviations were observed”.
performance
A measurable result. Here: how the system did on what was exercised, reported as the pass, fail and cannot-tell fractions of the covered criteria. Never merged with coverage. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.7.3: “measurable result”.
probe
A scenario tied to the acceptance criteria it exercises, derived from the Domain-Specific Ontology and checked for consistency before use, applied at one turn of a session. Its run yields one prompt and response pair; the response collected is evidence bearing on those criteria. A test plan may list its probes or choose them by a test strategy as the trajectory unfolds. Source: IEEE Computer Society, test case, p. 432 (fragment of the entry): “set of test inputs, execution conditions, and expected results developed for a particular objective”. Also: prompt, task, test case (SEVOCAB, tied to criteria).
provider
An organization that provides a product or a service. Two products change hands around an OG-CAIE evaluation, so there are two providers: the evaluation service provider, the testing organization, provides the evaluation as a service; the test item provider provides, and is accountable for, the AI system under test. They are different organizations when the evaluation is independent, and the record states that fact rather than assuming it. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.1.9: “organization that provides a product or a service”. Also: contractor, supplier. Ruling R-21.
recommendation
The advice the evaluation gives the sponsor about deploying the test item: its fitness (fit to deploy, fit with conditions, not fit) and the conditions or remediation, written by the evaluation operator, owned by the domain expert’s approval of the final report, and derived from every attestation the coverage used. Source: IEEE Computer Society, recommendation, p. 340 (ISO/IEC 14143-2:2011, 3.9): “provision that conveys advice or guidance”. Also: deployment fitness finding. Ruling R-50, R-51.
record
A document stating results achieved or giving evidence of activities performed: any document that records what was done or what was found, in the most general sense the standard gives the word. The evaluation record is its special case here, the record of one evaluation; the DSO release, the test plan and the report are records too, each in its own right. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.8.12: “document stating results achieved or providing evidence of activities performed”. Ruling R-48.
red teaming
A testing level in which human testers probe the system adversarially to see whether it stays within its guardrails. One way to run the probes in step four. Source: NIST AI 700-2, Appendix A, Red teaming, p. 17: “Testing level which evaluates whether applications adhere to guardrails in response to adversarial prompting or stress testing by human testers.”.
repeatability
How closely results agree when the same measurement is repeated under the same conditions: same system version, same procedure, same operator or policy, same session protocol, over a short period. Replicate sessions under repeatability conditions are what turn a probe into a pass rate. Source: IEEE Computer Society, repeatability (of results of measurements), p. 348 (ISO/IEC TR 14143-3:2003, 3.8; ISO/IEC 25021:2012, 4.15): “closeness of the agreement between the results of successive measurements of the same measurand carried out under the same conditions of measurement”. Ruling R-13, R-51.
report
The record’s delivered export: the summary of the testing performed, assembled by machine from the record with coverage and performance recomputed, a draft while gaps remain and final when it rests on a passed conformance verdict on the record and every criterion is covered; approved by a domain expert and delivered with the recommendation. Source: IEEE Computer Society, test completion report, p. 433 (ISO/IEC/IEEE 29119-2:2021, 3.26): “report that provides a summary of the testing that was performed”. Also: final report, test completion report. Ruling R-50.
reproducibility
How closely results agree when the measurement is repeated under changed conditions: different operators, different sessions, a different evaluation team. Two evaluations that share a requirement set and a DSO release are comparable to the extent their results reproduce. Source: IEEE Computer Society, reproducibility (of results of measurements), p. 349 (ISO/IEC TR 14143-3:2003, 3.9; ISO/IEC 25021:2012, 4.16): “closeness of the agreement between the results of measurements of the same measurand carried out under changed conditions of measurement”. Ruling R-13, R-51.
requirement
Something the system must do, or a condition it must meet, to be fit for its purpose. Broad requirements are decomposed into acceptance criteria that can be checked one at a time. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.5.1: “need or expectation that is stated, generally implied or obligatory”.
requirements traceability
The documented path linking what was required to how it was tested and what was found, so a coverage claim can be checked by anyone holding the record. Source: IEEE Computer Society, requirements traceability, p. 352 (ISO/IEC/IEEE 29148:2018): “identification and documentation of the derivation path (upward) and allocation/ flow-down path (downward) of requirements in the requirements set”. Ruling R-15.
risk-based testing
Testing in which what is tested, and how much, is chosen according to analyzed risk. Deployment sensitivity is the weight that makes coverage reflect consequence. Source: IEEE Computer Society, risk-based testing, p. 362 (ISO/IEC/IEEE 29119-2:2021, 3.16): “testing in which the management, selection, prioritization, and use of testing activities and resources are consciously based on corresponding types and levels of analyzed risk”.
scenario
A step-by-step description of a situation the system is put through: the user, their circumstances, and what they ask. A scenario becomes a probe once it is tied to the criteria it exercises. Source: IEEE Computer Society, scenario, p. 369 (ISO/IEC/IEEE 24765e:2015): “step-by-step description of a series of events that occur concurrently or sequentially”.
second-party conformity assessment activity
Evaluation performed by an organization with a user interest in the test item: a purchaser, a deploying agency, those who represent its users (ISO/IEC 17000 4.4; a regulator is an authority that uses results, not a user). The sponsor’s own activity in commissioning and accepting an evaluation is second-party; the independent testing organization it commissions performs a third-party activity (sheet 10-08, R-50). Source: ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 4.4 second-party conformity assessment activity: “conformity assessment activity that is performed by a person or organization that has a user interest in the object of conformity assessment”. Also: second party. Ruling R-21.
session
One pairing of one tester with one system under test, in which a sequence of turns is run. Sessions are stateful: what the system says at a later turn depends on everything said before, so evidence belongs to its session, not only to its probe. Source: NIST AI 700-2, Appendix A, Session, p. 17 (first sentence of the entry): “A single unit of ARIA testing, consisting of a pairing of one tester and one application.”. Ruling R-13.
stakeholder
A person or organization that can affect, be affected by, or believe itself affected by a decision or activity. The parties (the sponsor, the testing organization, the test item provider), the people who act for them, and the affected populations are the stakeholders named in this specification; customers and providers are kinds of interested party in the standard’s own examples. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.1.4: “person or organization that can affect, be affected by, or perceive itself to be affected by a decision or activity”. Also: interested party.
statement of work
The sponsor’s decisions about the work to be performed under the contract; here, for each affected population, whether it is interviewed, with its stakeholder needs documented as input, or represented by a member of the team. A scope-of-work judgment: representation is the common case, interviews are reserved for underrepresented stakeholders or underdocumented needs because of the effort they cost. Decided with the agreement, pinned at the contract, consumed by the scope step, and traceable either way. Source: IEEE Computer Society, statement of work, p. 406 (ISO/IEC 33202:2024, 3.25): “statement of the expected outcomes and outline of the work required to achieve the outcomes”. Also: SOW, scope of work. Ruling R-40.
sufficiency
The expert judgment that the evidence gathered is enough to support the claim being made. A quantity judgment, distinct from appropriateness. A judgment with no evidence cannot record a pass or a fail. Source: Hawkins, Section 3.3 Asserted solution, p. 10: “it is being asserted that the evidence put forward is sufficient to support the claim”. Ruling R-08.
technical expert
A person who provides specific knowledge or expertise to the evaluation team. Two kinds sit on the team. The domain expert is expert in the Domain-Specific Ontology: they supply or approve its release, attest that it is appropriate for the case being evaluated (this system under test, these requirements, this operational environment), and may be consulted on interpreting evidence. The evaluation operator is expert in performing the tests the Evaluation Process Ontology specifies, conditioned on the DSO, and in assembling the evidence. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.12.9: “person who provides specific knowledge or expertise to the audit team”. Also: domain expert, subject matter expert. Ruling R-10, R-50.
test coverage
How much of the declared acceptance criteria the evaluation actually reached, as a share weighted by deployment sensitivity. It says how much was evaluated, not how well the system did. Source: IEEE Computer Society, test coverage, p. 433 (ISO/IEC/IEEE 29119-2:2021, 3.28): “degree, expressed as a percentage, to which specified test coverage items have been exercised by a test case or test cases”. Also: coverage. Ruling R-01.
test item
The deployed AI system being evaluated, taken as a black box: only what it is given and what it produces are recorded. The prose calls it the system under test; the standard headword is test item. Source: IEEE Computer Society, test item, p. 435 (ISO/IEC/IEEE 29119-2:2021): “work product to be tested”. Also: SUT, application, system of interest, system under test, test object. Ruling R-19.
test plan
The statement, made before any test runs, of what the evaluation will test and how: which acceptance criteria are its objectives, which probes are its means, and what each probe is expected to show. Derived from the requirement set and the Domain-Specific Ontology in step three; every attestation is bound to the plan that produced its evidence. Source: IEEE Computer Society, test plan, p. 437 (ISO/IEC/IEEE 29119-2:2021, 3.50): “detailed description of test objectives to be achieved and the means and schedule for achieving them, organized to coordinate testing activities for some test item or set of test items”. Ruling R-12.
test strategy
The part of a test plan that decides which probe comes next from the trajectory so far, grounded in the Domain-Specific Ontology, rather than listing probes in advance. It closes a loop: the system’s outputs feed back into the choice of the next input. A human red teamer following their own judgment is one such strategy; a written policy is another. Source: IEEE Computer Society, test strategy, p. 439 (ISO/IEC/IEEE 29119-2:2021, 3.59): “part of the test plan that describes the approach to testing for a specific project, test level, or test type”. Also: semantic state feedback, state feedback policy, strategy. Ruling R-13.
test suite
A set of sessions run together, for example replicate sessions of one strategy, or the varied sessions of a sensitivity test or robustness battery. Evidence rolled up over a suite is evidence at the third level, above probe and session. Source: IEEE Computer Society, test suite, p. 439 (ISO/IEC/IEEE 29119-1:2022, 3.129): “set of test cases or test procedures”. Also: battery, robustness battery. Ruling R-18.
third-party conformity assessment activity
Evaluation performed by an organization independent of the provider of the test item and with no user interest in it, whoever commissions it (sheet 10-08, R-50). Independence of the testing organization from the test item provider is what makes an OG-CAIE result independent; the record states it as a fact about the parties rather than assuming it, and states separately whether the sponsor has a user interest. Source: ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 4.5 third-party conformity assessment activity: “conformity assessment activity that is performed by a person or organization that is independent of the provider of the object of conformity assessment and has no user interest in the object”. Also: independent testing, third party. Ruling R-21.
trajectory
The recorded sequence of prompt and response pairs of a session, in order. The system’s internal state is never observed; the trajectory is the observable realization, and plays the role a time series of measurements plays in system identification. A small change early in a trajectory can change everything after it. Source: IEC 60050-351:2013 International Electrotechnical Vocabulary, 351-41-10 trajectory: “representation of the solution x(t) of the state equation as connecting line of the ends of the vector x(t) in state space with time as parameter”. Also: history, realization. Ruling R-13.
verification
Confirming with objective evidence that specified requirements were fulfilled. Applied to the evaluation itself: did it exercise what it declared, and was its declared process followed. Both are mechanical checks over the record. Source: ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.11.12: “confirmation, through the provision of objective evidence, that specified requirements have been fulfilled”.

SEVOCAB definitions: Copyright © 2021 IEEE. Used by permission.

4Two cycles, and the canon they match

Work on an evaluation happens in two cycles, and neither is invented. The outer cycle is the contracting lifecycle, from the sponsor’s need to its acceptance of the delivery: its actors are the parties named above, and its six steps each cite their canon below, from the SEBoK’s (the Guide to the Systems Engineering Body of Knowledge) account of ISO/IEC/IEEE 15288 to ISO 9000’s contract and the ISO/IEC/IEEE 29119-2 test environment and completion report; the standards’ process groups stay coordinate; the cycle orders them under one agreement. The inner cycle is the Evaluation Process Ontology, performed under that agreement between access and delivery: its six steps are the 15288 stakeholder-needs and system-requirements processes, the 29119-2 test strategy and planning, test execution and test completion processes, and ISO/IEC 17000’s own function, review, decision and attestation (ruling R-31).

The distinction matters for what the record must say. What the contract pins is the first layer of assumptions: the counterparties, the test item, the frame of the requirements and the method. The evaluation takes those as given, documents them, and cannot change them. What the evaluation pins is the second layer: the operational environment, the requirement set, the DSO release, the plan. Every item kind in the record declares which layer fixes it, and a shape, one machine-checked rule over the record, checks that the first layer is closed before the second opens (ruling R-32). In the model the two cycles are one nested model (ruling R-33): the contracting lifecycle’s fulfil step is a black box typed by the evaluation process, its inputs the agreement, the statement of work and the access the contract pinned, its outputs the report, its approval and the recommendation the delivery carries; the evaluation process is that box opened, and it conforms to the interface above it by typing, with a shape checking that every input is fed and every output used. What this specification adds to the standards is only the executable form: each step’s inputs and outputs typed, each seam (a wire carrying one item kind from one step’s output to another’s input) wired, each record checked.

4.1The contracting lifecycle, C1 to C6

The outer cycle, contracting through delivery, whose actors are the parties and whose steps each cite their own canon, ordered under one agreement (Guide to the Systems Engineering Body of Knowledge, Stakeholder Needs (Acquisition and Supply), p. 353: “These needs and requirements are expressed in agreements between acquirers and suppliers.”; sheet 10-22).

StepMatchesAlso
C1 need: the sponsor states its mission and obligations towards the affected populations, the problem, and what success would look likeGuide to the Systems Engineering Body of Knowledge, Business or Mission Analysis, p. 543: “The purpose of Business or Mission Analysis is to understand a mission or market problem, threat, or opportunity, and to establish the goals, objectives and measures of success of a potential solution class.”ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.5.1 requirement: “need or expectation that is stated, generally implied or obligatory”
C2 propose: the sponsor seeks the service and the testing organization respondsGuide to the Systems Engineering Body of Knowledge, Enterprise Systems Engineering, contract products and services, p. 935: “Contract products and services often demand tailor-made system/service solutions which are typically specified by a single customer to whom the solution is provided. The supplier responds with proposed solutions.”ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 4.10 access: “opportunity for an applicant to obtain a conformity assessment service from a body under a conformity assessment scheme”; IEEE Computer Society, proposal, p. 329 (ISO/IEC/IEEE 24765:2014, SEVOCAB tag 24765c): “supplier’s offer to provide a system or service, usually including benefits, costs, risks, opportunities, and other factors applicable to decisions”; IEEE Computer Society, request for proposal (RFP), p. 350 (ISO/IEC/IEEE 24765:2017): “document used by the acquirer as the means to announce its intention to potential bidders to acquire a specified system, software product, or software service”
C3 agree: the service agreement fixes the counterparties, the test item, the requirements’ frame and the methodGuide to the Systems Engineering Body of Knowledge, Stakeholder Responsibilities, Acquirer/Supplier Agreements, p. 353 (INCOSE 2012, the acquisition process): “The acquisition process includes activities to identify, select, and reach commercial agreements with a product or service supplier.”ISO 9000:2026(en) Quality management — Fundamentals and vocabulary, 3.3.13: “binding agreement”; ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 4.9 conformity assessment scheme: “set of rules and procedures that describes the objects of conformity assessment, identifies the specified requirements and provides the methodology for performing conformity assessment”
C4 access: the test item provider (ISO/IEC 17000 first party) provides the test item and the test environmentIEEE Computer Society, test environment and data management process, p. 434 (ISO/IEC/IEEE 29119-2:2021, 3.37): “test process for establishing and maintaining a required test environment and corresponding test data”ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 4.2 object of conformity assessment: “entity to which specified requirements apply”; ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, Annex A, A.2 Selection (the functional approach: planning and preparation to obtain the inputs for determination); IEEE Computer Society, test environment, p. 434 (ISO/IEC/IEEE 29119-2:2021, 3.34): “environment containing facilities, hardware, software, firmware, and procedures needed to conduct a test”; IEEE Computer Society, test item transmittal report, p. 435 (ISO/IEC/IEEE 24765:2017): “document identifying test items”
C5 deliver: the authorized representative delivers the final report, its approval and the recommendation to the sponsorIEEE Computer Society, test completion report, p. 433 (ISO/IEC/IEEE 29119-2:2021, 3.26): “report that provides a summary of the testing that was performed”ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 7.3 attestation, Note 1 (fragment of the note): “is intended to convey the assurance that the specified requirements have been fulfilled”; IEEE Computer Society, delivery, p. 120 (ISO/IEC/IEEE 24765:2017): “release of a system or component to its customer or intended user”
C6 accept: the sponsor accepts the delivery as performance of the agreementIEEE Computer Society, acceptance, p. 4 (ISO/IEC/IEEE 24748-5:2017, 3.1): “action by an authorized representative of the acquirer by which the acquirer assumes ownership of products as partial or complete performance of an agreement”Guide to the Systems Engineering Body of Knowledge, Acceptance Criteria (glossary), p. 1437 (INCOSE 2011, Section 6.1.15): “The procurement specification, in the context of the overall agreement, should clearly state the criteria by which the acquirer will accept delivery from the supplier.”

4.2The evaluation, steps 1 to 6

The inner cycle, performed between access and delivery.

StepMatchesAlso
1 scope: every affected population represented, with its interview when the statement of work decided one; build or select the DSO release; declare the operational environmentGuide to the Systems Engineering Body of Knowledge, Stakeholder Needs and Requirements, p. 549 (Stakeholder Needs Definition, Concept Definition): “Stakeholder Needs Definition, the second process in Concept Definition, explores what capabilities are needed by various stakeholders for the system-of-interest (SoI) to accomplish the mission.”ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, Annex A, A.2 Selection (the functional approach)
2 declare the requirement set (operational envelope): requirements, acceptance criteria, weights; the domain expert assesses its appropriatenessGuide to the Systems Engineering Body of Knowledge, System Requirements Definition, Process for Generating System Requirements, p. 558: “The System Requirement Definition activities begin with the transformation of the integrated set of needs into a set of requirements for the SoI. These requirements must be appropriate to the level that the SoI exists within the system architecture and communicate “what” the SoI must do to meet the needs, avoiding requirements that state implementation of “how” to achieve the design realization of the physical SoI.”ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 5.1 specified requirement: “need or expectation that is stated”; Guide to the Systems Engineering Body of Knowledge, System Requirements Definition, p. 559: “A detailed analysis of a single need statement may result in multiple requirements expressing what the system must do to meet it, including definition of measurable performance criteria (INCOSE NRM 2022).”; IEEE Computer Society, system requirements specification (SyRS), p. 422 (ISO/IEC/IEEE 29148:2018, 4.1.29): “structured collection of the requirements (functions, performance, design constraints, and attributes) of the system and its operational environments and external interfaces”
3 plan: the test plan, its probes or strategy, consistency-checked and approvedIEEE Computer Society, test strategy and planning process, p. 439 (ISO/IEC/IEEE 29119-2:2021, 3.51): “test management process used to design the test strategy, complete test planning, and create and maintain test plans”IEEE Computer Society, test design and implementation process, p. 433 (ISO/IEC/IEEE 29119-2:2021, 3.32): “test process for deriving and specifying test cases and test procedures”
4 execute: sessions of turns against the test item; responses; evidence collected under the planIEEE Computer Society, test execution process, p. 435 (ISO/IEC/IEEE 29119-2:2021, 3.40): “dynamic test process for executing test procedures created in the test design and implementation process in the prepared test environment and recording the results”ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 6.2 testing: “determination of one or more characteristics of an object of conformity assessment, according to a procedure”; ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, Annex A, A.3 Determination (the functional approach)
5 determine and attest: determinations on evidence; attestations by the domain expertISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 7.2 decision: “conclusion, based on the results of review, that fulfilment of specified requirements has or has not been demonstrated”ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 7.1 review: “consideration of the suitability, adequacy and effectiveness of selection and determination activities, and the results of these activities, with regard to fulfilment of specified requirements by an object of conformity assessment”; ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, 7.3 attestation: “issue of a statement, based on a decision, that fulfilment of specified requirements has been demonstrated”; ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles, Annex A, A.4 Review, decision and attestation (the functional approach); IEEE Computer Society, evaluation, p. 155 (ISO/IEC 25001:2014, 4.1): “systematic determination of the extent to which an entity meets its specified criteria”
6 report: the record’s conformance checked by machine and the verdict recorded; the report assembled from the record and the verdict with coverage and performance recomputed, a draft with its gaps flagged or a final report resting on a passed verdict; its contents approved by a domain expert; recommendation written; handed to the authorized representative for deliveryIEEE Computer Society, test completion process, p. 433 (ISO/IEC/IEEE 29119-2:2021, 3.25): “test management process that aims to ensure that useful test assets are made available for later use, test environments are left in a satisfactory condition, and the results of testing are recorded and communicated to relevant stakeholders”

SEVOCAB definitions: Copyright © 2021 IEEE. Used by permission.

5Sources

The register is sources/sources.ttl. Committed snapshots live under sources/archive/; held-locally sources are referenced by content hash only. This work’s CC BY-SA licence does not extend to the sources.

RankSourcePostureSnapshotsLicence note
1ISO 9000:2026(en) Quality management — Fundamentals and vocabulary (Online Browsing Platform preview)heldLocally38ISO copyright; the preview is publicly browsable; short attributed quotations only
2IEEE Computer Society, Software and Systems Engineering Vocabulary (SEVOCAB), PDF export created 2026-09-02 (481 pp.)heldLocally1Each definition may be copied provided the IEEE statement remains with it, and the ISO/IEC definitions provided their source is cited; the PDF itself is not redistributed
3NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), January 2023committed1US Government work, public domain
3NIST AI 700-2, Assessing Risks and Impacts of AI (ARIA): ARIA 0.1 Pilot Evaluation Report (November 2025)committed1US Government work, public domain
4W3C, Evaluation and Report Language (EARL) 1.0 Schema, Working Group Note 2 February 2017committed2W3C Document Licence
4W3C, PROV-O: The PROV Ontology, Recommendation 30 April 2013committed2W3C Document Licence
4W3C, Shapes Constraint Language (SHACL), Recommendation 20 July 2017committed1W3C Document Licence
4W3C, SKOS Simple Knowledge Organization System Reference, Recommendation 18 August 2009committed1W3C Document Licence
5IEC 60050-351:2013 International Electrotechnical Vocabulary, Part 351: Control technology (Electropedia)citeOnly0IEC copyright; Electropedia is publicly browsable; short attributed quotations only; quotes transcribed from the browser and verified by Z on 2026-09-06 (rulings sheet 02)
5ISO/IEC 17000:2020(en) Conformity assessment — Vocabulary and general principles (Online Browsing Platform)citeOnly0ISO copyright; the vocabulary is publicly browsable; short attributed quotations only (the party terms, object, access, scheme, specified requirement, testing, review, decision, attestation and its note, surveillance)
5JCGM 200:2012 International vocabulary of metrology, basic and general concepts and associated terms (VIM), 3rd edition (BIPM)heldLocally1JCGM copyright; freely downloadable from BIPM; held locally
6INCOSE Guide to Writing Requirements v4, Summary Sheet (INCOSE-TP-2010-006-04, June 2023)heldLocally1INCOSE copyright restrictions; held locally
6IAASB, International Standard on Auditing 500: Audit Evidence (effective 15 December 2009)heldLocally1IFAC copyright; freely downloadable; held locally
6NIST Technical Note 1297, Guidelines for Evaluating and Expressing the Uncertainty of NIST Measurement Results, 1994 editioncommitted1US Government work, public domain
6Guide to the Systems Engineering Body of Knowledge (SEBoK) v2.14, released 18 May 2026heldLocally1CC BY-NC-SA; held locally per ruling R-04
7Gruber, A Translation Approach to Portable Ontology Specifications, Knowledge Acquisition 5(2), 1993 (author copy)heldLocally1Academic Press copyright; author-hosted copy; held locally
7Hawkins, Kelly, Knight and Graydon, A New Approach to Creating Clear Safety Arguments, Safety-Critical Systems Symposium 2011 (author copy, University of York)heldLocally1Springer proceedings; author-hosted copy; held locally
7Hogan et al., Knowledge Graphs, ACM Computing Surveys 54(4), 2021 (arXiv:2003.02320v3)heldLocally1arXiv non-exclusive licence; held locally
7Popper, The Logic of Scientific Discovery, London: Hutchinson, 1959citeOnly0cited as the canonical source named as the source of the definitions presented at the SciPy 2026 session; not quoted directly
8Hollek and Zargham, Building Scientific Approaches to Generative AI, Birds-of-a-Feather session, SciPy 2026 (public event)citeOnly0a public event on the conference record, cited as a fact about what was presented there; the definitions in the digest are the authors’ own words as presented; the organization logos shown at the session are reproduced at assets/hi-dsg-collaboration.png with the organizations’ agreement (ruling R-47, sheet 06-02)

6Sources cited

References
  1. International Organization for Standardization. (2026). ISO 9000:2026 Quality management — Fundamentals and vocabulary. International Organization for Standardization. https://www.iso.org/obp/ui/#iso:std:iso:9000:ed-5:v1:en
  2. IEEE Computer Society and ISO/IEC JTC 1/SC 7. (2026). IEEE Computer Society, Software and Systems Engineering Vocabulary (SEVOCAB), PDF export created 2026-09-02 (481 pp.). IEEE Computer Society and ISO/IEC JTC 1/SC 7. https://www.computer.org/sevocab
  3. Artificial Intelligence Risk Management Framework (AI RMF 1.0) (Techreport NIST AI 100-1). (2023). NIST. 10.6028/NIST.AI.100-1
  4. Amironesei, R., Godil, A., Greenberg, C., Greene, K., Hall, P., Jensen, T., Fiscus, J., & Schulman, N. (2025). Assessing Risks and Impacts of AI (ARIA): ARIA 0.1 Pilot Evaluation Report (Techreport NIST AI 700-2). National Institute of Standards and Technology. 10.6028/NIST.AI.700-2
  5. W3C. (2017). Evaluation and Report Language (EARL) 1.0 Schema. W3C. https://www.w3.org/TR/2017/NOTE-EARL10-Schema-20170202/
  6. W3C. (2017). Shapes Constraint Language (SHACL). W3C. https://www.w3.org/TR/2017/REC-shacl-20170720/
  7. IEC. (2013). IEC 60050-351:2013 International Electrotechnical Vocabulary, Part 351: Control technology. IEC. https://www.electropedia.org/iev/iev.nsf/index?openform=&part=351
  8. ISO/IEC. (2020). ISO/IEC 17000:2020 Conformity assessment — Vocabulary and general principles. ISO/IEC. https://www.iso.org/obp/ui/en/#iso:std:73029:en
  9. JCGM. (2012). JCGM 200:2012 International vocabulary of metrology, basic and general concepts and associated terms (VIM), 3rd edition. BIPM. 10.59161/jcgm200-2012
  10. INCOSE. (2023). Guide to Writing Requirements v4, Summary Sheet. INCOSE.
  11. International Auditing and Assurance Standards Board. (2009). International Standard on Auditing 500: Audit Evidence. International Auditing and Assurance Standards Board.
  12. Taylor, B. N., & Kuyatt, C. E. (1994). Guidelines for Evaluating and Expressing the Uncertainty of NIST Measurement Results (Techreport NIST Technical Note 1297). National Institute of Standards and Technology. https://nvlpubs.nist.gov/nistpubs/legacy/tn/nbstechnicalnote1297.pdf
  13. SEBoK Editorial Board. (2026). Guide to the Systems Engineering Body of Knowledge (SEBoK) v2.14, released 18 May 2026 (N. Hutchison, Ed.). The Trustees of the Stevens Institute of Technology. https://sebokwiki.org/
  14. Hawkins, R., Kelly, T., Knight, J., & Graydon, P. (2011). A New Approach to Creating Clear Safety Arguments. In C. Dale & T. Anderson (Eds.), Advances in Systems Safety: Proceedings of the Nineteenth Safety-Critical Systems Symposium (pp. 3–23). Springer. 10.1007/978-0-85729-133-2_1