Experiment evidence v1
Before a run, Apizr can describe what you intend to execute. After a run, it can record what was actually observed. Those are separate artifacts.
An Experiment Plan describes intended experiment inputs and controls. An Experiment Run records observed execution evidence. Neither proves scientific causality or reproducibility.
| Before execution: Plan | After execution: Run |
|---|---|
| Selected notebook or Python source | Source actually used and exact Plan digest |
| Intended dataset/input identities | Observed dataset/input identities |
| Intended parameters and randomness controls | Effective parameters and observed randomness controls |
| Declared or static environment evidence | Observed environment evidence |
| Execution kind, policy digest and declared controls | Status, optional timing, metrics and output artifacts |
These are development contracts in apizr.experiments. This version supplies no
experiment command, inspector, runner, capture mechanism, comparison or history
store. Producers must supply evidence explicitly. Reading a Plan or Run does not
execute its source or resolve its references.
An explicit Plan/Run pair
This fictional example demonstrates construction, not a captured execution. The example digests identify placeholder content; a real producer must supply the actual evidence. Runtime labels below describe the observations that producer would need to collect.
from apizr.experiments import (
EvidenceOrigin,
ExecutionIntent,
ExperimentPlan,
ExperimentRun,
InputArtifact,
Metric,
OutputArtifact,
Parameter,
RandomnessControl,
SourceIdentity,
plan_digest,
run_bytes,
run_digest,
validate_run_binding,
)
subject = SourceIdentity(
kind="notebook",
reference="notebooks/fraud_detection.ipynb",
digest="a" * 64,
)
plan = ExperimentPlan(
subject=subject,
inputs=(
InputArtifact(
name="train",
reference="data/train.csv",
digest="b" * 64,
size=4096,
origin=EvidenceOrigin.STATIC,
),
),
parameters=(
Parameter(
name="max_depth",
value=8,
origin=EvidenceOrigin.DECLARED,
),
),
randomness=(
RandomnessControl(
provider="sklearn",
name="random_state",
value=42,
origin=EvidenceOrigin.DECLARED,
),
),
execution=ExecutionIntent(kind="trusted-notebook", policy_digest="c" * 64),
)
# A future execution producer would supply the actual observed values.
run = ExperimentRun(
plan_digest=plan_digest(plan),
subject=subject,
status="success",
effective_parameters=(
Parameter(
name="max_depth",
value=8,
origin=EvidenceOrigin.RUNTIME,
),
),
metrics=(Metric(name="roc_auc", value=0.91, origin=EvidenceOrigin.RUNTIME),),
outputs=(
OutputArtifact(
name="fraud-model",
digest="d" * 64,
size=1024,
media_type="application/octet-stream",
origin=EvidenceOrigin.RUNTIME,
),
),
)
validate_run_binding(run, plan)
canonical_run = run_bytes(run)
identity = run_digest(run)
The reviewed Plan fixture and Run fixture include partial unknown input evidence and explicit environment/timing records. They are fictional test values, not reported scientific results.
Evidence origins and unknowns
EvidenceOrigin distinguishes declared, static, runtime and unknown.
A Plan permits declared, static and unknown input, parameter, randomness and
environment evidence. Execution controls are explicitly declared. A Run permits
runtime and unknown observations; finite JSON metrics, identified output artifacts,
timing and diagnostic records require runtime origin. No generic inferred-truth
state exists.
An unknown input may retain its selected logical name/reference, but its digest and size are null. An unknown parameter or randomness value is null; a declared or runtime JSON null is distinguishable by origin. Environment values and package versions are null exactly when their origin is unknown. Empty collections mean no records were supplied, not that capture established an empty environment. Absent optional categories do not claim observation.
A captured seed/control is evidence of a declared or observed randomness control. It is not proof that execution is deterministic.
Independent schemas and public values
The independent versions are apizr.experiment-plan/v1 and
apizr.experiment-run/v1, exposed as schema_version. The schemas are generated
directly from the typed models:
All nested models are frozen and reject unknown fields. Python construction uses
strict types, tuples and EvidenceOrigin members; JSON input uses arrays and
origin strings via model_validate_json. Parameters and metrics copy caller containers into
immutable mappings and tuples. Producers create new values to add observations;
they do not mutate a Plan into a Run.
| Public value | Evidence carried |
|---|---|
SourceIdentity |
kind (python or notebook), relative reference, SHA-256 digest, optional executable_digest, logical module, exact opaque capability_id |
InputArtifact |
Logical name, optional relative reference/digest/size, explicit origin |
Parameter |
Logical name, finite bounded JSON value and origin |
RandomnessControl |
Parameter fields plus explicit provider/source |
EnvironmentValue |
Python implementation/version or platform value and origin |
PackageEvidence |
Canonical package name, version when known and origin |
EnvironmentEvidence |
Optional Python/platform values, explicit package inventory and lock/config artifacts |
ExecutionIntent |
Generic explicit kind, optional policy digest, declared controls |
Metric |
Logical name, bounded finite JSON, optional unit for numeric scalars, runtime origin |
OutputArtifact |
Logical name, digest, optional relative reference/known size/media type, runtime origin |
RunDiagnostic |
Bounded stable lowercase code and runtime origin; no traceback |
RunTiming |
Optional observed start/end/duration and runtime origin |
Capability identity is optional, including for notebooks without an inference capability. If supplied, the exact logical identity participates in the Plan hash and binding. Selecting a capability here does not establish Readiness or authorize Exposure. Executable/transformed digests are optional explicit evidence; the contract does not transform notebooks or compute source fingerprints.
Run status is success, failed or cancelled. A successful Run cannot contain
failure diagnostics. Failed/cancelled Runs may retain partial observations without
inventing metrics or exception details. Timing requires at least one observed
field, rejects naive datetimes, normalizes aware datetimes to UTC (Z in JSON),
and requires end at or after start and finite non-negative duration. Duration may
come from a monotonic clock and need not equal a wall-clock subtraction.
Identity, ordering and limits
plan_bytes and run_bytes revalidate even unchecked model_construct or
model_copy values and nested records before dumping typed JSON. They sort JSON
object keys, emit UTF-8 with ensure_ascii=False, compact separators,
allow_nan=False and exactly one final newline. plan_digest and run_digest
return lowercase SHA-256 hex over those exact bytes. Canonical documents are
bounded to 4 MiB. No timestamp, UUID or host directory is automatically added.
Canonical serialization means the same artifact content produces the same bytes and digest. It does not mean separate experiment executions produce the same Run digest.
Named collections sort by name; randomness sorts by (provider, name);
diagnostics sort by code. Package names normalize lowercase and runs of -, _
and . to -. Duplicate identities are rejected after normalization. JSON
object order is irrelevant, while arrays inside parameter and metric values retain their
meaningful order. Integer and floating-point values retain their JSON number
representation; 1 and 1.0 need not have the same artifact identity.
| Boundary | Maximum |
|---|---|
| Inputs, parameters, metrics, outputs per artifact | 256 each |
| Randomness records, execution controls | 128 each |
| Packages, environment lock/config artifacts, diagnostics | 4096 / 64 / 32 |
| Logical names, versions, media types, diagnostic codes | 128 characters |
| Python/platform evidence values, metric units | 256 / 64 characters |
| Project-relative references, capability identity | 1024 / 512 characters |
| One parameter/control/metric JSON value | 64 KiB, depth 16, 4096 value nodes |
| One JSON array/object, string, object key | 1024 entries / 8192 / 256 characters |
| Integer values, sizes | Signed 64-bit; sizes non-negative |
Text must be valid UTF-8. Logical text fields reject ASCII controls. Local references must be portable project-relative paths: no absolute/drive paths, backslashes, empty/dot/parent segments, NUL, URI schemes or home expansion. References do not name registry objects or remote stores. Moving a project leaves identity unchanged when its logical evidence stays the same. Arbitrary parameter strings are data, not interpreted paths or Python expressions.
JSON Schema Draft 2020-12 describes structural validation. Cross-field origin and status consistency, named uniqueness/order, UTF-8, portable-reference semantics, recursive depth/node/byte budgets and Run-to-Plan binding are additionally enforced by the typed API. Consumers must use that validation; schema validation alone is not artifact admission. Regenerate schemas with:
uv run --locked python scripts/export_experiment_schemas.py
Binding and trust boundaries
validate_run_binding(run, plan) validates caller-provided artifacts, the exact
Plan digest and the whole source identity, including absent/present capability,
logical reference/module and executable digest. It performs no external lookup.
Effective inputs, parameters and environment can differ from intent and remain
explicitly recorded without changing the Plan. A valid binding validates the
claim's internal consistency; it does not authenticate a producer or prove that
execution occurred. Use run_bytes to enforce the complete canonical byte budget
before handing the artifact to another system.
A Run can show that two executions used different inputs, parameters or environments. It does not prove that one of those changes caused a metric difference.
Matching metric names alone do not establish comparable experiments. There are no
causality, determinism or reproducibility flags. Producers must never put secrets
or machine-local absolute paths in parameter values. These contracts have no
credentials, authentication structure, secret store or arbitrary environment
variable map. Names such as token_count remain legitimate; the model does not
guess secrets from parameter names.
Experiment Run is an attestable artifact; it is not an attestation format.
The experiment domain depends on existing primitive contracts, bounded descriptor
access in workspace.files, and the
existing Pydantic runtime. No Attest, OCI, MCP, FastAPI, MLflow, W&B, DVC or
DS framework defines these values. External systems may consume their canonical
bytes later. The package composition rule
records this dependency boundary.
Additive input evidence in 0.4.5 development
261 adds three optional fields to InputArtifact: content_origin, format_hint
and uri. Absent values are omitted from canonical serialization. Old valid v1
JSON still validates, and the original #260 Plan and Run golden bytes and SHA-256
values are unchanged. Both schema versions remain v1; their regenerated schemas
add optional properties without adding required fields.
origin describes the input/reference record. content_origin, when supplied,
describes how digest/size were established. A declared local selection can have
origin=declared and content_origin=static after fingerprinting. An absent
content origin retains legacy semantics: it makes no additional provenance claim.
unknown content requires absent digest and size; a known content origin requires
at least one of them. The producer always observes both digest and size together.
Plan evidence rejects runtime content origins, including environment artifacts;
Run observations permit runtime/unknown content origins. This prevents hiding a
runtime observation inside a declared Plan record or static evidence inside a Run.
reference retains its portable project-relative validation. uri is a separate,
mutually exclusive remote reference; see the exact URI and format vocabulary in
experiment inputs. A remote URI is not evidence
of verified bytes. Future runtime producers may observe remote bytes independently;
261 never fetches them. Format hints are canonical evidence, not parser guarantees.
JSON Schema describes structure. No schema alone establishes symlink safety, exact-byte fingerprinting, changed-during-read detection or remote no-fetch. These are responsibilities of the reviewed producer, with executable safety tests.
Randomness and environment producers (#262)
discover_randomness(source, source_reference=...) returns bounded controls,
producer diagnostics and versionless distribution candidates. Controls use the
existing (provider, name) identity; repeated equal controls collapse, while
conflicting or dynamic values remain unknown. Import binding authority and bounded
AST parsing are shared with input discovery. Static presence, including presence
inside a function, does not establish execution. Torch/TensorFlow seed evidence
also retains unknown accelerator_determinism; no deterministic/reproducible
flag, score or causal conclusion is produced.
capture_runtime_environment(distributions=..., root=...) observes the current
trusted process's Python implementation/version, platform, machine architecture
and selected installed distribution versions, always including Apizr. It does not
import selected packages, enumerate all distributions, execute project source,
probe accelerators, invoke subprocesses, install packages or contact a network.
No environment-variable names or values are captured or hashed.
EnvironmentEvidence.architecture is an additive optional EnvironmentValue,
omitted when absent. It participates in the existing Plan/Run origin validators.
The v1 schema identifiers and original canonical golden bytes are unchanged.
discover_environment_specs(root) returns static file evidence; runtime capture
with an explicit root returns runtime observations. Both inspect only root-level
uv.lock, poetry.lock, requirements*.txt, requirements*.lock,
environment.yml, environment.yaml, conda.yml and conda.yaml. Admission is
bounded to 32 matches and 4096 entries. Each uses the same no-follow regular-file
fingerprint reader as input capture, with a default 16 MiB file limit. Successful
content origin matches the enclosing producer, static or runtime; failed content
remains unknown. Contents, absolute paths and arbitrary host state are not emitted.
A lock digest proves byte identity, not that packages were installed from it.
The seeds and environment guide defines exact callable/provider identities, accepted literals, import mappings, bounds and diagnostics. These are producers for the existing contracts, not new experiment execution, history, metrics or comparison orchestration.
Metric signals and observed results (#263)
capture_metric(name, value, unit=...) constructs a canonical runtime Metric
from the caller's explicit observation, without execution or tracking state.
Metric.value now uses the existing bounded, deeply immutable FiniteValue
contract. Numeric, string, boolean, null, array and object values are deliberate;
non-finite numbers and arbitrary Python objects are rejected without stringifying
objects. Units are permitted only on integer/float scalars.
discover_metrics(source, source_reference=...) returns producer-side
MetricSignal occurrences for six sklearn callables. Semantic metric names are
independent of source positions. Each occurrence retains its resolved callable,
source, line, column and static origin, without a fabricated value. Static call
syntax, including inside a function that may never execute, cannot populate
ExperimentRun.metrics. Future execution correlation belongs to #264.
discover_outputs recognizes a bounded joblib literal-filename subset as
OutputSignal values. Each holds an OutputDeclaration plus static provenance;
it has no digest, size or runtime authority. Explicit callers can instead create
a declaration directly or use the pure name=relative-reference parser.
fingerprint_output and fingerprint_outputs revalidate declarations/policy
before reading selected regular files through _files, the same internal
no-follow streamed digest primitive used by inputs and environment specifications.
A successful observation produces the existing generic OutputArtifact, including
runtime origin, exact digest/size and its portable reference. Safety failures
produce diagnostics and no artifact. The primitive has no input/output/environment
origin policy: that remains with each producer.
OutputArtifact.reference is optional and omitted when absent. Metric.value
expands additively to finite JSON; old numeric JSON is unchanged. The Plan schema
is unchanged; the Run schema changes only these two properties. Both v1 schema
identifiers and original Plan/Run golden bytes remain unchanged. Canonical Run
serialization still revalidates unchecked nested models, orders names, rejects
duplicates and preserves exact metric/artifact evidence.
Capture proves current bytes, not which statement wrote them or whether a run succeeded. #264 will own capture at the reviewed post-success point. There is no runner, lifecycle boolean, automatic post-run capture, output deserialization, copy, upload, model registry or scientific inference in these producers. Optional external consumers can use the resulting canonical values later.
The metrics and outputs guide documents exact callables, argument forms, limits, diagnostic codes and explicit capture examples.