Skip to content

Know exactly which local data you used

import pandas as pd

df = pd.read_parquet("data/train.parquet")

Apizr can recognize this literal reference from Python syntax. A separate, explicit fingerprint operation can record data/train.parquet, its exact size and SHA-256, with static content provenance. Neither operation imports pandas or reads rows into a DataFrame.

SHA-256 evidence identifies bytes. Apizr does not upload, cache, version, restore or distribute the dataset. It identifies experiment inputs; it does not own them. Changing training data changes Experiment Plan identity. It does not change the logical capability ID by itself.

Use the Python API

This runnable example uses a temporary local file. It performs discovery without dataset access, then explicitly selects an input, fingerprints it, and constructs a new frozen Plan. No experiment CLI exists yet.

from hashlib import sha256
from pathlib import Path
from tempfile import TemporaryDirectory

from apizr.experiments import (
    ExecutionIntent,
    ExperimentPlan,
    SourceIdentity,
    discover_inputs,
    fingerprint_inputs,
    parse_input_declaration,
    plan_digest,
    select_inputs,
)

source = b'import pandas as pd\npd.read_csv("data/train.csv")\n'
found = discover_inputs(source, source_reference="train.py")
selection = select_inputs(found, (parse_input_declaration("training=data/train.csv"),))

with TemporaryDirectory() as directory:
    root = Path(directory)
    (root / "data").mkdir()
    (root / "data/train.csv").write_bytes(b"feature,label\n0.25,0\n")
    observed = fingerprint_inputs(root, selection.artifacts)
    assert not observed.diagnostics
    plan = ExperimentPlan(
        subject=SourceIdentity(
            kind="python",
            reference="train.py",
            digest=sha256(source).hexdigest(),
            capability_id="python:train:train",
        ),
        execution=ExecutionIntent(kind="training"),
        inputs=observed.artifacts,
    )
    identity = plan_digest(plan)
    assert plan.inputs[0].origin.value == "declared"
    assert plan.inputs[0].content_origin.value == "static"

InputDeclaration(name="training", reference="data/train.parquet") and fingerprint_input(root, declaration) also work independently of discovery. The parser accepts name=reference, splitting at the first = without trimming. --input training=data/train.parquet is future CLI syntax, not an executable command delivered here. Invalid parser input raises the bounded error explicit_input_invalid, without echoing paths or credentials.

All producer result objects are frozen, strict typed values. InputResult contains artifacts and diagnostics; it is not a partial Experiment Plan. Check diagnostics before deciding which evidence to put in a Plan. Fingerprinting returns new artifacts; it never mutates an existing Plan. Failed local observations retain the reference with unknown content and no digest or size. Supplied hash/size claims are discarded before every observation, including failed or remote observations.

What the static recognizer understands

Loader Canonical format hint
pandas.read_csv csv
pandas.read_parquet parquet
numpy.load numpy
joblib.load joblib

Only a literal string in the first positional argument is recognized. Additional loader options are not interpreted. No keyword path forms are supported in v1. Ordinary imports, import pandas as pd, from pandas import read_csv, and from pandas import read_csv as read are recognized without importing the library. Clearly inherited module/function bindings work inside functions, including async functions. Local unambiguous imports work as well.

Lexical analysis intentionally rejects ambiguous bindings: repeated/rebound or deleted aliases, star/conditional/relative imports, parameters and other local shadowing, nested global/nonlocal declarations, and explicit attribute writes. It does not infer library identity from a variable's spelling. A direct call must follow its import syntactically. Whole-scope conservatism can discard an earlier valid read if the alias is later rebound. Class namespaces do not grant authority to methods or nested comprehension scopes. Generic definitions with PEP 695 type parameters are outside this initial recognizer and are ignored in full. Unsupported loaders are ignored; indirect get_loader()(...) calls produce unsupported_input_reference.

These observations describe syntax, not proof a call executes or a branch is reachable. Arbitrary Python side effects cannot be modeled statically. No source import, evaluation, subprocess, environment-variable resolution or network access occurs. Notebook extraction belongs to the future notebook analysis integration; its normalized Python can use this same source-only API.

pd.read_parquet(DATA_PATH)

Apizr does not guess the value of DATA_PATH, even if a constant assignment appears elsewhere. Variables, indexing, f-strings, computed Path expressions, starred arguments and keyword-only calls retain dynamic_input_reference.

A declaration such as training=data/train.parquet declares experiment intent; it does not prove the Python expression evaluates to that path. select_inputs retains the discovery diagnostics while allowing explicitly named inputs into the selection. Explicit names replace static candidates for the same reference and keep declared origin. Different explicit names may select identical bytes; digests do not determine logical input names. Duplicate names are rejected before any batch reads, never resolved with last-write-wins.

Static names equal the exact local reference or reviewed URI. References longer than the 128-character name bound produce input_name_required; use an explicit short name. Repeated reads of the same reference collapse. Incompatible loader hints clear the hint and add input_format_conflict. Candidates and diagnostics are sorted deterministically. Diagnostic locations contain the relative source, 1-based line, 0-based UTF-8 column and recognized loader, without source excerpts.

Remote references stay unverified

pd.read_parquet("s3://bucket/train.parquet")

Apizr can record the reference but does not download remote data merely to manufacture a content fingerprint. uri carries the remote reference; reference remains reserved for local project-relative paths. Their fields are mutually exclusive. Static and declared remote records retain their reference origin, content_origin=unknown, absent digest/size and remote_content_unverified. No GET, HEAD, cloud SDK, credential lookup or local filesystem traversal occurs.

The deliberately narrow v1 URI syntax supports lowercase https://, s3:// and gs://; an ASCII lowercase alphanumeric authority may contain internal dots and hyphens separated by alphanumeric runs. Optional nonempty path segments contain only ASCII letters, digits, ., _, ~ or -; standalone . and .. segments are rejected. Maximum length is 1024 characters. No trailing/doubled slash, userinfo, password, port, query, fragment, percent escape, Unicode authority, control character or file:// scheme is accepted. This is a conservative evidence syntax, not a general URL validator. It cannot detect secrets placed in ordinary path segments; producers must not insert sensitive reference strings.

Local fingerprint safety and limits

One input represents one regular file under an explicit project root. References are portable POSIX-style relative paths, maximum 1024 characters. Absolute paths, .., ., empty segments, backslashes, drive/colon prefixes, home expansion and control characters are rejected without normalization. Absolute source literals produce redacted nonportable_input_reference diagnostics.

The producer reuses workspace.files.directory_fd for no-follow root traversal and opens child components relative to directory descriptors. Input-file symlinks and symlink parents are rejected, even when they point inside the project. Existing macOS /tmp and /var system aliases follow the shared root policy. Ordinary hard links are allowed: identity uses the logical reference and bytes, not inode. Directories, file sets, devices, FIFOs and sockets are unsupported. Nonblocking, no-follow opens and descriptor regular-file checks prevent special-file waits and symlink swaps from granting read authority.

Hashing streams exact file bytes with SHA-256 in chunks of at most 1 MiB. FingerprintPolicy.max_file_bytes defaults to 1 GiB, configurable from 1 byte to a hard maximum 1 TiB per file. The batch maximum is 256 inputs (therefore at most 256 times the selected per-file limit); no directory traversal or implicit file selection occurs. Source discovery separately admits at most 1 MiB of Python, 100,000 AST nodes, depth 128, 256 unique candidates and 512 diagnostics. Exceeding a discovery budget returns only input_discovery_limit, not a silently partial scan. Invalid Python/source identity produces input_source_invalid.

Before/after fstat snapshots compare device, inode, size, nanosecond mtime and ctime. The final directory-relative path identity and exact bytes-read count must also agree. Mutation discards the entire digest with input_changed_during_read; exceeding the read budget produces input_too_large. This detects observable mutation; it is not a filesystem snapshot or protection against privileged adversaries defeating filesystem metadata. Other failures distinguish input_missing, input_symlink, input_not_regular and input_unreadable. Root traversal errors are sanitized through the same codes; an ENOTDIR from a no-follow root/parent open is reported as input_not_regular.

Dataset bytes, absolute roots, inode/device, timestamps and permission bits never enter evidence or diagnostics. Identical bytes and logical references at different roots produce identical artifacts and Plan digests. Only the selected file is read; no neighboring files, sidecars, dataset copies, writes, caches or watch services are involved.

An explicit selection yields origin=declared, while successfully observed local bytes yield content_origin=static. Static selections keep origin=static and also receive content_origin=static after hashing. Format hints for declarations come only from exact lowercase suffixes .csv, .parquet, .npy, .npz and .joblib; unknown extensions have no hint. No content sniffing occurs. A canonical format hint can change Plan identity even when bytes match; it is not proof of parser behavior.

Interoperability and contract boundary

DVC, MLflow, W&B and external catalogs can become future evidence adapters. They are not core dependencies. No dataset store, runtime capture, metric collection, randomness/environment capture, experiment execution, history or CLI is added. The minimal core needs no pandas, NumPy, joblib, pyarrow or remote storage SDK.

The additive InputArtifact fields preserve old Plan/Run v1 canonical bytes when absent. See experiment evidence v1 for provenance constraints and schema limits. Fingerprinted inputs are reusable Plan evidence; future Run producers must separately supply runtime observations.