Frontier AI Assurance
Connect a configured evaluation to its declared inputs, policy, and reported outcome.
Verification checks record integrity and attribution; source truth, effectiveness, and coverage remain separate questions.
Red-teaming, adversarial testing, and safety evaluations produce real artifacts inside the lab — transcripts, eval harness outputs, model-behavior traces. At the handoff, those artifacts may not bind the declared test, configuration, signer, and reported result in a form another reviewer can check. Signed records can add that integrity layer without establishing the evaluation’s truth or completeness.
The EU AI Act requires “appropriate logging.” The White House voluntary commitments mention “red teaming.” None of them say what counts as proof.
This gap exists industry-wide: evaluations produced inside the lab rarely travel intact to the regulator or customer.
GLACIS creates signed evidence for in-scope safety evaluations, with configured receipts carrying commitments rather than raw test payloads.
For a configured red-team test or safety evaluation, GLACIS can instrument the in-scope event and generate a signed record of the reported control outcome.
Configured prompts and outputs can be hashed locally, with the receipt carrying their commitments rather than the raw content.
The operator signs the record. Portal-minted receipts may also carry a Glacis service-operated witness countersignature and inclusion proof; SDK or self-hosted receipts may be operator-signed only. Either path can be checked at /verify.
from glacis import Glacis glacis = Glacis() # Your red team evaluation receipt = glacis.attest( service_id="safety-eval", operation_type="red_team_test", input={"prompt": adversarial_prompt}, # Local hash on this configured path output={"response": model_output}, # Verify surrounding data flows separately metadata={ "model": "llama-3-70b", "test_suite": "harmbench", "evaluator": "safety-team" } ) # Share this with auditors, regulators, the public print(receipt.verification_url) # → https://www.glacis.io/verify/att_7f3k...
A signed report that declared adversarial prompts were evaluated by the named model at a recorded time, with the covered result. The signature binds the report; trusted collection and coverage evidence establish what actually ran.
Auditors can verify without seeing your test data.
A safety claim can link to a signed attestation identifying the declared benchmark, model, time, and reported result. Reviewers can check those fields and their signer without treating the record as independent proof of the test’s validity or completeness.
Model cards with teeth.
Record the covered model, data commitment, and time in a form others can check. Timestamped, signed, verifiable.
For papers, audits, or your own records.
Commit to declared training-data provenance while excluding the underlying dataset from the portable record. The commitment does not establish origin, rights, or completeness and should be paired with the evidence required for a dataset review.
Make a declared provenance claim checkable without copying the dataset into the record.
A data-minimizing deployment keeps prompts and outputs local while portable records carry only the bounded evidence a reviewer needs. The deployment’s actual data flow remains part of the review.
Prompts and outputs
Protected content can stay local; the record carries a cryptographic commitment and bounded metadata.
Training data
Can remain local on a configured path; verify model, telemetry, and storage flows.
Model weights
Can remain local on a configured path; verify model, telemetry, and storage flows.
A signed, bounded record that can exclude the underlying test content.
Open-source SDK. Supported signing and verification. Explicit evidence boundaries.