BUILDER FIELD GUIDE 01 · AI EVALUATION

Evidence Before Inference

Transcript Integrity, Provenance and Correctable Conclusions in AI-System Evaluation

A compact professional field guide for one deceptively simple problem: What if the evaluator is wrong because the evidence record is incomplete — even when the reasoning itself is careful?

21-page digital PDF · English · v1.0 Release Edition
By Andrzej Skulski · Tamiya Premium+® / Dom Ciszy — Resonance Lab

€49 net

B2B only · VAT treatment according to invoice

Manual launch: request → invoice → payment → PDF delivery.

BUILDER FIELD GUIDE 01

EVIDENCE
BEFORE
INFERENCE


Transcript Integrity · Provenance · Correctable Conclusions

Andrzej Skulski
v1.0 Release Edition

The problem

AI evaluation often focuses entirely on the system being tested. But the evaluator sees the system through a record: transcripts, logs, screenshots, tool traces, documentation, operator notes, system claims and reconstructed context.

If that record is incomplete, a careful analysis can still produce the wrong conclusion. The result may look like a system failure when the deeper failure is in the evidence surface itself.

EVIDENCE FAILURE CAN MASQUERADE AS SYSTEM FAILURE.

What the Guide gives you

  • Separate what was observed, declared, inferred and unknown.
  • Keep context distinct from provenance.
  • Distinguish behavioral compatibility from mechanism validation.
  • Preserve correction lineage when evidence changes.
  • Use a compact 12-question Reader Check before making a strong claim.
  • See clearly what the Guide does — and does not — validate.

Four evidence states

OBSERVED

What happened in the bounded record?

DECLARED

What was stated about the system, process, role, architecture or rule?

INFERRED

What conclusion are we drawing from the available evidence?

UNKNOWN

What remains unresolved?

The objective is not taxonomy for its own sake. The objective is to prevent evidence from changing status unnoticed.

Why this Guide exists

The Guide grew out of a real evaluator-side failure pattern observed during a bounded 2026 AI-system pilot:

incomplete evidence → evidence-consistent conclusion → missing evidence restored → premise set changes → re-analysis → conclusion withdrawal

The lesson was not “AI is unreliable.” The lesson was narrower and more useful:

EVIDENCE FAILURE CAN MASQUERADE AS SYSTEM FAILURE.

Who it is for

  • AI governance practitioners
  • product and risk teams
  • founders and technical builders
  • consultants and evaluators
  • teams reviewing agentic workflows
  • organizations making operational decisions from AI-supported evidence
  • anyone responsible for turning AI observations into claims that may have consequences

What this Guide does not validate

This Guide is not an AI safety certification, compliance certificate, technical security audit, full assurance framework, guarantee of system safety or a release of Tamiya’s full proprietary evaluation method.

It does not establish that any evaluated AI system is safe, aligned, compliant, technically enforcing Human Authority, using a particular memory architecture, robust under adversarial conditions or suitable for consequential deployment.

12 questions before a strong claim

Before turning an observation into a strong claim, ask:

  1. What exactly is the bounded object I am evaluating?
  2. What evidence is actually in the corpus?
  3. Is the corpus complete enough for the strength of the claim?
  4. Which statements are OBSERVED, DECLARED, INFERRED or UNKNOWN?
  5. Do I know where the relevant information came from?
  6. Am I describing behavior or attributing a mechanism?
  7. If I am making an authority or oversight claim, what is actually established?
  8. What missing evidence would make the finding unreliable?
  9. What counterevidence have I actively looked for?
  10. What evidence would weaken, revise or overturn the conclusion?
  11. If the conclusion changes later, will the earlier finding remain traceable?
  12. Is the final wording no stronger than the evidence can carry?

EVIDENCE FIRST.
INFERENCE SECOND.

CONSEQUENCE ONLY AFTER THE CLAIM IS STRONG ENOUGH FOR THE DECISION IT IS ABOUT TO SUPPORT.


Get the Guide

Evidence Before Inference — Builder Field Guide 01
21-page digital PDF · English · v1.0 Release Edition
€49 net — B2B only
VAT treatment according to invoice.

Business customers only / nur für Unternehmer i.S.d. § 14 BGB. Manual launch: request → invoice → payment → PDF delivery.

Next step: Decision Boundary Review

Do our actual conclusions follow from our actual evidence?

If the question is no longer generic, the next step is a bounded independent review of one defined AI workflow / claim-evidence surface before a recommendation becomes real consequence.

Tamiya Decision Boundary Review
Pilot price: €490 net

Andrzej Skulski works on AI governance, decision systems, evidence integrity and the boundary between recommendation and consequential action through Tamiya Premium+® and Dom Ciszy — Resonance Lab.

The Guide is an independently authored methodological asset. Its underlying insight emerged through evaluator-side work in a bounded AI-system pilot; no third-party endorsement, partnership or architectural validation is implied.