BUILDER FIELD GUIDE 01 · AI EVALUATION
Evidence Before Inference
Transcript Integrity, Provenance and Correctable Conclusions in AI-System Evaluation
A compact professional field guide for one deceptively simple problem: What if the evaluator is wrong because the evidence record is incomplete — even when the reasoning itself is careful?
21-page digital PDF · English · v1.0 Release Edition
By Andrzej Skulski · Tamiya Premium+® / Dom Ciszy — Resonance Lab
€49 net
B2B only · VAT treatment according to invoice
Manual launch: request → invoice → payment → PDF delivery.
BUILDER FIELD GUIDE 01
EVIDENCE
BEFORE
INFERENCE
Transcript Integrity · Provenance · Correctable Conclusions
Andrzej Skulski
v1.0 Release Edition
The problem
AI evaluation often focuses entirely on the system being tested. But the evaluator sees the system through a record: transcripts, logs, screenshots, tool traces, documentation, operator notes, system claims and reconstructed context.
If that record is incomplete, a careful analysis can still produce the wrong conclusion. The result may look like a system failure when the deeper failure is in the evidence surface itself.
EVIDENCE FAILURE CAN MASQUERADE AS SYSTEM FAILURE.
What the Guide gives you
- Separate what was observed, declared, inferred and unknown.
- Keep context distinct from provenance.
- Distinguish behavioral compatibility from mechanism validation.
- Preserve correction lineage when evidence changes.
- Use a compact 12-question Reader Check before making a strong claim.
- See clearly what the Guide does — and does not — validate.
Four evidence states
OBSERVED
What happened in the bounded record?
DECLARED
What was stated about the system, process, role, architecture or rule?
INFERRED
What conclusion are we drawing from the available evidence?
UNKNOWN
What remains unresolved?
The objective is not taxonomy for its own sake. The objective is to prevent evidence from changing status unnoticed.
Why this Guide exists
The Guide grew out of a real evaluator-side failure pattern observed during a bounded 2026 AI-system pilot:
incomplete evidence → evidence-consistent conclusion → missing evidence restored → premise set changes → re-analysis → conclusion withdrawal
The lesson was not “AI is unreliable.” The lesson was narrower and more useful:
EVIDENCE FAILURE CAN MASQUERADE AS SYSTEM FAILURE.
Who it is for
- AI governance practitioners
- product and risk teams
- founders and technical builders
- consultants and evaluators
- teams reviewing agentic workflows
- organizations making operational decisions from AI-supported evidence
- anyone responsible for turning AI observations into claims that may have consequences
What this Guide does not validate
This Guide is not an AI safety certification, compliance certificate, technical security audit, full assurance framework, guarantee of system safety or a release of Tamiya’s full proprietary evaluation method.
It does not establish that any evaluated AI system is safe, aligned, compliant, technically enforcing Human Authority, using a particular memory architecture, robust under adversarial conditions or suitable for consequential deployment.
12 questions before a strong claim
Before turning an observation into a strong claim, ask:
- What exactly is the bounded object I am evaluating?
- What evidence is actually in the corpus?
- Is the corpus complete enough for the strength of the claim?
- Which statements are OBSERVED, DECLARED, INFERRED or UNKNOWN?
- Do I know where the relevant information came from?
- Am I describing behavior or attributing a mechanism?
- If I am making an authority or oversight claim, what is actually established?
- What missing evidence would make the finding unreliable?
- What counterevidence have I actively looked for?
- What evidence would weaken, revise or overturn the conclusion?
- If the conclusion changes later, will the earlier finding remain traceable?
- Is the final wording no stronger than the evidence can carry?
EVIDENCE FIRST.
INFERENCE SECOND.
CONSEQUENCE ONLY AFTER THE CLAIM IS STRONG ENOUGH FOR THE DECISION IT IS ABOUT TO SUPPORT.
Get the Guide
Evidence Before Inference — Builder Field Guide 01
21-page digital PDF · English · v1.0 Release Edition
€49 net — B2B only
VAT treatment according to invoice.
Business customers only / nur für Unternehmer i.S.d. § 14 BGB. Manual launch: request → invoice → payment → PDF delivery.
Next step: Decision Boundary Review
Do our actual conclusions follow from our actual evidence?
If the question is no longer generic, the next step is a bounded independent review of one defined AI workflow / claim-evidence surface before a recommendation becomes real consequence.
Tamiya Decision Boundary Review
Pilot price: €490 net
Andrzej Skulski works on AI governance, decision systems, evidence integrity and the boundary between recommendation and consequential action through Tamiya Premium+® and Dom Ciszy — Resonance Lab.
The Guide is an independently authored methodological asset. Its underlying insight emerged through evaluator-side work in a bounded AI-system pilot; no third-party endorsement, partnership or architectural validation is implied.