# Evaluate your own workflow

This is a nine-task, public reference exercise. No model has been run or scored by SD in this package.
It supports repeatable evaluation setup; it is not a representative or held-out AI benchmark.

1. Keep a clean copy of this package. Record the model/provider, exact version, UTC run date,
   prompt, parameters and any tool use. Use evaluation-submission-template.json for metadata.
2. Provide evaluation-inputs.json to your chosen model or workflow. This file contains the
   questions and cited source values, without the answerGuide or expected-results files.
   Do not provide evaluation-tasks.json: that is the answer key. Preserve the model's raw output.
3. Map the output to the template without correcting its answers. Use null for unanswered items.
   Save it as my-run.json. Run: python evaluate.py my-run.json
4. Inspect evaluation-report.json. All nine tasks stay in the denominator, including omissions.
   A task passes only when its typed value matches and all required source locators are present,
   with no unrecognized citation locators. False is not the integer zero. Invalid/duplicate task IDs
   are rejected. Metadata is self-reported; the scorer does not attest which model produced answers.
5. Review the original prose separately: source support, invented claims, temporal context,
   missing-history handling, and correct separation of arithmetic from legal obligations.
   Exact locator matching does not establish citation entailment. Save reviewer decisions and reasons.
6. If reporting results, include every task, the raw outputs, prompt, version, date, package hash,
   failures and manual review. Label them results on these nine public reference tasks only.
   Do not describe 9/9 as general accuracy, independent validation or evidence of real-world benefit.

For a stronger study, define a larger independent held-out sample and scoring protocol before
running models, include difficult and missing-data cases, and compare against a documented baseline.
Record timing only when measured. Never convert record counts into model-performance claims.
