Reshot

Open benchmark methodology

Release Truth Benchmark Methodology

A preregistered method for evaluating linkage between changed product outcomes, verification, Decisions, documentation, support, media, and release records.

The Release Truth Benchmark is a preregistered method for evaluating whether public software release processes connect changed product outcomes to verification, Decisions, documentation, support, training, demos, media, and portable release evidence.

This page publishes methodology only. No company scores or observations are published until the sample, raw data, evidence snapshots, and reproducible evaluator are complete. LLMs may classify sourced evidence; they may not simulate missing observations.

Research question

How completely does a public release process link changed user outcomes to exact verification evidence, attributable decisions, outward-content readiness, unresolved risk, and durable machine-readable release records?

Unit of analysis

One publicly documented software release with an identifiable version or date, source or change record where public, release notes, verification evidence where public, and outward guidance affected by the release.

Do not compare private internal evidence or infer unavailable practices.

Proposed sample

Use open-source SaaS or developer tools with reproducible public releases and explicit project licenses. Define inclusion and exclusion before collection. Record product, repository, release identifier, observation date, source URLs, and whether evidence is absent or merely non-public.

Absence of public evidence is not proof that an internal process does not exist.

Dimensions

Score separate evidence dimensions, never one opaque quality score:

  • release identity;
  • changed outcome identification;
  • Journey or workflow context;
  • functional verification linkage;
  • visual verification linkage;
  • accessibility evidence linkage;
  • attributable decision or approval;
  • risk and exception transparency;
  • documentation/support impact;
  • training/demo/media impact;
  • delivery receipts or observable versioning;
  • machine-readable integrity and stable history.

Each dimension uses defined evidence levels such as absent, referenced, directly linked, and machine-verifiable.

Evidence rules

Prefer primary project sources: repositories, release pages, documentation, changelogs, test reports, signed artifacts, public schemas, and official status. Record retrieval time and bounded excerpts. Preserve source URLs and hashes where licensing permits.

Do not score marketing claims without supporting artifacts. Do not infer a failed practice from a broken source until rechecked.

LLM role

An LLM can extract candidate claims, map source passages to dimension definitions, identify conflicts, and prepare a reviewer packet. A separate evaluator verifies source entailment. High-risk or ambiguous observations require two reviewers or deterministic evidence.

The model cannot create a result for an unavailable source, estimate private coverage, or approve its own analysis.

Reproducibility

Publish:

  • sample manifest;
  • source dossier per release;
  • dimension rubric and version;
  • raw observations;
  • reviewer decisions and disagreement;
  • derived tables;
  • scripts and environment;
  • limitations;
  • dated changelog.

Stable dataset URLs identify the current benchmark; versioned files preserve historical releases.

Validation

Double-code at least a preregistered subset. Report agreement by dimension and disclose adjudication. Test the evaluator against positive, absent, ambiguous, stale, and contradictory examples. Fail when a source does not entail the claim.

Derived reporting

Publish dimension distributions, evidence-level counts, source coverage, reviewer agreement, and missing-public-evidence rates. Avoid a single leaderboard unless the rubric proves that aggregation is meaningful. If an overall index is ever added, publish the weights, sensitivity analysis, and every component score.

Audience adaptations for engineering/QA, documentation/support, and product marketing must derive from the same frozen dataset. They may emphasize different dimensions but cannot create separate observations.

Correction process

Allow projects to submit a primary-source correction. Record the request, reviewer decision, affected observation, and benchmark version. Correct material errors without rewriting historical released datasets. Publish a new version and changelog.

Competitor or project outreach must describe the methodology accurately and cannot offer favorable scoring in exchange for links, access, or promotion.

Limitations

Public evidence favors projects with open engineering practices. Release complexity differs. Machine-verifiable artifacts do not guarantee good product outcomes. Lack of public records does not establish lack of internal discipline. The benchmark cannot measure private incident response, customer impact, or team effort without permissioned data.

Publication gate

Do not publish rankings until:

  1. the sample is frozen;
  2. the rubric and code are public;
  3. every observation has a primary source;
  4. independent review is complete;
  5. raw and derived data reconcile;
  6. limitations are visible;
  7. corrections and appeals have a documented process.

Reshot disclosure

Reshot is both the publisher and a system that implements Release Book contracts. Its own entry must use the same rubric, cite public evidence, disclose internal dogfood, and separate local implementation from production/external receipts. The benchmark must not be designed to make Reshot the inevitable winner.

Planned first dataset

Version 1 will be created only after the collection harness can capture and hash current public sources and reviewers can reproduce results. Until then, this methodology remains a citable preregistration, not a benchmark result.

Every correction remains attributable and dated.

Historical versions remain available.

Use the Open Release Book Manifest, Release Evidence for Product Journeys, and Release Book Verifier to inspect the reference implementation boundaries.

Start with one release-critical Journey.

Define the outcome, declare its state, run it, inspect the evidence, and record the Decision before expanding coverage.