# Verification-shaped prose is not verification

## An aggregate report from a staged audit of model-generated cross-domain hypotheses

**Version:** 1.0.1  
**Released:** 2026-07-30  
**Author and rights holder:** Acashic Research Division  
**License:** [CC BY 4.0](LICENSE.md) for this report and its authored metadata

## Abstract

We audited a staged language-model workflow for proposing and evaluating
cross-domain research hypotheses. The bound portion contains 224 initially
generated candidates that received a uniform first-pass scorecard. Thirty-eight
dependent derivatives were then produced by prompts requesting additional
mechanism, measurement, falsification, and source-shaped detail. Every
derivative received a higher merit score from its later panel: the descriptive
cross-panel increase ranged from 15 to 39 points (mean 25.01; median 25.75).
The panels changed, so these differences do not identify an elaboration effect.

The decisive gate was source checking. All eight derivatives marked
`FEASIBILITY_ELIGIBLE` and all seven qualified reserves were checked; 0 of 15
passed. Outcomes were eight `DEMOTE_REPAIR`, five `UNRESOLVABLE`, and two
`CONTRADICTED`. These annotations were unattributed and time-unbound, so they
are neither authoritative verification nor empirical falsification. They do,
however, establish a local process failure: score-improving,
verification-shaped prose did not survive the workflow's own grounding gate.

A later source-first screen scheduled 1,200 probes. Its narrative reports 523
candidate packets, 1,346 dataset mentions, and zero exact source locators, but
no row-level receipt binds those reported observations. We therefore report
them only as narrative evidence, not as one measured population. The practical
lesson is methodological: use models to implement and attack frozen,
human-selected targets, not to supply hypotheses of record.

## Study design

The 224 initial records were generated through four provider groups:

| Initial generator route | Records |
| --- | ---: |
| DeepSeek through DeepInfra | 117 |
| Google | 32 |
| OpenAI | 38 |
| xAI | 37 |
| **Total** | **224** |

All 224 received the same first-pass scorecard. A later two-scorer stage covered
153 candidates. Thirty-eight records from its repair track were elaborated and
regated. The elaboration prompt explicitly requested fields resembling the
evaluation criteria. The stage is therefore an operational probe of whether
the scoring apparatus distinguished stronger evidence from better-shaped
prose, subject to the panel-change limitation below.

The before and after scores came from different panels for every derivative.
Accordingly, score changes are descriptive only; no randomization, same-rater
counterfactual, or causal effect estimate is claimed.

## Results

| Stage result | Count |
| --- | ---: |
| Prompt-conditioned derivatives | 38 |
| Positive cross-panel merit differences | 38 |
| `FEASIBILITY_ELIGIBLE` after regate | 8 |
| Qualified reserve after regate | 7 |
| Other repair/reject decisions | 23 |
| Individually source-checked | 15 |
| Source-check gate passes | 0 |

The 38 score increases ranged from 15 to 39 points. Their mean was 25.01 and
their median was 25.75. Later source checks yielded:

| Source-check outcome | Count |
| --- | ---: |
| `DEMOTE_REPAIR` | 8 |
| `UNRESOLVABLE` | 5 |
| `CONTRADICTED` | 2 |
| **Passed gate** | **0** |

Across the 38 selected derivatives, later-panel scores were higher. Among the
15 source-checked candidates, none passed the source-check gate. These are
separate observations over selected records; they do not estimate an
elaboration effect, measure before-and-after grounding, establish a universal
property of models, or establish a general verified-novelty base rate.

The source-first continuation does not repair the evidentiary gap. It
contributes a hash-bound 1,200-row schedule and only narrative-reported
downstream counts: 523 candidate packets, 1,346 dataset mentions, and zero
exact source locators. Those values must not be added to the 224 released
initial rows as though the result were one uniformly observed population.

## Interpretation

The workflow recorded no source-grounded actionable candidate. Later panels
assigned higher scores to the 38 derivatives; the separate grounding stage
recorded zero passes among the 15 selected candidates. Because selection and
scoring panels differ, the retained annotations do not authorize a causal or
thematic claim.

The negative result changes the division's allocation rule:

1. Humans select instrument targets from documented real pain.
2. A real instrument and frozen protocol precede the claim.
3. Models may implement, adversarially review, and blind-score; they do not
   supply the target of record.
4. At most two targets remain in flight.

This is not a claim that language models cannot contribute to research. It is a
claim about role assignment under a high verification bar: implementation and
hostile review are better-supported roles here than de-novo ideation.

## Limitations

- The 224 candidates are a designed campaign, not a random sample of all model
  ideas or scientific domains.
- All 38 elaboration comparisons changed scoring panels, precluding causal
  attribution.
- The 15 source checks have unrecorded reviewer identity, no persisted evidence
  access time, and no authoritative verification status.
- The DT2 downstream observations are narrative-only until bound receipts
  exist.
- Provider and third-party rights do not support public release of the rows.
- No inference should be made from these records about clinical validity,
  individual model rankings, or a population-wide rate of scientific novelty.

## Evidence and data availability

The private evidence candidate is ARD-004 negative corpus RC3:

- manifest SHA-256:
  `ead8d2c026c36139ff5c4483052a1487c686094af87bb6aa1cd5c67d70108b8c`
- package SHA-256:
  `af3bb2b10e2eb1e480c8181b66f6c9fd81b04de0141f72ccf8473fabe1611814`
- bound records: 224 initial, 38 derivative, 15 source-check, and one DT2
  aggregate annotation

The row-level files are not distributed. Their routed model-output terms chain,
third-party evidence prose, and reviewer provenance do not support a blanket
public license. This report contains no row text or model-output excerpt and
does not grant rights in the private evidence package.

## Transparency statement

Language models assisted with implementation, aggregation, drafting, and
adversarial review. Acashic Research Division selected the claims, verified the
reported aggregates against the frozen evidence, accepted responsibility for
the text, and authorized this release. Provider names are used only to describe
the study routes; no affiliation or endorsement is implied.

## Citation

Acashic Research Division. (2026). *Verification-shaped prose is not
verification: An aggregate report from a staged audit of model-generated
cross-domain hypotheses* (Version 1.0.1).
