The errors in a generated research report are not spread evenly. They cluster where the decision is, and the rule you already use to catch them is aimed at the wrong variable.
TL;DR: A generated research report is right only if every link in it holds, so accuracy falls away as the steps pile up. In one benchmark’s public filings, the same models score 83% on answering what a document says and 35% on answering what follows from it. Verifying the recent and the obscure misses it; decisions rest on public material. The mode that most resembles diligence fabricates the most citations. Mark every claim retrieved or composed, and re-source only the composed ones before money moves.