21 August 2026
Checking the sentences, not just the citations
Suppose every citation in an answer is real, every quote is verbatim, and every pinpoint is right. The answer can still be wrong.
Section 105I empowers the Authority to fix fees, and to charge them retrospectively. [Cap. 132, s. 105I] "may fix fees for the use of any facility"
The citation is genuine. The quote is genuine. The second half of the sentence is not in the passage, and nothing about the citation apparatus reveals that.
This is the gap between verified and supported, and it is where a lawyer gets hurt: the citation machinery is exactly what makes the sentence look checked.
What the check does
Before the answer is published, each sentence that states law is taken together with the evidence cited in that sentence, and judged against it. Three things get reported:
| overreach | Cites evidence on the right subject but asserts more than it establishes — a broader rule, an extra element, a certainty the passage does not carry |
| contradicted | The cited evidence says something materially different |
| unsupported | Asserts a proposition of law with nothing behind it that bears on that proposition |
And, just as deliberately, what is not reported: a fair paraphrase worded differently; a sentence already hedged to what the evidence supports; a heading, a signpost or a restatement of the question; a sentence the answer already marks as unverified; anything the checker merely suspects is incomplete.
Silence is the correct output for a sound answer.
It rewrites rather than warns
This is the part that took us longest to get right.
The earlier design ran after the answer was finished and appended a hedge to anything it doubted — not confirmed against the source. Two problems. It hedged correct statements: in one measured run, three of four labels landed on sentences that were right, which teaches a reader to distrust the whole answer. And a hedge is a worse outcome than a correction, because the sentence was still fixable.
So the check moved earlier, to the moment the answer is composed and the mapping from sentence to evidence still exists exactly. A flagged sentence goes back with an instruction to cite evidence that does establish it, narrow it to what the evidence supports, or drop it — explicitly not to add a hedge.
One more property matters: the answer that passed every deterministic check is held while the rewrite is attempted. An improvement must never be able to destroy the thing it was improving. If the rewrite is slow, fails, or never arrives, the held answer ships.
What it changed
Two arms, same branch, same hour, differing only in whether the check ran (2026-08-20):
| check off | check on | |
|---|---|---|
| claims the run's own evidence did not back | 22.9% | 5.1% |
| answers that reached the end with no evidence at all | 23 | 0 |
| supported claims per answer | 8.53 | 8.67 |
| quality gate | 14/15 | 14/15 |
The third row is the one we watch hardest. It would be trivial to drive unbacked claims to zero by saying less; supported claims went up, so the answers did not get thinner. It costs about 34 seconds and roughly a cent a turn.
The honest bound
"Unbacked" means the run never grounded the sentence. It does not mean the sentence is false. The checker sees the evidence this chat holds, which is precisely what any check has unless it re-opens every source document — and re-opening them costs the same clock the research needs.
So this catches the answer asserting more than it read. It does not catch the answer being wrong about Hong Kong law in a way its own evidence cannot reveal, and it is not a substitute for a lawyer who has read the authorities.
None of this is legal advice, creates a solicitor–client relationship, or replaces reading the authorities.