Su Tung.
Sign in

← Blog

21 August 2026

Checking the sentences, not just the citations

A real case, a verbatim quote and a correct pinpoint still do not tell you whether the sentence in front of them follows from the passage. That is a separate check, and it runs before the answer reaches you.

Suppose every citation in an answer is real, every quote is verbatim, and every pinpoint is right. The answer can still be wrong.

Section 105I empowers the Authority to fix fees, and to charge them retrospectively. [Cap. 132, s. 105I] "may fix fees for the use of any facility"

The citation is genuine. The quote is genuine. The second half of the sentence is not in the passage, and nothing about the citation apparatus reveals that.

This is the gap between verified and supported, and it is where a lawyer gets hurt: the citation machinery is exactly what makes the sentence look checked.

What the check does

Before the answer is published, each sentence that states law is taken together with the evidence cited in that sentence, and judged against it. Three things get reported:

overreachCites evidence on the right subject but asserts more than it establishes — a broader rule, an extra element, a certainty the passage does not carry
contradictedThe cited evidence says something materially different
unsupportedAsserts a proposition of law with nothing behind it that bears on that proposition

And, just as deliberately, what is not reported: a fair paraphrase worded differently; a sentence already hedged to what the evidence supports; a heading, a signpost or a restatement of the question; a sentence the answer already marks as unverified; anything the checker merely suspects is incomplete.

Silence is the correct output for a sound answer.

It rewrites rather than warns

This is the part that took us longest to get right.

The earlier design ran after the answer was finished and appended a hedge to anything it doubted — not confirmed against the source. Two problems. It hedged correct statements: in one measured run, three of four labels landed on sentences that were right, which teaches a reader to distrust the whole answer. And a hedge is a worse outcome than a correction, because the sentence was still fixable.

So the check moved earlier, to the moment the answer is composed and the mapping from sentence to evidence still exists exactly. A flagged sentence goes back with an instruction to cite evidence that does establish it, narrow it to what the evidence supports, or drop it — explicitly not to add a hedge.

One more property matters: the answer that passed every deterministic check is held while the rewrite is attempted. An improvement must never be able to destroy the thing it was improving. If the rewrite is slow, fails, or never arrives, the held answer ships.

What it changed

Two arms, same branch, same hour, differing only in whether the check ran (2026-08-20):

check offcheck on
claims the run's own evidence did not back22.9%5.1%
answers that reached the end with no evidence at all230
supported claims per answer8.538.67
quality gate14/1514/15

The third row is the one we watch hardest. It would be trivial to drive unbacked claims to zero by saying less; supported claims went up, so the answers did not get thinner. It costs about 34 seconds and roughly a cent a turn.

The honest bound

"Unbacked" means the run never grounded the sentence. It does not mean the sentence is false. The checker sees the evidence this chat holds, which is precisely what any check has unless it re-opens every source document — and re-opening them costs the same clock the research needs.

So this catches the answer asserting more than it read. It does not catch the answer being wrong about Hong Kong law in a way its own evidence cannot reveal, and it is not a substitute for a lawyer who has read the authorities.

None of this is legal advice, creates a solicitor–client relationship, or replaces reading the authorities.

Checking the sentences, not just the citations — Su Tung