AI & Automation

Making an LLM verifiable on French accounting law: what we measured

We built a dated corpus, a tool-calling agent and a benchmark for French accounting law. Citations that resolve: 98–100% with the harness, 57–75% without.

Why accounting law is a hard case for RAG

A rule in this domain is never just a rule. It is an article, in a version, in force between two dates. The corpus therefore stores one record per period an article was in force, and every query carries a reference date — the session’s, or one read from the question. A dedicated tool returns, for an identifier and a date, the applicable version.

Retrieval is hybrid: a lexical channel (BM25 on normalised, stemmed French) and a dense channel (multilingual embeddings), fused by reciprocal rank, filtered by the nature of the question (accounting, tax, audit, IFRS), then reranked by a cross-encoder. On 426 development questions the baseline recall@10 is 0.869. The breakdown is more instructive than the average: 0.968 when the question names the article, 0.929 when it uses professional vocabulary, 0.692 when it is phrased the way a non-specialist would. The difficulty is not search. It is translating everyday wording into regulatory vocabulary.

The harness: tools, a context file, a forced output format

The harness is everything built around the model:

  • Eight tools the model calls itself: search, browse the table of contents, fetch a record, resolve a free-text citation to an identifier, follow cross-references, get the version in force on a date, read and set the session date. Every tool answers with a typed status, so a malformed call is readable by the model and fixable on the next turn.
  • A context file sent with the first message: a map from everyday phrasing to regulatory vocabulary, a decision procedure, a date rule, an answer format.
  • A loop that allows six tool calls, after which the model must conclude with what it has read.
  • An imposed output format whose citations are record identifiers — not free text.

That last point is what produces verifiability. A model that may only cite identifiers it has opened cannot cite an article that does not exist.

Forty-nine levers, seven kept

Every change to the system — a lever — was measured alone, against a criterion fixed before the measurement: a paired bootstrap over questions (10,000 draws), adoption only if the probability of improvement reaches 0.95 and no nature or category drops by more than 0.05. A metric is never swapped for another after the fact.

Forty-nine levers were measured; seven were adopted. The largest gains come from the question sent to the engine: rewriting it to add the regulatory terms it implies lifts recall@10 from 0.672 to 0.852 (+0.180), and reading the date in the question so the search runs on the texts in force that day adds +0.099. The cross-encoder reranker adds +0.066. Two negative results are worth keeping:

  • The full context file does not beat a minimal one at the fixed criterion (+0.022 on 77 questions, p = 0.78), and a per-question excerpt of the vocabulary map degrades the mark.
  • On exam papers and real documents, the only documentary lever that carries is a memory of real cases (+0.167 on documents never seen, confirmed once on the frozen split at +0.054). Generating fictitious cases does not.

Rejected levers stay in the code at their neutral value, so a rejection can be re-measured.

Thirteen models, one harness

Fifty rubric questions, the same harness, the same judge; 49 questions scored. Answers are graded blind: the judge sees the question, the rubric and the answer — never the expected citation or the model’s name. The judging chain is calibrated against thirty hand-graded answers.

Mark /20
Best model (GPT-5.6 Luna)18.3
Reference arm (Qwen3.8-27B)17.9
Most expensive model (Claude Sonnet 5)17.3
Lowest mark6.4

The top seven models sit between 15.8 and 18.3 — within two repeat margins of each other. We know the margin because one arm was replayed identically a few hours later: the same measurement moved by 0.07 on a 0–1 scale. Differences smaller than that are read as ties. The spread comes mostly from models that cannot drive a tool loop: one averages 16.5 tool calls per question and uses 2.5 times the tokens of the others for the lowest mark; another stops after 1.4 calls.

On price, the rank of the mark correlates with the list price of output tokens (Spearman 0.62, p = 0.011) but not with the measured cost per question (0.23, p = 0.23). With thirteen models and price axes examined after the measurement, that trend is indicative, not established.

One practical note: about 90% of the input tokens of a question are the context file, re-sent on every turn of the loop. When the provider caches that prefix, it is billed at a tenth.

With and without the harness

The two best models were run in four setups: alone, with web search, with the harness, with both.

Answers citing the expected textGPT-5.6 LunaQwen3.8-27B
Model alone44%8%
With web search76%57%
With the harness78%90%

Every setup beats the model alone on the mark; harness and web search tie (16.9 vs 17.6 for Luna, 17.9 vs 18.8 for Qwen — below the repeat margin). The gap is in what can be checked: with the harness, 98 to 100% of citations resolve in the corpus; without it, 57 to 75%.

What are the citations that do not resolve? Partly an artefact — the resolver reads free text where the harness cites identifiers — which is why we re-read them one by one: corrected, 66 to 78% of citations without the harness point to a corpus text. The rest are account numbers cited as if they were texts (11 to 18%), and, for 11 to 16%, articles of the chart of accounts under their pre-2025 numbering, out-of-scope texts, and references that cannot be found.

The same experiment on two cheap models, chosen by a rule fixed beforehand: alone, 9.8 and 9.6 out of 20, citing the expected text in 12% of answers; with the harness, 15.6 and 16.3, and 73% and 76%. Web search lifts one of them to 15.5 but it cites the expected text in only 48% of answers, and barely helps the other (10.4; 14%).

What this does not show

  • Model comparisons rest on 49 scored questions. Two models less than 0.07 apart are indistinguishable.
  • Verifiability here means the citation resolves in the corpus. It does not say the cited text supports the answer.
  • Abstention — refusing to answer when the texts do not allow it — is not measured in these experiments: none of the scored questions calls for one.
  • Both hypotheses were formulated after the measurements that settle them; only the levers were judged against a criterion fixed in advance.
  • List prices change; token counts are published alongside them.

The next measurement follows from the second limit: whether the citations that resolve actually support the answer they are attached to.

Full method, all tables and the answer explorer: the paper, in French.

Back to the light read (2 min).