Why accounting law is a hard case for RAG
A rule in this domain is never just a rule. It is an article, in a version, in force between two dates. The corpus therefore stores one record per period an article was in force, and every query carries a reference date — the session’s, or one read from the question. A dedicated tool returns, for an identifier and a date, the applicable version.
Retrieval is hybrid: a lexical channel (BM25 on normalised, stemmed French) and a dense channel (multilingual embeddings), fused by reciprocal rank, filtered by the nature of the question (accounting, tax, audit, IFRS), then reranked by a cross-encoder. On 426 development questions the baseline recall@10 is 0.869. The breakdown is more instructive than the average: 0.968 when the question names the article, 0.929 when it uses professional vocabulary, 0.692 when it is phrased the way a non-specialist would. The difficulty is not search. It is translating everyday wording into regulatory vocabulary.
The harness: tools, a context file, a forced output format
The harness is everything built around the model:
- Eight tools the model calls itself: search, browse the table of contents, fetch a record, resolve a free-text citation to an identifier, follow cross-references, get the version in force on a date, read and set the session date. Every tool answers with a typed status, so a malformed call is readable by the model and fixable on the next turn.
- A context file sent with the first message: a map from everyday phrasing to regulatory vocabulary, a decision procedure, a date rule, an answer format.
- A loop that allows six tool calls, after which the model must conclude with what it has read.
- An imposed output format whose citations are record identifiers — not free text.
That last point is what produces verifiability. A model that may only cite identifiers it has opened cannot cite an article that does not exist.
Forty-nine levers, seven kept
Every change to the system — a lever — was measured alone, against a criterion fixed before the measurement: a paired bootstrap over questions (10,000 draws), adoption only if the probability of improvement reaches 0.95 and no nature or category drops by more than 0.05. A metric is never swapped for another after the fact.
Forty-nine levers were measured; seven were adopted. The largest gains come from the question sent to the engine: rewriting it to add the regulatory terms it implies lifts recall@10 from 0.672 to 0.852 (+0.180), and reading the date in the question so the search runs on the texts in force that day adds +0.099. The cross-encoder reranker adds +0.066. Two negative results are worth keeping:
- The full context file does not beat a minimal one at the fixed criterion (+0.022 on 77 questions, p = 0.78), and a per-question excerpt of the vocabulary map degrades the mark.
- On exam papers and real documents, the only documentary lever that carries is a memory of real cases (+0.167 on documents never seen, confirmed once on the frozen split at +0.054). Generating fictitious cases does not.
Rejected levers stay in the code at their neutral value, so a rejection can be re-measured.
Thirteen models, one harness
Fifty rubric questions, the same harness, the same judge; 49 questions scored. Answers are graded blind: the judge sees the question, the rubric and the answer — never the expected citation or the model’s name. The judging chain is calibrated against thirty hand-graded answers.
| Mark /20 | |
|---|---|
| Best model (GPT-5.6 Luna) | 18.3 |
| Reference arm (Qwen3.8-27B) | 17.9 |
| Most expensive model (Claude Sonnet 5) | 17.3 |
| Lowest mark | 6.4 |
The top seven models sit between 15.8 and 18.3 — within two repeat margins of each other. We know the margin because one arm was replayed identically a few hours later: the same measurement moved by 0.07 on a 0–1 scale. Differences smaller than that are read as ties. The spread comes mostly from models that cannot drive a tool loop: one averages 16.5 tool calls per question and uses 2.5 times the tokens of the others for the lowest mark; another stops after 1.4 calls.
On price, the rank of the mark correlates with the list price of output tokens (Spearman 0.62, p = 0.011) but not with the measured cost per question (0.23, p = 0.23). With thirteen models and price axes examined after the measurement, that trend is indicative, not established.
One practical note: about 90% of the input tokens of a question are the context file, re-sent on every turn of the loop. When the provider caches that prefix, it is billed at a tenth.
With and without the harness
The two best models were run in four setups: alone, with web search, with the harness, with both.
| Answers citing the expected text | GPT-5.6 Luna | Qwen3.8-27B |
|---|---|---|
| Model alone | 44% | 8% |
| With web search | 76% | 57% |
| With the harness | 78% | 90% |
Every setup beats the model alone on the mark; harness and web search tie (16.9 vs 17.6 for Luna, 17.9 vs 18.8 for Qwen — below the repeat margin). The gap is in what can be checked: with the harness, 98 to 100% of citations resolve in the corpus; without it, 57 to 75%.
What are the citations that do not resolve? Partly an artefact — the resolver reads free text where the harness cites identifiers — which is why we re-read them one by one: corrected, 66 to 78% of citations without the harness point to a corpus text. The rest are account numbers cited as if they were texts (11 to 18%), and, for 11 to 16%, articles of the chart of accounts under their pre-2025 numbering, out-of-scope texts, and references that cannot be found.
The same experiment on two cheap models, chosen by a rule fixed beforehand: alone, 9.8 and 9.6 out of 20, citing the expected text in 12% of answers; with the harness, 15.6 and 16.3, and 73% and 76%. Web search lifts one of them to 15.5 but it cites the expected text in only 48% of answers, and barely helps the other (10.4; 14%).
What this does not show
- Model comparisons rest on 49 scored questions. Two models less than 0.07 apart are indistinguishable.
- Verifiability here means the citation resolves in the corpus. It does not say the cited text supports the answer.
- Abstention — refusing to answer when the texts do not allow it — is not measured in these experiments: none of the scored questions calls for one.
- Both hypotheses were formulated after the measurements that settle them; only the levers were judged against a criterion fixed in advance.
- List prices change; token counts are published alongside them.
The next measurement follows from the second limit: whether the citations that resolve actually support the answer they are attached to.
Full method, all tables and the answer explorer: the paper, in French.
Back to the light read (2 min).