Large language models answer accounting, tax and audit questions fluently. Fluency tells you nothing about whether the answer can be trusted. In this field an answer depends on a text, on a version of that text, and on a date. A plausible answer resting on the wrong article is a professional risk.
So we asked a narrow question: under what conditions does an LLM give a reliable answer in French accounting law, and at what cost? We defined reliable as two measurable properties. An answer is correct when it establishes what a grading rubric, written before any model answered, expects. It is verifiable when every text it cites exists, in the version in force on the date of the facts, and can be opened by a reader.
To measure that, we built three things: a corpus of 113,105 dated regulatory texts from public sources (the French chart of accounts and ANC regulations, BOFiP tax doctrine, audit standards, the Commercial and Tax codes, EU-adopted IFRS); an agent with eight tools over that corpus — what we call the harness; and a benchmark of 567 practitioner questions with verified citations, plus 763 past-exam questions and 214 questions on real accounting documents.
Three results stand out.
The harness makes citations verifiable, and costs nothing in correctness. With the harness, 98 to 100% of the citations in an answer resolve to a text in the corpus, whatever the model. Without it — model alone or model with web search — 57 to 75% do (66 to 78% once the unmatched citations are re-read one by one). On the marks, the harness and a plain web search are tied, and combining them adds nothing. RAG for accounting law does not make the model smarter; it makes its answer checkable.
The model still matters — but price per token is a poor guide. Across thirteen models on the same harness, marks range from 6.4 to 18.3 out of 20. Marks follow the list price per token, not the measured cost per question, because a weak model burns tokens in loops: the lowest-scoring one makes 16.5 tool calls per question. The best-scoring model costs fourteen times less per question than the most expensive one.
Cheap models become usable. Two inexpensive models that score under 10/20 on their own reach 15.6 and 16.3 with the harness, for €5 and €7 per thousand questions.
One caveat we state plainly: a citation that resolves is not yet a citation that supports the answer. That check has not been measured.
The full paper (in French), with every figure and an explorer of all graded answers, is here: RAG pour le droit comptable français.
Want the sources, timeline and detail? Read the deep dive (5 min).