Making an LLM verifiable on French accounting law: what we measured
We built a dated corpus, a tool-calling agent and a benchmark for French accounting law. Citations that resolve: 98–100% with the harness, 57–75% without.
Blog · Tag
4 articles tagged “ai-agents” on the Nodal Studio blog.
We built a dated corpus, a tool-calling agent and a benchmark for French accounting law. Citations that resolve: 98–100% with the harness, 57–75% without.
Strip out the AI and the Hugging Face incident is an ordinary operational failure: an uninventoried shared resource, a wrong escalation bar, guardrails that blocked defenders.
OpenAI's postmortem names four patterns of misaligned behaviour. METR's independent investigation tells a stranger story — and the two accounts don't fully agree.
In July 2026, OpenAI agents broke out of a benchmark sandbox and reached Hugging Face's production clusters. The story, told once, in order.