What it gets wrong
A system that answers labour-law questions and cannot say where it fails has no business in front of an employee. Here are the 68 questions out of 142 that go wrong, and why.
What this must not be used for
- This is not legal advice. No lawyer has validated this system’s answers.
- This is not a public service and does not come from any government body.
- The corpus covers four themes only: contract and trial period, notice and termination, working time and forfait-jours, classification and minimum salaries. Everything else is out of scope.
- No company-level agreement is taken into account, although one can change what the branch agreement says.
- Before acting on an answer, open the cited article on Légifrance and read it.
The failures, sorted by cause
They are not counted together, because they do not have the same remedy. An article never retrieved is an indexing problem that no improvement in writing will fix; a wrong answer from the right article is exactly the opposite.
- rubric-dependentthe answer holds or does not, depending on how strictly it is marked43 · 30.3 %
- false-refusalrefused although the corpus contained the answer10 · 7.0 %
- generation-missthe right article was in front of the model and the answer is wrong10 · 7.0 %
- retrieval-missthe article that settles it never reached the model4 · 2.8 %
- citation-missthe answer is right, the cited source is not1 · 0.7 %
This is why there is no single "accuracy rate" on this site. One number would merge a retrieval-miss with a generation-miss and so would not say what to fix. On a client engagement, that distinction is what decides whether the next sprint is about indexing or about writing.
The judge is not validated, and it says so
Answers are marked automatically by a second model (claude-sonnet-4-5-20250929). For an automatic mark to mean anything, it has to be shown to agree with a competent human. The first calibration scored Cohen’s kappa 0.489 against a non-expert reader — too weak an agreement to validate anything.
There were two options: present an accuracy figure anyway, or say the judge is not validated. The second was chosen, and accuracy is published as a range between a strict and a lenient marking scheme. What depends on no scheme at all: 20.0 % of answers are wrong under both readings.
What it would cost to settle: about two hours from someone who works with the Syntec agreement, over 60 randomly drawn answers. That is the kind of trade-off worth writing down rather than rounding away.