A small specialist model solving a difficult proof is an interesting result. A handful of such answers is the beginning of an evaluation, not the end of one.
The recovered House of Senn draft compared QED-Nano with Claude on fifteen mathematical tasks. Its opening celebrated the specialist’s five introductory results. Its combined table told a more complicated story: the reported average was 5.3 out of 10 for QED-Nano and 6.1 for Claude. Neither score has been independently reproduced for this edition.
Separate the research from our pilot
The QED-Nano team’s publication is dated February 15, 2026, rather than the April release date given in our draft. It describes a 4B model for natural-language mathematical proofs, trained with supervised examples, dense rubric-based rewards, and a reasoning cache. Its strongest scaffolded results used substantial test-time computation. That published setup is different from a short local completion.
The original draft also called the reward binary. The researchers describe partial-credit rubrics. That distinction matters: rewarding intermediate progress is different from accepting or rejecting an entire answer.
Read the complete result
The pilot used five initial tasks and ten harder follow-ups. According to its table, QED-Nano won six tasks, Claude won seven, and two tied. Some wins were low-scoring partial answers. A score of five beating a score of three does not establish a correct proof.
The draft contains another warning: one chapter heading names a different group-theory problem from the prompt underneath it. A reproducible benchmark needs stable problem identifiers and exact prompts, not just a narrative summary.
Claims about either model’s hidden parameter count, training examples, or internal reasoning capacity cannot be inferred from these outputs. An unsuccessful long proof could reflect a timeout, truncation, prompting, or an actual reasoning error. The trace must distinguish them.
A blind judge is still a judge
The pilot hid candidate identity from a separate Claude instance. This removes an obvious cue, but it does not establish independent mathematical correctness or eliminate preferences associated with style and model family.
Keep each complete response, finish reason, token allowance, runtime version, and grading rationale. Have a qualified reviewer inspect disputed proofs. Where formal checking is practical, report exactly which proposition was checked and which assumptions remain outside the checker.
Swapping answer order and repeating grading can reveal instability. Repeating generation can reveal whether an attractive example was typical.
Keep rubric scores separate from checked proofs
The QED-Nano publication uses problem-specific partial-credit rubrics and substantial test-time scaffolding. Its rubric scale and evaluation setup differ from the archive’s informal ten-point pilot. Do not place those scores on a single chart as if they measured the same quantity.
For a fresh pilot, define the rubric before generation and preserve its version with each problem. Include complete-correctness status alongside partial credit. A method that makes useful progress but leaves a critical lemma unproved can receive partial credit without being counted as a solved proof.
Where a proof is formalized, Lean’s kernel-checking boundary provides a separate check of the formal statement. Record which assumptions and formalization choices remain to be reviewed. A natural-language solution does not become formally verified merely because it resembles a theorem that could be written in Lean.
Audit the fifteen tasks as a case series
Treat this small collection as a source of failure cases. For each task, retain the exact statement, output, termination reason, assigned score, and disputed step. Separate mathematical mistakes from truncation, tool failure, and grading disagreement.
Have a qualified reviewer assess disputed arguments without model identity, and repeat a subset with answer order reversed to investigate judging instability. Select a larger held-out collection before tuning on these cases.
Report the original six wins, seven losses, and two ties as historical observations. The new method does not change those counts or imply that this edition reran either model. Its value is a clearer path from an interesting example to an evaluation another researcher could reproduce.
Design the next comparison around a decision
The practical question might be whether a local specialist is useful as an additional candidate generator. Test that directly: compare a single baseline, the specialist alone, and a combined system under stated time and cost budgets.
A specialist that improves a narrow workflow can be valuable without replacing a general model. Publish its failed cases with its successful ones. That gives the next builder something more useful than a winner’s headline: a reasoned boundary around where the tool helps.
Keep a good idea close.
Follow Signal Tower

