Introduction

When an LLM benchmark unexpectedly returned a 100% failure rate, the researchers initially suspected model collapse. Instead, a rigid parser was rejecting valid outputs. This post walks through how a format-level bug was identified, isolated, and fixed without compromising the research integrity.

What Happened

The team built a benchmark to test whether large language models could generate project risk registers from planning documents. Across 21 real-world projects and three major models, one cell showed 42 of 42 runs failing, yielding a recall of 0.0. Breaking down the raw failures revealed the model was producing correct content, but the parser was rejecting it over technicalities: the model did not restate a project ID that the code already knew, and the JSON extraction regex demanded a single object anchored at the response end. In total, 30 failures stemmed from the missing project ID restatement, 5 from JSON array formatting, and 7 from the absence of a FINAL JSON: marker. The model's output was sound; the harness was too strict.

Why This Matters

Scoring an LLM down for format issues that do not reflect actual content quality undermines the validity of benchmark results. If researchers adjust their scorer after seeing the numbers, they risk overfitting the evaluation to quirks of their own harness rather than model ability. The fix had to preserve every raw output, change only the parsing logic, and add guardrails so future changes cannot quietly lower standards.

Key Takeaways

  • Save every raw response. Never overwrite generation output; keep it separate from parsing and scoring so you can re-score without re-calling the API.
  • Never make the model restate what your code already knows. IDs, project names, and run metadata belong in your code, not in the model's output.
  • Log finish_reason and completion-token counts. A length finish reason is a harness event, not a model answer. Count it separately.
  • Split failures into format and content. Only format failures are yours to fix. Write the rule down before touching code.
  • Break failures down by root cause, per cell. An overall failure rate hides the one cell that is on fire.
  • Put the rule in a test. If a fix must not loosen something, write the test that fails when it does.
  • Report the bug. A disclosed, corrected harness bug makes your numbers more credible, not less.

Conclusion

The researchers recovered 39 of 88 parse failures by tightening format rules and preserving raw data. They also uncovered a second, separate issue: hidden reasoning tokens on free-tier APIs eating into the output budget and truncating answers. For anyone running LLM evaluations, the takeaway is clear: treat extreme success or failure as a potential harness bug, inspect raw outputs, and fix the parser—not the model—when format is the only thing standing between data and results.