Back to notes
Measurement2026-09-092 min read

What goes here, and what does not

An opening note. Only things that can be reproduced — every number with its basis, and every correction stated along with whether it flattered us or not.

This is a placeholder post describing what this column is for. The first real articles will be written by the team.

Only things that can be reproduced

We say publicly that an accuracy figure without its measurement basis is not a number. That claim needs somewhere to be honoured.

A post that belongs here can answer at least:

  • Which dataset, how many questions, which categories were excluded
  • The exact model used for extraction, for answering, and as judge
  • How many runs, and how far apart the runs were
  • If a comparison appears — did we run the competitor ourselves, or are we quoting their self-reported number

Comparisons we cannot pin a basis to do not get published.

Corrections get published too

The easiest thing to suspect about a changed basis is that someone picked the arithmetic that suited them.

So a correction has to state its direction: did this make our number better or worse. Once we had counted the adversarial category in the denominator while five public harnesses all exclude it. Recomputing on the industry basis lowered our absolute score — and slightly raised our lead. Both halves get written down; publishing only the second is picking the flattering half.

And the things that broke

A system breaking is not the problem. A system breaking and saying nothing is. This column covers silent failures: no error, no alert, the run "finishes normally", and the result is wrong.

Writing them down is not a display of candour. It is that this class of problem only gets remembered once it is written — we added the guardrail for one of our own incidents only after it went into the README.

What does not go here

  • Roadmaps with no code behind them
  • Competitor comparisons without a basis
  • Customer names, unless they have explicitly agreed

Run it on the same ruler first. Everything else comes after.

If you have an eval set, we plug into it. If you do not, we use a shared third-party harness where neither the judge nor the answer prompt is ours to choose. Only a result both sides can reproduce is worth comparing.

© 2026 Reglos. All rights reserved.