How it works
Can I Publish is an alternative to that LLM you upload your paper to, hoping what it says makes sense. It is built from real reviews and measured the way an instrument is measured: with preregistration, a held-out validation set, and a long process of inter-rater agreement.
If you are a researcher, we know you need data, not promises. This page is the whole method: where the corpus comes from, how the instrument was built, how it scores in the benchmarks and what it cannot do. It is long on purpose, it has to be. If you only want the summary, it is on the home page.
Where the data comes from
The corpus
Our data comes from two open-review sources: Routledge Open Research (38 papers · 87 reports) and Open Research Europe (80 · 205). Public reviews, citable and checkable by anyone.
We now know what a review report actually contains
They agree on the destination, not on the route
This is the finding that governs the whole design: what a reviewer criticises is decided more by the reviewer than by the manuscript; how serious the whole thing is, the manuscript decides. So the goal is not to imitate one reviewer's list — it is to cover the space of replicable criticism and get the aggregate judgement right.
How the instrument was built
The method, in three steps
- 01InductiveFour independent passes over the material, with no prior categories: which themes emerge on their own.
- 02CurationA researcher in the field prunes, merges and draws boundaries. From 22 categories to 7 families.
- 03DeductiveApplied blind to reports that were not used to build it. A separate step, on purpose.
Inter-coder reliability
A category system is only worth anything if two independent coders, reading the same thing, agree. That is measured with Krippendorff's α and published with its confidence interval — ours, by bootstrap over 3,000 resamples.
The aspect row shows the 7 families, which is the level the product works and reports at. At the fine level of 22 categories agreement stopped at 0.53 — not enough — and so it is not used. We would rather say so.
How the intervals are computed, and what coder stability tells us
The 95% intervals come from a bootstrap of 3,000 resamples over the coded units.
The same coder, rereading the same subsample in three passes, repeats himself at α = 0.952 · 0.966 · 0.971. This is the figure that separates «the scheme is ambiguous» from «the coder is noisy»: if the disagreement between two people were coder noise, this number would drop too. It does not, so the disagreement is ambiguity in the scheme — and that can be fixed.
At the level of the 7 families, agreement reaches α = 0.684. At the level of the 22 fine categories it stays around 0.53, which is why the system does not work at that level.
The check almost nobody runs
Collapsing 22 categories into 7 raises agreement all by itself. What if the whole improvement were just that — arithmetic? We tested it against chance: 3,000 random groupings of 22 into 7.
The family structure corresponds to something real in the domain, not to the effect of having fewer categories.
How it analyses your manuscript
Inside the analysis
Every version comes in with its design already detected — quantitative, qualitative or mixed — and its text extracted. The original file is deleted at that moment.
Cross-checking citations against your reference list is done by code, not a model: the candidates are computed deterministically and the model only verifies them. That is the difference between «I think this reference is missing» and «this reference is missing».
What it still does not do
It does not replace a reviewer. One in four of the points a reviewer would raise still gets past it.
Calibrated to the social sciences. Outside that field we have no evidence.
Evidence from open review. About double-blind review we can claim nothing.
It does not predict the editorial decision. And that decision is less firm than it looks:
What the benchmarks say
Against the best generic model, preregistered
The question every researcher asks: «why not just paste the PDF into a chatbot?». We answered it by measuring — not by having an opinion: generic models from two companies — Claude and ChatGPT, each with two prompts — against our system, over the same 20 papers (~800 real reviewer comments as the yardstick), with the prompts frozen before we saw a single result, half the papers held out as a validation set, and the matching judged by a model from a different family than ours.
The coverage metric is the Atomic Recall from PRAIB, reimplemented over our social sciences corpus. Consistency between the papers already seen and the held-out ones was within ±2 points: the result is not inflated by familiarity.
Anchoring: the one place we beat the reviewer instead of imitating
Human reviewers quote from memory: only 42% of their quotes from the manuscript can be found verbatim in the text. Here anchoring is not a habit — it is a rule, checked by code on every finding.
Validity framework
All the evidence is framed within the standard of psychometric measurement — Messick (1995) and the AERA/APA/NCME Standards (2014). It is not enough for an instrument to correlate: you have to document what it is made of, how it behaves and what consequences it has for whoever uses it.
What you get, and what you do not
The anatomy of a finding
«theoretical saturation was reached after the twentieth interview»
Saturation is claimed, but not the criterion used to determine it. What counted as a new code, and how many interviews without new codes were taken as the threshold? Worth spelling out in Method.
Illustrative example
The overall verdict
It is a summary of how serious what we found is, and a recommendation. It does not predict whether your paper will be accepted, and we do not present it as such.
Corpus and reliability verified on 29 July 2026; benchmark and anchoring, on 31 July 2026 under preregistration. Spotted an error, or want the full methodological detail? [email protected]