Skip to content

How it works

Can I Publish is an alternative to that LLM you upload your paper to, hoping what it says makes sense. It is built from real reviews and measured the way an instrument is measured: with preregistration, a held-out validation set, and a long process of inter-rater agreement.

75%of what human reviewers raise, detected — the best generic model with a good prompt: 67%
94%of its findings are echoed by what real reviewers said about the same paper
90%of quotes from the manuscript, verified by code — human reviewers: 42%

If you are a researcher, we know you need data, not promises. This page is the whole method: where the corpus comes from, how the instrument was built, how it scores in the benchmarks and what it cannot do. It is long on purpose, it has to be. If you only want the summary, it is on the home page.

Where the data comes from

The corpus

118full papers
292real review reports
100reports coded
1,423reviewer comments, coded one by one

Our data comes from two open-review sources: Routledge Open Research (38 papers · 87 reports) and Open Research Europe (80 · 205). Public reviews, citable and checkable by anyone.

We now know what a review report actually contains

criticismneutral commentspraise
90%of the problems are fixed by writing
3-4%cannot be fixed without redoing the study
of the report asks for nothing: summary, scaffolding and overall assessment

They agree on the destination, not on the route

65%of the time two reviewers of the same paper agree on the VERDICT (chance: 38%)
≈ chanceis how much they agree on WHICH aspects to criticise (overlap 0.26 against 0.24 by chance)

This is the finding that governs the whole design: what a reviewer criticises is decided more by the reviewer than by the manuscript; how serious the whole thing is, the manuscript decides. So the goal is not to imitate one reviewer's list — it is to cover the space of replicable criticism and get the aggregate judgement right.

How the instrument was built

The method, in three steps

  1. 01
    InductiveFour independent passes over the material, with no prior categories: which themes emerge on their own.
  2. 02
    CurationA researcher in the field prunes, merges and draws boundaries. From 22 categories to 7 families.
  3. 03
    DeductiveApplied blind to reports that were not used to build it. A separate step, on purpose.
Building the category system and applying it were separate steps, so that the scheme could not justify itself.
22categories
7families
3%falls into «other» — the scheme leaves no gaps

Inter-coder reliability

A category system is only worth anything if two independent coders, reading the same thing, agree. That is measured with Krippendorff's α and published with its confidence interval — ours, by bootstrap over 3,000 resamples.

0.667 — conventional thresholdPolarity0.8100.747Fixability0.5570.696Aspect (7 families)0.6840.000.250.500.751.00
Round 1 (n=100)Round 2 (n=75), with 95% confidence interval

The aspect row shows the 7 families, which is the level the product works and reports at. At the fine level of 22 categories agreement stopped at 0.53 — not enough — and so it is not used. We would rather say so.

How the intervals are computed, and what coder stability tells us

The 95% intervals come from a bootstrap of 3,000 resamples over the coded units.

The same coder, rereading the same subsample in three passes, repeats himself at α = 0.952 · 0.966 · 0.971. This is the figure that separates «the scheme is ambiguous» from «the coder is noisy»: if the disagreement between two people were coder noise, this number would drop too. It does not, so the disagreement is ambiguity in the scheme — and that can be fixed.

At the level of the 7 families, agreement reaches α = 0.684. At the level of the 22 fine categories it stays around 0.53, which is why the system does not work at that level.

The check almost nobody runs

Collapsing 22 categories into 7 raises agreement all by itself. What if the whole improvement were just that — arithmetic? We tested it against chance: 3,000 random groupings of 22 into 7.

none of the 3,000 random groupings got past this point0.530median at random0.684our 7 families

The family structure corresponds to something real in the domain, not to the effect of having fewer categories.

How it analyses your manuscript

Inside the analysis

01 / 06Manuscript

Every version comes in with its design already detected — quantitative, qualitative or mixed — and its text extracted. The original file is deleted at that moment.

Cross-checking citations against your reference list is done by code, not a model: the candidates are computed deterministically and the model only verifies them. That is the difference between «I think this reference is missing» and «this reference is missing».

What it still does not do

It does not replace a reviewer. One in four of the points a reviewer would raise still gets past it.

Calibrated to the social sciences. Outside that field we have no evidence.

Evidence from open review. About double-blind review we can claim nothing.

It does not predict the editorial decision. And that decision is less firm than it looks:

64 out of 100. That is how often two human reviewers agree with each other on the same manuscript.Cicchetti (1991). This is why no tool can predict the editorial decision: the criterion it would be trying to predict is not stable either.

What the benchmarks say

Against the best generic model, preregistered

The question every researcher asks: «why not just paste the PDF into a chatbot?». We answered it by measuring — not by having an opinion: generic models from two companies — Claude and ChatGPT, each with two prompts — against our system, over the same 20 papers (~800 real reviewer comments as the yardstick), with the prompts frozen before we saw a single result, half the papers held out as a validation set, and the matching judged by a model from a different family than ours.

Can I Publish75%ChatGPT67%Claude58%% of what reviewers raise, detectedCan I Publish94%ChatGPT92%Claude90%% of findings echoed by real reviewers
Can I Publishthe best generic configuration from each company, with a good prompt

The coverage metric is the Atomic Recall from PRAIB, reimplemented over our social sciences corpus. Consistency between the papers already seen and the held-out ones was within ±2 points: the result is not inflated by familiarity.

Anchoring: the one place we beat the reviewer instead of imitating

Can I Publish90%Human reviewers42%
Of the quotes from the manuscript each one makes, how many can be found verbatim in the text. The system's are checked by code, one by one; any that cannot be located is shown with a warning.

Human reviewers quote from memory: only 42% of their quotes from the manuscript can be found verbatim in the text. Here anchoring is not a habit — it is a rule, checked by code on every finding.

Validity framework

All the evidence is framed within the standard of psychometric measurement — Messick (1995) and the AERA/APA/NCME Standards (2014). It is not enough for an instrument to correlate: you have to document what it is made of, how it behaves and what consequences it has for whoever uses it.

What you get, and what you do not

The anatomy of a finding

Methodological rigour·analytic transparency·important

«theoretical saturation was reached after the twentieth interview»

Saturation is claimed, but not the criterion used to determine it. What counted as a new code, and how many interviews without new codes were taken as the threshold? Worth spelling out in Method.

Illustrative example

Axiswhich dimension it speaks to
Subdimensionthe specific problem
Polarityproblem or strength
Severityblocking · important · minor
Anchorthe literal quote that justifies it
Commentwhat is happening and what it would take

The overall verdict

It is a summary of how serious what we found is, and a recommendation. It does not predict whether your paper will be accepted, and we do not present it as such.

Corpus and reliability verified on 29 July 2026; benchmark and anchoring, on 31 July 2026 under preregistration. Spotted an error, or want the full methodological detail? [email protected]