How to Run a Literature Review With AI Tools (Five Stages, Each With a Stopping Point)
Priyanka's review collapsed at 214 sources in a folder. The tools had worked. The method had not — she had treated a review as collecting, and collecting has no natural end.
Priyanka’s literature review collapsed in week six. Not from a shortage of sources — she had 214 of them, in a folder, with filenames like paper (3) final FINAL.pdf. She could not remember which she had read, which she had only skimmed, or why she had downloaded about forty of them.
The tools had worked. The method had not. She had treated a review as a collecting exercise, and collecting has no natural end.

What follows is the method she rebuilt it around — five stages, each with a defined stopping point, and AI doing the parts it is genuinely good at rather than all of them.
Stage 1 — Settle the question before you search
The most expensive mistake is searching too early. A vague question returns everything, and everything is unreviewable.
Before opening a search box, write the question as a sentence you could answer yes, no or “it depends” to. Then run it through an evidence tool like Consensus and see what comes back. You are not looking for the answer yet. You are looking for one of three outcomes:
- Settled — the literature broadly agrees. Narrow the question or change it.
- Contested — the literature disagrees. This is the good outcome. Contested is where a review has something to say.
- Empty — almost nothing comes back. Either you have found a gap, or your terminology is wrong. It is usually the terminology.
Stop when: you can state the question in one sentence and you know which of those three you are in.
Stage 2 — Search wide, deliberately, once
Now cast the net. Semantic Scholar for coverage, ResearchRabbit to map outward from two or three papers you already trust, and your discipline’s own database because the general tools do not have everything.
The discipline here is to log your searches as you run them — the terms, the database, the date, the number of results. This feels like bureaucracy for about a week and then saves you, because in month three someone will ask how you found your sources and “I searched for a while” is not a method.

Stop when: new searches return papers you have already seen. That saturation point is the signal, not a target number.
Stage 3 — Screen ruthlessly, in two passes
This is where Priyanka’s review had failed. She had downloaded first and decided later, so nothing was ever excluded.
Screen in two passes with written criteria fixed before you start:
Pass one — title and abstract only. In or out. No downloading, no “maybe” pile. A maybe pile is a decision you have deferred and will have to make again.
Pass two — full text on whatever survived, against the same criteria.
Record the number excluded at each stage and why. Two hundred and fourteen sources is not a strong review; it is an unscreened one. Most good reviews end up somewhere between twenty and sixty.
Stop when: every paper is in or out, with a reason.
Stage 4 — Extract into one table
This is where AI earns its place. Decide the fields you need from every paper — population, method, sample size, outcome measure, finding, limitations, whatever your question requires — then use an extraction tool such as Elicit to pull them into a single table.

Two rules, and neither is negotiable:
- Every field must be traceable to a page in the source. If you cannot point at where a number came from, it does not go in the table.
- Spot-check by hand. Take a random ten per cent, open the papers, and verify the extracted values yourself. If the error rate is low, continue. If it is not, extract manually and accept the time cost.
Extraction tools are good. They are not good enough to skip verification, and the table is what your conclusions rest on.
Stop when: the table is complete and your sample check has passed.
Stage 5 — Synthesise, which is the part that is yours
Sort the table by finding rather than by author. Papers that agree cluster; papers that disagree stand out; gaps show up as empty columns. That grouping is the review.

A model can group things for you and it is useful for a first cut. What it cannot do is explain why two studies disagree — that one used a clinical population and the other undergraduates, that the measurement instrument changed in 2019, that the effect only appears in one geography. Those judgements are the contribution, and they come from having read the papers.
If your synthesis reads like a list of summaries, you have described the literature rather than reviewed it.
Keeping track without a second collapse
The organisational failure is as common as the analytical one. Three things prevent it:
- A reference manager from day one, not from the moment things get messy. Zotero is free and enough.
- A single sheet as the spine — one row per paper, columns for screening decision, read status, extraction status. Everything else hangs off it.
- Consistent filenames, author-year-keyword. Trivial, and it is what separates a folder from a pile.

What Priyanka’s rebuild looked like
She started again from the question, which turned out to be two questions, and dropped one. Her search log filled a page. Screening took the 214 down to 180, then to 41. Extraction took two evenings and her ten per cent check found two mistakes, both in the same field, which told her to re-check that column across the set.
The review was finished in five weeks and was shorter than the one she had abandoned. Forty-one sources, every one of them read, every one of them there for a reason she could state.
The tools did not save her. The stopping points did.
For choosing between the individual tools mentioned here, see our comparison of the best academic search engines and research tools. And before your reference list goes anywhere near a submission, check what is happening with fabricated citations — the rate has moved sharply, and it is now something reviewers look for.
Frequently asked questions
How many sources should a literature review have?
Most well-screened reviews land between twenty and sixty. A very large number usually indicates screening has not happened rather than that the review is thorough.
When should I stop searching?
When new searches keep returning papers you have already seen. Saturation is the signal, not a target count.
Can AI do the screening for me?
It can help sort and prioritise, but the include-and-exclude decision should be yours against written criteria fixed before you start. Screening is where the review’s boundaries are set.
Are AI extraction tools accurate enough?
Good enough to save substantial time, not good enough to skip checking. Verify a random ten per cent by hand against the sources, and if the error rate is high, extract manually.
What is the most common reason a literature review fails?
Searching before the question is settled, and collecting without screening. Both produce a large folder and no argument.
Which reference manager should I use?
Zotero is free and sufficient for most reviews. The important part is starting from day one rather than after the folder becomes unmanageable.
General guidance on research method. Requirements for systematic reviews vary by discipline and by journal — check your field’s reporting standards before relying on any workflow.
Sign up to our news alerts
The day's business headlines in your inbox each morning.
Unsubscribe from any email.


