Data Extraction
Pull your defined fields from every included paper into a table, verify each value against its citation, and confirm the stage to generate your report.
What is Data Extraction?
The final working stage. Paperguide AI reads every included paper and pulls out the fields you defined in your protocol, assembling them into a table with a citation behind every value. A person verifies that table before it becomes the basis of your report.
Before you start: Your extraction fields must be defined in the Protocol, and Full Text Screening must be confirmed. Assign your verifier before starting, not after.
How this stage works
- A Full Access member starts extraction. AI reads each included paper and fills every extraction field.
- Results build into a table. One row per paper, one column per field, with a citation behind each value.
- A verifier checks the values. They open each citation, compare the value against the supporting statement, and correct anything that is wrong.
- The verifier confirms the papers. Only then can your report be generated.
This is the stage where your review stops being a set of decisions and becomes data you can analyse.
What gets extracted
Every paper marked Included in full text screening moves here. Excluded and ineligible papers do not.
If you disabled full text screening when you created the review, papers arrive here straight from abstract screening. Any without a PDF are ineligible for extraction, since there is no full text to read.
Before you start
Define your extraction fields
Extraction fields tell AI what to pull from each paper and how to present it. Each field has a title, instructions, and an output format. The output setting controls the format, so your instructions do not need to repeat it.
| Output | Returns |
|---|---|
| Answer | Free text, written as your instructions specify |
| Yes / Maybe / No | One of three fixed options, for judgments you want recorded consistently |
| Specified answers | One option from a list you define |
Free text is flexible but harder to compare across papers. Fixed options are comparable, but only if the list is complete and the options do not overlap.
Writing instructions AI can act on
Extraction quality follows directly from how specific your instructions are. Write them so a new reviewer could apply the field without asking you what you meant.
A clear instruction does three things:
- States what to extract, in your review's terminology, anchored to your population or setting so off-topic content is ignored.
- Lists the details to capture, including sub-components, units, and expected categories.
- Says how to present the answer, so values are comparable across every paper.
Building better extraction fields
Assign a verifier
Assign the person responsible for verifying this stage before extraction starts. Assignment determines who can act on the results and who gets notified when the run finishes.
| If the stage is | Who verifies it |
|---|---|
| Assigned to a member | That member. They receive the notification when extraction finishes. |
| Not assigned to anyone | Any Full Access or Review member. All of them are notified and can work through the results together. |
View members can never verify, assigned or not.
Running extraction
Only Full Access members can start extraction. Review members verify results but cannot trigger the run.
Before you begin, check that the extraction fields defined in your protocol are the ones you want. Changing fields after a run means extracting again.
Select Start extraction. For each paper, AI works through your fields one at a time, following each field's instructions to locate the relevant material, interpret it, and produce the output in the structure you specified. Results build into the table as papers are processed.
Paperguide sends an email when the run completes: to the assigned verifier, or to every Full Access and Review member if the stage is unassigned.
PDF quality affects extraction accuracy. Image-only scans and PDFs without usable text layers give AI less to work with than a clean digital PDF.
Working with your extracted data
There are two ways to look at the results.
Grid view
The default. A table showing every included paper at once: columns for the paper's details and for each of your extraction fields, one row per paper.
Grid view is for comparing across papers. It is where you notice that a field is thin across the whole set, or that two studies report an outcome differently. Use search, filtering, and sorting to work through a large table.
Detailed view
Open a paper to see it in full: the PDF and its metadata on one side, your extraction fields on the other.
Detailed view is for verifying a single paper properly. You have the source and the extracted values side by side, so checking a value means reading the paper rather than trusting the table.
Verifying and editing
Extracted values are proposals. A confirmed value is one a person has checked.
Every value carries a citation link. Select it and Paperguide takes you into the full text PDF and shows the supporting statements the value was drawn from. This is what makes the table verifiable: any value in it traces back to a specific sentence in a specific paper in a few seconds, by you now or by a reviewer later.
To verify a paper, work through its fields one at a time. For each value, open the citation and read the supporting statement in context. Ask whether the value answers what your field instructions asked for, not just whether it appears in the paper.
To correct a value, edit it directly. Your edit is saved over the AI value, and both are kept, so the review holds a record of what AI extracted and what a person changed.
To confirm, select Confirm once you have verified a paper's fields. Work through papers one at a time, or select Confirm All to confirm the full table at once when you have reviewed it.
Full Access members can review and override extracted values after verification is complete, including values confirmed by another member.
Watch for a field that comes back thin across many papers. That usually means the field instructions were unclear rather than the literature being silent, and it is worth fixing the field rather than filling values by hand.
Confirming the stage
When every paper has been verified, the extraction stage completes.
Confirming makes the report available. Paperguide synthesizes your extracted data into the systematic review report, including the PRISMA diagram.
The report draws only on confirmed extraction data. Anything you have not verified does not reach it.
Exporting your data
Export your extraction data to CSV or Excel.
This is the output most likely to leave Paperguide. Extracted data is what feeds a meta-analysis, a statistical package, or an evidence table in your manuscript, and the export carries the same values you verified.
If your protocol or papers change
If you change your extraction fields in the protocol, extraction may need to run again so your table matches the fields you are now using. How much needs re-extracting depends on what changed. This is the strongest reason to settle your fields before starting: an extraction table you have verified by hand is expensive to rebuild.
If you add papers to the collection, they move through screening first. Any that are included arrive here, and only those papers are extracted.