Extract DataUse Cases

What you can do with Extract Data

Screen large corpora, build standard evidence tables, prepare meta-analysis inputs, run gap analyses, and assemble regulatory evidence packs — all from papers in your Paperguide library.

Extract Data reads every paper in your set and fills one table: a row per paper, a column per question, every value clickable back to the sentence it came from. This page covers the twelve field families that work in any discipline, fifteen use cases with worked examples, and four starter packs you can lift wholesale.

For ready-to-copy field sets organised by research discipline, see Extract Data examples by discipline.

For how to word a field, see Tips for creating better extraction fields.

The twelve field families

Nearly every extraction table is built from the same twelve families. Only the vocabulary changes between disciplines.

FamilyClinicalCS / MLQualitativeMaterials / engineeringSocial science / economics
IdentityTrial name, registration IDModel or method nameStudy settingMaterial or systemDataset or survey
DesignRCT, cohort, case-controlExperimental setup, ablationEthnography, interviewsTest protocolPanel, cross-section, RCT
SampleParticipants analysedDataset size, train/test splitParticipants to saturationSpecimens testedRespondents, observations
Population / contextAge, sex, comorbidityDomain, data sourceCommunity, roleComposition, conditionRegion, demographic
Intervention / methodDrug, dose, procedureArchitecture, training regimeProcess conditionPolicy or programme
ComparatorPlacebo, standard careBaselines, ablationsControl conditionControl group, counterfactual
Outcome / metricEndpoint, event rateAccuracy, F1, AUCThemesMeasured propertyCoefficient, elasticity
ResultEffect size, CI, pBenchmark scoreRepresentative findingsValue with toleranceEstimate with SE
Duration / timingFollow-up monthsEpochs, runtimeFieldwork periodAgeing, cyclingWave, period
Quality / biasRisk-of-bias domainsReproducibility, seedsTrustworthiness criteriaStandard complianceAttrition, weighting
ProvenanceFunding, registrationCode and data availabilityEthics approvalStandard referencedData source, funding
LimitationsAuthor-stated limitationsStated limitationsReflexivity notesTest limitationsStated limitations

Pick six to eight families that answer your question. A narrow question deserves few sharp columns, not padding.

What you can do with Extract Data

Fifteen jobs researchers hire extraction for. Each one is a table you can build today.

A. Before you read anything

1. Screen a large set without reading in full

You have 300 hits and no idea which 40 matter. Add one or two Yes/Maybe/No columns, run them across everything, sort, and keep what survives.

Example: Randomised? · Adults only? · Reports the outcome I care about? Three columns, 300 rows, then filter. You end up with: a defensible shortlist and a record of why each paper was kept or dropped — every verdict backed by a sentence you can click.

Yes/Maybe/No is the type built for this. One rule decides whether a screening column works: define all three verdicts in the instructions, including what the information's absence means. The extractor follows your definitions; when they're silent about absence, a paper that says nothing comes back maybe — never no.

Which way should absence go? Ask: would a paper meeting this criterion necessarily say so? An RCT always announces its randomisation, so "no mention of random allocation anywhere in the Methods" is a legitimate no — write it into the instructions. But ethics approval can be obtained and never stated, so there absence stays maybe — a no would assert misconduct from a formatting omission.

A worked example — separating human studies from everything else:

FieldTypeInstructions
Human participants?Yes/Maybe/NoDoes this study collect or analyse data from living human participants — patients, volunteers, survey respondents, or their clinical samples? yes if the Methods describe recruiting, enrolling, consenting, or analysing data or samples from humans, including secondary analysis of an existing human dataset. no if the Methods describe only a non-human study population — animal model, cell line, in vitro, in silico, simulation — or if the paper reports no empirical study at all (review, editorial, protocol). maybe if both human and non-human work is reported, or if the population is not described clearly enough to tell.

Note the shape: every verdict has its own condition, and between the no clauses and the maybe clauses there is nowhere undefined left for a paper to fall. A one-line instruction like "no if it does not use human participants" leaves everything else — including silence — to the default, and the default is maybe.

The strongest case for defining absence as no is a criterion like adverse-event reporting: a paper with no safety data anywhere contains no sentence that could ever "rule it out", so without an absence clause every failing paper returns maybe and the column separates nothing. "no if no adverse-event data appears anywhere in the paper, including tables and supplementary material" is what makes that screen work.

Want to see the corpus rather than just filter it? Use Specified instead: Human participants, Human tissue/cells only, Animal, In vitro/cell line, In silico/simulation, No empirical study. Same single pass, but you learn the composition of your set instead of a keep/drop flag — and you can still filter on it.

2. Triage a mixed corpus by type

A Specified column (RCT, Cohort, Review, Protocol, Case report) separates primary studies from everything else in one pass, before you invest in a full evidence table.

B. Building the table your review needs

3. The standard evidence table

The table every literature review, SLR, thesis chapter, and grant background section needs: design, population, intervention, comparator, outcome, findings, limitations.

Example topic: interventions for post-operative delirium in older adults. You end up with: a Word- or Excel-ready characteristics-of-included-studies table, exported in one click.

4. Compare methods, not just findings

Understanding how studies did the thing is what explains why they disagree. Extract design, sampling, measurement instrument, and analysis approach side by side.

Example topic: how physical activity was measured across cohort studies — self-report vs accelerometer. You end up with: the heterogeneity paragraph of your discussion, already sourced.

5. Map who was studied

Age range, sex distribution, country, setting, comorbidities, inclusion boundaries — one row per study — shows instantly where the evidence is thin.

Example topic: which populations appear in trials of digital mental health apps. You end up with: a generalisability or equity section, and a clear statement of who the evidence does not cover.

6. Pull intervention and protocol detail

Dose, frequency, duration, delivery mode, provider, adherence, co-interventions — everything you need to replicate or compare what was actually delivered.

Example topic: dose and duration of mindfulness programmes across RCTs. You end up with: a protocol design brief grounded in what has already been tried.

7. Inventory measurement instruments

Before you pool anything, know which scale, assay, or metric each study used.

Example topic: which depression scales appear in adolescent intervention trials. You end up with: a defensible decision about what can be combined and what can't.

C. Getting to numbers

8. Prepare meta-analysis inputs

Sample size per arm, event counts, means and SDs, effect estimates with CIs, follow-up duration — the numeric sheet your statistical software wants.

Example topic: effect of SGLT2 inhibitors on heart failure hospitalisation. You end up with: a CSV that goes straight into RevMan, R, or Stata — with every number traceable to its sentence when a reviewer questions it.

9. Benchmark models, methods, or materials

The comparison table at the heart of every technical survey: dataset, baseline, metric, result, conditions.

Example topic: accuracy and F1 across transformer variants on a named benchmark. You end up with: the survey table, plus an honest view of which papers report compute and which don't.

10. Extract economic parameters

Cost items, currency and price year, perspective, time horizon, utilities, ICER, discount rate.

Example topic: cost-effectiveness inputs across published models in a therapy area. You end up with: a parameter table for your own model, or a payer dossier appendix.

11. Track change over time

Add publication year alongside the variable you care about, sort, and read the trend.

Example topic: how reported model accuracy on a benchmark moved year over year. You end up with: the trend figure in your survey's introduction.

D. Judging what you can trust

12. Appraise quality and risk of bias

Randomisation, allocation concealment, blinding, attrition, selective reporting, funding — one Specified column per domain (Low, Some concerns, High) gives you a grid you can colour.

Example topic: risk of bias across trials informing a clinical guideline. You end up with: a risk-of-bias summary where every judgement links to the sentence that justifies it.

13. Check reporting, transparency, and provenance

Registration ID, protocol availability, funding source, conflicts of interest, data and code availability, preprint status.

Example topic: how many trials in a field are prospectively registered. You end up with: a meta-research result, or a fast answer to the reviewer who asks who funded what.

E. Finding what isn't there

14. Gap analysis — read the empty cells

This is the use case people don't think of, and the one a spreadsheet can't do. Extract what each paper covers, then look at the column that came back mostly "Not reported".

Example topic: which outcomes are never reported in studies of a given intervention. You end up with: the "gap in the literature" paragraph for a grant proposal or a thesis contribution statement — with a table proving it, not an assertion.

Because absent data is marked honestly (Not reported, Not applicable, Unclear) rather than guessed, the empty cells are trustworthy evidence.

F. Deliverables for industry and regulators

15. Assemble regulatory and market-access evidence

Device or product studied, indication, sample, endpoints, adverse events, favourable and unfavourable outcomes, funding — the appraisal a CER, PER, HTA submission, or value dossier requires, with traceability built in.

Example topic: clinical performance of a diagnostic assay across the published literature. You end up with: an appraisal table where every claim has a sentence-level source, which is exactly what an auditor or notified body asks for.

Four starter packs

Lift one wholesale, then tune.

Screening (3 columns, any discipline)

FieldTypeInstructions
Primary study?Yes/Maybe/NoDoes this paper report original data collected or generated by the authors? yes if Methods describe the authors' own data collection or generation. no for reviews, editorials, commentaries, and protocols. maybe if the design is unclear.
Population matchYes/Maybe/NoDoes the study population match [your population]? yes only if the paper explicitly describes that population. no if a different population is described. maybe if the population is not described clearly enough to tell.
Reports target outcomeYes/Maybe/NoDoes the paper report [your outcome] as a measured result? yes only if a result is reported, not merely discussed. no if the outcome appears nowhere in the results, including tables. maybe if it is discussed without a measured result.

Note that each instruction defines all three verdicts, including what absence means — a paper reporting your outcome would necessarily say so, which is why "appears nowhere" is a legitimate no there. See the worked example in Part 2 §A before adapting them.

Evidence table (8 columns)

Design · Sample · Population · Intervention · Comparator · Primary outcome · Primary result · Author limitations. Field wording as in the clinical medicine set on the examples by discipline page.

Meta-analysis inputs (8 columns)

FieldTypeInstructions
Spread statisticSpecified — Mean (SD), Mean (SE), Mean (95% CI), Median (IQR), Median (range)Classify which summary statistic the paper reports for the primary continuous outcome. Read this column first: it tells you how many studies need converting before you pool.
N interventionAnswerNumber analysed in the intervention arm. Number only. Format: "148".
N controlAnswerNumber analysed in the control arm. Number only. Format: "146".
Events interventionAnswerNumber of primary-outcome events in the intervention arm with denominator. Format: "28/148".
Events controlAnswerNumber of primary-outcome events in the control arm with denominator. Format: "42/146".
Mean (SD) interventionAnswerPost-treatment mean and SD for the primary continuous outcome in the intervention arm, exactly as printed. If only median (IQR) is given, answer "Not reported: median (IQR) only".
Mean (SD) controlAnswerSame for the control arm.
TimepointAnswerTimepoint at which the extracted values were measured. Format: "12 weeks".

Extract Data copies what the paper prints and never computes — so median-to-mean conversions (Wan et al. 2014, Luo et al. 2018) belong in your analysis script, not in a cell. Extract the medians, quartiles, and per-arm n, export, and convert there, where the formula is visible to a reviewer.

Transparency audit (6 columns)

FieldTypeInstructions
Registration IDAnswerTrial or review registration identifier exactly as printed, with the registry name. If absent, answer "Not reported".
Prospectively registeredYes/Maybe/NoWas registration before enrolment or before analysis began? yes only if dates make this explicit. no if the dates show registration after enrolment began or the paper says it was retrospective. maybe if timing cannot be determined.
Protocol availableYes/Maybe/NoDoes the paper state that a full protocol is publicly available? yes if a protocol link, citation, or availability statement is given. no if none appears anywhere. maybe if a protocol is mentioned with no access route.
FundingSpecified — Industry, Public/government, Charity/foundation, Mixed, None statedClassify the funding source from the funding statement.
Conflicts declaredAnswerConflicts of interest as declared, summarised in one line. If the paper states none, answer "None declared".
Data or code availableYes/Maybe/NoDoes the paper state that underlying data or code is available? yes only if a statement or link is given. no if no availability statement appears anywhere. maybe if available on request.

Where the papers come from

Any set of references in your Paperguide library works: papers you uploaded as PDFs, papers imported from a search, or public papers added to your Reference Manager. Up to 100 papers and 20 columns per table — filter by source, sort, select the rows you want, and export the whole table or just the selection to CSV or Excel. Share a read-only link when someone needs to check your work.