Reference · Hub

Real-world data, real-world insight, real-world evidence.

Three terms, one direction of travel, and they are not interchangeable. Data is the record that routine care left behind. Insight is the reading of that record: what it appears to show, and whether it is worth pursuing. Evidence is the conclusion that survives a specified method, a stated limitation and an independent reviewer. This page sets out the sequence, then works through every data type, what each one actually contains, and the questions each one can and cannot answer.

Start here

The sequence, in the order it happens.

Most confusion in this field comes from collapsing three stages into one word. They are sequential. Each stage costs more than the one before it, and each is defensible against a different audience.

STAGE ONE STAGE TWO STAGE THREE Real-world data The record itself. Generated because a health system ran, not because a researcher asked a question. Claims, records, dispensing, labs structure, clean, describe Real-world insight What the record appears to show, once it is described properly. Directional, fast, and enough to make a decision. Sizing, pathways, trends, gaps specify, control, test, disclose Real-world evidence A conclusion about use, benefit or risk, produced by a pre-specified method and stated limitations. Submission, publication, defence Evidence changes the question, and the next read starts again from the record
01

Data is a source, not a finding

Routinely collected information about health status or the delivery of care: claims, dispensing, hospital activity, clinical records, laboratory values, registries, devices. Nobody designed it to answer your question. That is simultaneously its greatest strength, because it is unselected and at scale, and its central difficulty, because the thing you want measured may never have been recorded.

02

Insight is what a competent description yields

Once the data is structured, de-duplicated and correctly denominated, it will tell you how large a population is, what is being used and in what order, where the gaps sit and which direction the trend runs. This is descriptive work, it is fast, and for most commercial and medical planning decisions it is sufficient. It is not proof of anything, and should not be presented as though it were.

03

Evidence is what survives a method

A causal or comparative claim about use, benefit or risk, generated from specified data by a method fixed before the data was seen, with confounding addressed, sensitivity analyses run and limitations declared. Change the method and you change the result, which is why the method note, not the dataset, is the part a regulator, a payer or a journal actually assesses.

The distinction is not pedantry. A stakeholder who asks for "real-world data" almost always wants insight or evidence. A supplier who delivers a dataset has technically complied and practically failed. Equally, a descriptive read presented as evidence is the most common way a well-run analysis loses its credibility in review.

Crossing the stages

Five steps separate a record from a defensible conclusion.

Each step is where a study is usually lost. None of them is the analysis, which is the fastest part of almost every programme.

1

Specify the question

Population, exposure, comparator, outcome and time window, written down before anyone looks at the data. A question that cannot be written in that form cannot be answered from routine data.

2

Test feasibility

Does the cohort exist at usable size, is the outcome actually captured in this data type, and is the follow-up long enough. This is where a question gets refused, and refusing it early is the cheapest outcome available.

3

Define the cohort

Codes, algorithms and washout periods turned into a reproducible definition. Two teams using different definitions on the same data will report different numbers, and both can be right.

4

Address confounding

Nobody was randomised. A difference between groups may reflect the treatment, or it may reflect why each patient received it. Matching, adjustment and negative controls are how that is tested rather than assumed.

5

State the limits

What the data could not see, which results moved under sensitivity analysis, and where the conclusion stops. An analysis with no declared limitation has not been examined closely enough.

Where it usually goes wrong. Not in the modelling. In an outcome that was never recorded, a denominator that was a projection rather than a count, a cohort defined after the first result was seen, or a comparison drawn across two markets whose records cannot be joined.

The raw material

What each data type contains, and what it can answer.

"Real-world data" is not one thing. It is at least twelve, and they differ more from each other than most buyers expect. The single most common error in this field is asking a data type for something it never recorded. Read the last column first.

Data type What the record actually contains Questions it answers well Where it stops
Claims, government A payment record created whenever a subsidised medicine is dispensed or a subsidised service delivered. Item, date, quantity, benefit paid, prescriber type, patient category. Market size and share on a counted denominator, treatment sequencing, persistence from dispensing gaps, cost and utilisation, the signature of a policy change. It records what a scheme paid for, not why. Usually no diagnosis on a prescription, no clinical result, and nothing about privately funded care.
Claims, insurance The same shape of record with a private or employer-based insurer as payer. Billing codes, dates, amounts, and frequently attached laboratory values. Utilisation and cost in the insured population, treatment patterns outside public funding, budget impact, comparative effectiveness where labs are attached. The covered population is not the whole population. Employer schemes skew to working age, so elderly and unemployed patients are under-represented.
EMR and EHR What clinicians wrote down: diagnoses, orders, results, notes, often free text. The richest clinical detail available in routine care. Standard of care as practised rather than as guidelines describe it, clinical outcomes including response and progression, cohorts defined clinically, reason for discontinuation. Coverage is institutional, not national. A patient who moves between hospitals usually restarts. Free text is often excluded or needs separate permission.
Hospital activity Counts of what hospitals did: admissions, procedures, separations, length of stay, day cases, emergency presentations. Usually facility or jurisdiction level. Where a procedure is performed and at what volume, facility-level targeting, peer benchmarking, capacity as a proxy for site feasibility, procedure trends. Tells you what was done, not what it cost the hospital, who decided it, or what happened to the patient afterwards. Rarely patient level.
Dispensing The record created when a medicine is physically handed over at a pharmacy. Closer to real use than a prescription, which may never be filled. Actual uptake rather than intent to treat, refill and persistence patterns, within-class switching, seasonality and stockpiling, week-by-week launch tracking. Shows the medicine leaving the pharmacy, not the patient taking it. Under-co-payment dispensing is incompletely captured in some schemes.
Laboratory results Numeric values from diagnostic tests, with dates, so trajectories rather than single points. Response and progression rather than treatment alone, biomarker-defined cohorts and eligible volume, testing rates and testing gaps, severity stratification, early safety signals. Reference ranges and units differ between laboratories, so harmonisation is real work. Where only a selected subgroup is tested, the tested population is not the treated population.
Registry A collection built on purpose for one disease, procedure or product. Captures fields nobody records elsewhere: stage, grade, histology, biomarker status. Disease epidemiology at a defined completeness, staging and histology, outcome and survival research with a validated denominator, quality benchmarking, cohort enrichment. Contains only what it was built to capture, for only the population it enrols. Reporting lags are usually longer than claims because cases are curated.
Biobank Biological samples linked to health information about the people who gave them: questionnaire, physical measurement, genotype, and often follow-up. Genetic risk and polygenic scores, biomarker discovery and validation against real outcomes, target identification, exposure analysis alongside clinical data. Participants volunteer, so the cohort is not a random sample. Sample sizes limit rare-event work, and follow-up intervals are set by the study, not by care.
Genomics and precision medicine Sequencing and molecular profiling results, ideally joined to the clinical record of what was then done and what followed. Population-specific variant frequency, biomarker-eligible population sizing, testing pathway analysis, matching of molecular subtype to treatment and outcome. A sequence with no linked outcome answers a biology question, not a clinical one. Linkage, not sequencing, is almost always the binding constraint.
Linked data Two or more sources joined at the person level under an approved linkage, most often claims with hospital, or a registry with dispensing. Full pathway questions that no single source can answer: diagnosis through treatment to outcome, cost across settings, and true incident-case identification. Requires a custodian-approved linkage and usually a formal approval pathway, which sets the timeline. Records cannot be joined across national systems.
Population data Census, demographic, socioeconomic and geographic reference data published by statistical agencies. Denominators, per-capita normalisation, catchment and territory design, equity and access analysis, standardisation for age and sex. No clinical content at all. It is the denominator layer that makes other data types comparable, and useless on its own for a clinical question.
Epidemiological studies Published prevalence, incidence, survival and burden estimates, derived from surveys, cohorts and surveillance systems. Baseline burden and unmet need, sanity checks against a claims-derived estimate, market sizing where record-level access does not exist. Published aggregates cannot be re-cut. If your question needs a subgroup the authors did not report, the estimate cannot answer it.

Two data types beat one. Claims describe what was funded, laboratory values describe what happened, and a registry describes who the patient was. Any two of those answers materially more than either alone, which is why the linkage question, rather than the volume question, decides what a programme can realistically deliver.

Use cases

Which question goes to which data type.

The same dataset serves different functions differently. What follows is the mapping we use when scoping: the decision on the left, the data types that can carry it, and whether the honest output is an insight or evidence.

Decision The question, stated properly Data types that can carry it Honest output
Unmet need How many patients reach a given line of therapy each year, and what happens to those who fail it. Claims, dispensing, registry, linked data Insight, becoming evidence once the outcome is captured rather than inferred
Market sizing How many treated patients exist in this indication, on a counted rather than modelled denominator. Claims, dispensing, population data Insight
Treatment pathway In what order are therapies actually used, for how long, and where do patients switch or stop. Claims, dispensing, EMR, linked data Insight, robust when built on dispensing gaps rather than prescriptions
Comparative effectiveness Do patients on A experience different outcomes from comparable patients on B. Linked data, EMR with laboratory results, registry Evidence, and only with a pre-specified confounding strategy
Safety in practice What is the rate of an event among treated patients, against a defensible denominator. Claims, dispensing, EMR, laboratory results Evidence for rates, insight for signals
Trial feasibility How many eligible patients exist under these criteria, and at which sites. Claims, EMR, registry, hospital activity Insight, and it is the cheapest study you will ever run
External comparator What would have happened to these patients under standard care. Registry, linked data, EMR with outcomes Evidence, subject to the outcome being measured the same way
Health economics What does this pathway cost in this system, and what changes if the mix shifts. Claims, hospital activity, linked data, population data Evidence for cost, insight for the counterfactual
Reimbursement submission What does the local population look like, and what do local inputs do to the global model. Claims, registry, epidemiological studies, population data Evidence, and the local inputs are usually the part that is challenged
Biomarker strategy Who is tested, who is not, and how many patients are eligible but never identified. Laboratory results, genomics, EMR, registry Insight, becoming evidence where testing and treatment are linked
Field and territory design Where is the procedure performed, at what volume, and by which facilities. Hospital activity, claims, population data Insight
Launch performance Is uptake tracking to plan, in which segments, and against what baseline. Dispensing, claims Insight, week by week rather than quarter by quarter
The honest comparison

Real-world data against controlled trial data.

Neither is superior. They answer different questions, and the most common error in the field is asking one of them to do the other's job. A trial establishes whether a treatment can work under controlled conditions. Real-world data establishes what happens when it meets an actual population.

Real-world dataControlled clinical trial
What it measuresEffectiveness. What happened in practiceEfficacy. What happened under protocol
SettingRoutine clinical practiceControlled research conditions
Who is includedNo strict inclusion criteria. The population as it isStrict criteria, applied before enrolment
Who drives the dataPatient-centred. Care happened, a record followedInvestigator-centred. Data collected on purpose
Comorbidities and interactionsPresent, because real patients have themIncluded only where the protocol allows
Role of the physicianMultiple physicians, chosen by the patientA designated investigator
ComparatorWhatever the market and the physician actually chosePlacebo or a defined standard of care
TreatmentVariable, as determined by practiceFixed, according to the protocol
Response monitoringVariable, whenever care happened to occurContinuous, on a schedule
Follow-upDetermined by clinical practiceDefined by the protocol
Assignment to treatmentChosen by clinician and patient, so confounded by indicationRandomised, which is what makes the comparison causal
Strongest useGeneralisability, long-term outcomes, cost, and populations a trial excludedEstablishing that an effect exists at all, under control

The practical consequence. Because nobody randomised anyone, a difference between two groups in real-world data may reflect the treatment, or it may reflect why each patient received that treatment in the first place. Handling that is the entire craft: cohort definition, covariate control, sensitivity analysis, and the honesty to say when the design cannot settle the question.

Who reads it, and what they test it against

Six audiences, six different standards of proof.

The same analysis is judged differently depending on who receives it. Knowing which audience the output is for determines how far up the sequence the work has to go.

Patients

Better access to information has raised awareness of both trial and real-world findings, and with it the reporting of safety concerns, comorbidity effects and long-term outcomes. It has also raised willingness to take part in real-world studies at all, which is what keeps registries and biobanks viable.

Clinicians

A trial population rarely resembles a clinic list. Real-world evidence describes what happens to older patients, patients on four other medicines, and patients who would have been excluded from the registration study. The standard applied is whether the cohort looks like their patients.

Regulators

Post-market safety and effectiveness at population scale, and a growing role in label extensions and single-arm settings where a randomised comparator is not feasible. The standard applied is whether the method was fixed before the data was seen.

Payers and HTA bodies

They ask what a treatment does in their own population, at their own prices, in their own care pathway. That is a real-world question by definition. The standard applied is whether the local inputs are local, and they will check.

Industry

Unmet need, comparator selection, feasibility, launch tracking, safety context and reimbursement evidence. Most commercial questions inside a life sciences company are real-world questions wearing a different name. The standard applied is whether a decision can be made on it.

Health systems

Service planning, quality improvement and resource allocation, using the record the system already generates rather than commissioning new collection. The standard applied is whether the finding holds at facility level, where the budget sits.

Why standardisation has become the live issue. Registered real-world studies have grown steadily, and with them the importance of standardising how such work is conducted and reported. An unregistered, unspecified real-world analysis is difficult to distinguish from an exercise in finding the answer somebody wanted. Pre-specification and a published method note matter more here than in almost any other setting.

So what

The data almost always exists. The question is whether you can get an answer out of it.

  • Name the stage you actually need. Most decisions need insight, not evidence, and paying for evidence when insight would do is the most common overspend in this field. The reverse, presenting insight as evidence, is the most common credibility failure.
  • The method note is the product. Two competent teams analysing the same dataset can reach different conclusions. What makes one defensible is that the method was specified before the data was seen.
  • Match the question to the data type. Asking claims for a clinical outcome, or a registry for a market size, wastes a quarter before anyone notices the record never held it.
  • Governance decides your timeline. In this region approval pathways rather than analysis time set the schedule, and they differ market by market.
  • Someone has to be willing to say no. The most valuable output of a feasibility read is often the finding that the question, as posed, cannot be answered from routine data.

Have a question that might be a real-world question?

Describe the decision you are trying to make. We will tell you whether routine data can answer it, from which data type, in which market, and how long it would take.

De-identified patient-level data and publicly available data, through local alliance partners. De-identification at source. Analysis in-market, under local law. Never identifiable records.