Reference · Hub

The twelve data types, in plain English.

Every real-world data type we work with, with every acronym spelled out. What each one is, what it genuinely answers, where it stops, and which of our markets it exists in. Written for someone who has to brief a colleague, not for someone who already knows.

One thing to hold on to before you read further. No data type is better than another. Each one records a different moment in a patient's care, so the right question is never "which is the best data" but "which moment does my question live in". A claim records a payment. A laboratory value records a measurement. A registry records a diagnosis somebody curated on purpose. Ask what your question needs to have been written down, and the data type chooses itself.

The twelve

Each one, and what it is actually for.

Availability differs by market. A data type that answers your question in one country may not exist, or may not be releasable, in the next, which is why the last column matters as much as the first.

01

Claims, government

A claim is the record a provider submits to get paid. Government claims come from a public scheme.

Every time a subsidised medicine is dispensed or a subsidised service is delivered, a payment record is created. Australia's PBS (Pharmaceutical Benefits Scheme, the national medicine subsidy programme) and MBS (Medicare Benefits Schedule, the national list of subsidised medical services) are the clearest examples. Korea and Taiwan run single-payer systems where near the whole population appears.

What it answers
  • Market sizing and share, with a denominator that is a count rather than a projection
  • Treatment pathways and sequencing, inferred from what was funded and when
  • Persistence and adherence, from the gaps between dispensings
  • Cost and utilisation, because a claim is by definition a payment record
  • Policy impact, because scheme rule changes leave a visible signature
Where it stops

It records what a scheme paid for, not why. There is usually no diagnosis attached to a prescription, no clinical result, and nothing at all about care the scheme does not fund. Private prescriptions and over-the-counter volume are invisible.

Where we have it

Australia, South Korea, Taiwan, New Zealand

02

Claims, insurance

The same shape of record, but the payer is a private or employer-based insurer.

Where private insurance carries a meaningful share of care, the insurer's billing records become a major real-world data source. Japan's corporate and employee health-insurance claims are the largest example in the region, and several carry laboratory values alongside the billing detail.

What it answers
  • Utilisation and cost in the privately insured population
  • Treatment patterns where the public scheme does not fund the therapy
  • Health economics and budget impact for an employed population
  • Comparative effectiveness where laboratory values are attached
Where it stops

The covered population is not the whole population. Employer-based schemes skew heavily to working age, which means elderly and unemployed patients are under-represented, and generalising from them to the whole country is a mistake reviewers will catch.

Where we have it

Japan, China

03

EMR and EHR

EMR is the Electronic Medical Record, the chart one organisation keeps. EHR is the Electronic Health Record, a record designed to follow the patient across organisations.

This is what clinicians actually wrote down: diagnoses, notes, orders, results, and often free text. It is the richest clinical detail available in routine care and, for the same reason, the messiest. Coding systems you will see include ICD (International Classification of Diseases) and SNOMED-CT (Systematized Nomenclature of Medicine, Clinical Terms), and many sources are mapped to a CDM (Common Data Model) such as OMOP (Observational Medical Outcomes Partnership) so that the same analysis can run across several hospitals.

What it answers
  • Standard of care as it is actually practised, not as the guideline describes it
  • Real clinical outcomes, including response and progression
  • Cohort identification using clinical criteria rather than billing proxies
  • Reason for discontinuation and specific indication, which claims cannot carry
  • Validation studies, where a claims-based definition is checked against the chart
Where it stops

Coverage is institutional, not national. If a patient moves between hospitals their history usually restarts. Free text is often excluded or requires separate application, and completeness varies field by field in ways that have to be measured rather than assumed.

Where we have it

Hong Kong, Singapore, China, Japan, South Korea, Taiwan, India

04

Hospital activity

Counts of what hospitals did: admissions, procedures, length of stay, day cases, emergency presentations.

Administrative activity data, usually published at facility or jurisdiction level rather than patient level. In Australia this includes elective surgery waiting-list data recorded against the intended procedure, plus emergency, admitted-patient, safety and cost measures for each facility.

What it answers
  • Where a procedure is actually performed, and at what volume
  • Facility-level targeting and territory design for a field team
  • Peer benchmarking of one hospital against comparable hospitals
  • Capacity as a proxy for where a trial site or a service could be supported
  • Trend in a procedure over time, which often precedes a change in therapy mix
Where it stops

Activity data tells you what was done, not what it cost the hospital, who decided it, or what happened to the patient afterwards. Published series are usually annual, so it cannot answer a question that turns on a month.

Where we have it

Australia, New Zealand, Hong Kong, Singapore, plus network-level data in India

05

Dispensing

The record created when a medicine is physically handed to a patient at a pharmacy.

Closer to real use than a prescription, because a written prescription that is never filled leaves no dispensing record. In subsidised systems the dispensing record is also the payment record, which is why dispensing and government claims often arrive together.

What it answers
  • Actual uptake rather than intent to treat
  • Refill patterns, from which persistence and adherence are estimated
  • Switching between molecules within a class
  • Seasonality and stockpiling effects, which distort a quarter if you do not model them
  • Launch tracking, week by week rather than quarter by quarter
Where it stops

It shows the medicine leaving the pharmacy, not the patient taking it. Under-co-payment dispensing may be incompletely captured in some schemes, which understates low-cost generics.

Where we have it

Australia, South Korea, Taiwan, Japan, China

06

Laboratory results

The numeric values from diagnostic tests, over time.

The difference between knowing a patient was treated and knowing whether they responded. Laboratory data is what makes biomarker-defined cohorts possible, and it is the field most often missing from pure claims sources.

What it answers
  • Response and progression, rather than treatment alone
  • Biomarker-defined cohort construction and eligible patient volume
  • Testing rates and testing gaps, which is a commercial question as much as a clinical one
  • Disease severity stratification, where a claims code cannot distinguish it
  • Safety monitoring, where a laboratory value is the earliest signal
Where it stops

Reference ranges and units differ between laboratories, so harmonisation is real work rather than a mapping table. Where only a selected subgroup is tested, the tested cohort is biased by construction and the size of that bias has to be stated.

Where we have it

Japan, Hong Kong, Singapore, Taiwan, China, New Zealand

07

Registry

A collection built on purpose, for one disease, one procedure or one product.

Because a registry is designed rather than accumulated, it captures fields nobody bothers to record elsewhere: stage, grade, histology, biomarker status, and structured outcome. Taiwan's cancer registry, which covers more than 98 per cent of cancer patients and links to national claims, is the strongest example in the region.

What it answers
  • Epidemiology of a specific disease, at a defined completeness
  • Cancer staging and histology, which claims cannot carry
  • Outcome and survival research with a validated denominator
  • Guideline development and quality-of-care benchmarking
  • Cohort enrichment, where a registry identifies patients and claims describe their treatment
Where it stops

A registry only contains what it was built to capture, and only for the population it enrols. Reporting lags are often longer than claims because cases are traced and corrected before release.

Where we have it

Taiwan, New Zealand, Singapore, India

08

Biobank

A collection of biological samples, linked to health information about the people who gave them.

The bridge between what is in a person's biology and what happened to them clinically. Taiwan Biobank holds 267,000 participants with questionnaire, physical and blood examination data, followed every two to four years, with multi-omics layers on subsets.

What it answers
  • Genetic risk and polygenic risk score research
  • Biomarker discovery and validation against real outcomes
  • Drug target identification and validation
  • Lifestyle and environmental exposure analysis alongside clinical data
  • Validation of self-reported conditions against measured values
Where it stops

Participants volunteer, so a biobank cohort is not a random sample of the population. Sample sizes can be limiting for rare events, and follow-up intervals of two to four years are too long for some questions.

Where we have it

Taiwan, Singapore, Australia

09

Genomics and precision medicine

Sequencing and molecular profiling data, linked to what happened to the patient.

Terms you will meet here: NGS (Next-Generation Sequencing, high-throughput sequencing of many genes at once), WGS (Whole-Genome Sequencing), SNP (Single Nucleotide Polymorphism, a single-letter variation in DNA) and GWAS (Genome-Wide Association Study, which tests millions of variants against a trait). A large gene panel typically covers four hundred to seven hundred genes.

What it answers
  • Biomarker prevalence in a real population, not a trial population
  • Eligible patient volume for a targeted therapy
  • Testing gap analysis: who should have been tested and was not
  • Outcome by molecular subgroup
  • Companion diagnostic strategy and precision-medicine positioning
Where it stops

Tested populations are selected by definition, usually towards more advanced or better-resourced patients. Panel composition differs between sources, so a variant absent from one dataset may simply never have been looked for.

Where we have it

China, Taiwan, Singapore, Australia, India

10

Linked data

Two or more sources joined at the patient level, so one journey can be followed across settings.

The most valuable data type and the most heavily governed, because linkage is exactly what privacy law is designed to control. Done properly, personal identifiers are separated from clinical content and the join is performed by an accredited linkage authority, not by the analyst.

What it answers
  • Patient journeys across primary care, hospital and pharmacy
  • Survival and mortality outcomes, where a death registry is one of the linked sources
  • Full treatment pathway with clinical detail and cost together
  • Health economic models that need utilisation and outcome in one record
  • Real-world comparative effectiveness with adequate covariate control
Where it stops

Linkage is never automatic. It requires approval, it usually adds months, and match quality is itself a variable that has to be reported. Records cannot be joined across national health systems, so multi-market work is harmonisation rather than linkage.

Where we have it

Australia, New Zealand, Taiwan

11

Population data

Census, demographic and vital statistics. Not clinical, but indispensable.

The denominator. Without it a count is a number rather than a rate, and rates are what let you compare a small state with a large one, or this year with last year after the population has grown.

What it answers
  • Per-capita rates, so a big market is not mistaken for a strong one
  • Age and sex standardisation, without which two populations are not comparable
  • Incidence and prevalence estimation from a case count
  • Market potential, expressed as treated share of an estimated eligible population
  • Geographic normalisation for territory design
Where it stops

It describes populations rather than patients, and it is usually published annually with a lag. It cannot tell you anything about an individual, which is the entire point of it.

Where we have it

All nine markets

12

Epidemiological studies

Existing cohort, surveillance and survey studies conducted by academic or public health groups.

Often the fastest route to a defensible incidence or prevalence figure, because somebody has already done the hard work of assembling and validating a cohort. Also the route to lifestyle and behavioural variables that routine health data does not record.

What it answers
  • Incidence and prevalence with a published, citable method
  • Risk factor analysis, including behavioural factors
  • Burden of disease estimation for a health economic case
  • Natural history of disease, where a cohort has long follow-up
  • Benchmarking your own real-world finding against an independent estimate
Where it stops

The study answers the question it was designed to answer. Reusing it for a different question means inheriting its inclusion criteria, its time period and its geography, and saying so.

Where we have it

All nine markets, availability varies by disease area

The other acronyms

The ones that turn up in every meeting.

RWD and RWE

Real-World Data is the data itself, collected in routine care. Real-World Evidence is the conclusion you draw from analysing it. People use them interchangeably and should not.

HTA

Health Technology Assessment. The body that decides whether a health system will pay for a treatment, and at what price. Every market has one and none of them want the same evidence.

HEOR

Health Economics and Outcomes Research. The discipline that quantifies what a treatment costs and what it achieves, usually to support an access or reimbursement case.

ICD and ATC

International Classification of Diseases codes conditions. Anatomical Therapeutic Chemical classification codes medicines. Both change over time, which quietly breaks trend analysis if nobody accounts for it.

OMOP and CDM

A Common Data Model is a shared structure that lets one analysis run across differently shaped databases. OMOP is the most widely used one in observational research.

IRB and ethics

An Institutional Review Board approves research involving human subjects or their data. In most Asia-Pacific markets its calendar, not your analysis, decides your timeline.

De-identification

Removing or obscuring the fields that could identify a person. Performed at source by the party holding the records, before anything reaches an analyst, and never reversed.

Small-cell suppression

Hiding counts below a threshold so an individual cannot be inferred from a rare combination. It is why some cells in a published table say "not published" rather than zero.

PPV

Positive Predictive Value. Of the patients a definition flags as having a condition, the share who genuinely have it. The number to ask for whenever somebody hands you a case definition.

Choosing between them

Most programmes fail because the data type was chosen before the question was written down.

  • Write the question as a specification first. Population, exposure, comparator, outcome and time window. If you cannot write those five, no data type will rescue you.
  • Ask which of the five has to have been recorded. That single test eliminates most of the twelve immediately.
  • Then ask whether it is permitted in that market. Existing and permitted are different questions, and the second one decides your timeline.
  • Assume you will need two types, not one. Claims plus laboratory, or registry plus claims, answers far more than either alone.
  • Decide what you will do if the answer is no. Knowing that in advance is what turns a failed feasibility into a redesign rather than a write-off.

Not sure which data type your question needs?

That is the first piece of work, not a prerequisite for it. Describe the decision and we will tell you which of the twelve can answer it, and in which market.

De-identified patient-level data and publicly available data, through local alliance partners. De-identification at source. Analysis in-market, under local law. Never identifiable records.