Original Research and Data Studies: Complete Content Specification
Build an original research study with a transparent methodology, defined sample, honest limitations, verifiable data, and findings people can confidently cite.
An original research page publishes evidence the organization created or directly collected, explains how that evidence was produced, and lets another person inspect the basis of its claims. Its contract is question, method, sample, findings, limitations, and verification. If any of those pieces is hidden, the page may still be promotion or commentary, but it is not a dependable study.
Primary evidence attracts links through usefulness and trust; linkability is not the research objective.
The format belongs in the SEO post types system when an awareness-stage audience needs new evidence rather than another opinion, definition, or compilation.
Questions it answers
The page should let a careful reader answer these questions without contacting the author:
- What question did the research test, what was measured or asked, and during which dates?
- Where did the records or participants come from, what was included or excluded, and how did cleaning change the sample?
- What analysis produced each finding, how uncertain is it, and what cannot be concluded?
- Where can readers inspect the data, instrument, codebook, or analysis?
The headline finding should be understandable on its own, but it must not outrun the method. “Pages with shorter response times were cited more often in this dataset” is a bounded association. “Speed makes AI engines cite a page” is a causal claim and needs a design capable of isolating causation.
Study vs. survey: the instrument is not the conclusion
A data study analyzes observed, measured, experimental, transactional, or administrative records. A survey collects self-reported answers from a sampled group using a questionnaire or interview. A survey is one research method, not a synonym for research, and it measures what respondents report under the survey conditions.
That distinction changes the headline. A behavioral dataset may support “42 of 60 audited pages contained X” if the audit rules are reproducible. A survey may support “42 of 60 respondents said they use X.” It cannot silently become “70% of companies use X,” because respondents may not represent companies, the sample may not represent the market, and reported behavior may differ from observed behavior.
When to use this post type
Use original research when an important question remains unanswered, the organization has lawful access to suitable evidence, and the team can publish enough method detail for scrutiny. “We queried our database” is not a method until the page defines its records, coverage, transformations, exclusions, and biases.
Choose a confusable sibling when the reader’s actual job is different:
| Post type | Choose it when | Evidence ownership | Boundary from original research |
|---|---|---|---|
| Original research | A new answer to a bounded question | Publisher creates or collects it | This is the reference type |
| statistics roundup | Current numbers from existing sources | External studies | Cite originals; do not relabel curation as new research |
| case study | Proof from one implementation | Customer or project evidence | One intervention is not a population estimate |
| myth-busting post | A belief tested and replaced | New or existing evidence | A disputed claim, not a research question, organizes it |
| ultimate guide | Comprehensive topic coverage | Mostly synthesis | A study supports a chapter; the guide must not duplicate it |
Do not commission a survey merely because a percentage makes an easy headline. Use the method that matches the construct: the underlying idea being measured, such as trust, adoption, performance, or visibility.
Best for these business types
This ranking weighs distinctive data access, editorial demand, and safe publication.
- Media publishers and affiliates . Research can become a recurring editorial franchise and primary source. Disclose commercial relationships, selection criteria, and sponsor influence.
- SaaS . Product events, audits, benchmarks, and experiments can reveal category behavior. Prevent re-identification and never present customers as the whole market.
- B2B services . Specialists can code audits, project records, interviews, or procurement patterns. Keep the method separate from the sales conclusion.
- Marketplaces . Transaction, supply, availability, and search data show behavior at scale. Geography, marketplace rules, and off-platform activity constrain generalization.
- Ecommerce . Tests, returns, reviews, inventory, and purchases can answer practical questions. State customer mix, seasonality, and test controls.
- Finance, fintech, and insurance . Proprietary records and surveys can illuminate cost, access, and behavior, but privacy, regulation, selection, and advice boundaries demand stronger review.
Healthcare can also produce valuable research, but public-facing original studies involving patient data or clinical conclusions need research governance, consent, privacy protection, and specialist review beyond an ordinary content workflow.
Search intent
Search intent is the outcome a person wants from a query. Original research usually serves awareness-stage evidence seeking: a journalist needs a source, a practitioner needs a benchmark, an analyst needs a number, or a prospective buyer wants to understand a market problem before comparing solutions.
A search engine results page may mix academic papers, company reports, press releases, news coverage, and derivative roundups. Inspect whether titles lead with a question, year, population, or finding, then check whether each page supplies a method, sample, limitations, and data.
AI answers often detach a number from its population, dates, units, or uncertainty. Write every key result as a self-contained claim: finding + population + period + qualifier. Keep a method note beside charts and a stable source URL. Never leave “users” to mean respondents, customers, visitors, or all adults by inference.
Record the query, language, country, device, date, signed-in state, and interface. The capture documents an observed result shape; it does not prove stable demand or permanent ranking behavior.
Page structure
Target 2,200–3,500 words for a focused study, excluding appendices, instruments, codebooks, and machine-readable data. Complexity should expand the method and limitations, not the promotional introduction.
| Section | Word range | Purpose | Status |
|---|---|---|---|
| Hero and headline finding | 80–140 | Name question, population, period, and bounded result | Required |
| Key findings | 120–220 | Preview three to five qualified results | Required |
| Research question and expectations | 100–220 | Define questions, constructs, hypotheses, and planned comparisons | Required; preregistration conditional |
| Methodology | 350–650 | State design, sources, dates, instrument, variables, cleaning, exclusions, and analysis | Required before detailed findings |
| Sample and coverage | 180–320 | Show selection, sample flow, subgroups, and coverage limits | Required |
| Findings | 600–1,000 | Present results with charts or tables, uncertainty, and interpretation | Required |
| Robustness checks | 150–350 | Test alternative definitions, exclusions, or models | Conditional; required for high-stakes or model-dependent claims |
| Limitations | 180–350 | State selection, measurement, confounding, missingness, generalization, and time limits | Required |
| Data and verification | 120–260 | Link data, documentation, analysis, licence, and version or explain restrictions | Required |
| Implications | 150–300 | Translate evidence into bounded actions | Required |
| Sources, authorship, and updates | 100–220 | Credit inputs, roles, dates, versions, corrections, and conflicts | Required |
| FAQ | 300–550 | Resolve method, sample, reuse, data, schema, and update questions | Required; 5–8 questions |
| CTA | 40–90 | Offer data inspection, topic tracking, or citation | Required |
“Methodology first” does not require hiding the headline result until halfway down the page. State the bounded finding in the hero, preview the results, then make the full method and sample available before detailed analysis. That sequence supports scanning without asking readers to trust charts they cannot yet evaluate.
Required elements
| Element | Status | Exact position | Why it belongs there |
|---|---|---|---|
| Direct answer block | Always | First content | State a bounded result with population and period |
| Quick overview and table of contents | Always | After key findings | Route directly to method, sample, limitations, data, and findings |
| Chart block | Conditional | Beside its finding | Show a pattern or comparison that prose cannot convey as efficiently |
| Stat band | Conditional | After hero or method summary | Highlight only fully qualified results with units and sample context |
| Sources block | Always | After data access | Separate external inputs from original evidence and map claims to sources |
| Freshness stamp | Always | Near headline and method | Distinguish fieldwork, analysis, publication, and review dates |
| Update log | Always | After data access | Preserve corrections and material method, data, or conclusion changes |
| Related content block | Conditional | Before FAQ | Route to definitions or applications without duplicating the report |
| FAQ structure | Always | Before CTA | Match method and reuse questions to structured records |
| CTA block | Always | Final block | Ask readers to inspect, cite, monitor, or apply the research |
Every chart needs a descriptive title, labels, units, sample base, relevant uncertainty, accessible text, source line, and exposed values. A chart screenshot is not data access.
Frontmatter
Follow the frontmatter specification
and set entity = "post-type-original-research" for this specification. For a published study, use a stable research identity such as study-ai-citations-and-response-time-2026; do not use the headline claim as the entity because the conclusion may change after correction or replication.
Use schemaType = "Article" for the editorial report. Add Dataset markup only when the page exposes a genuine dataset with a stable name, description, creator, temporal and spatial coverage where relevant, licence, distribution or access route, version, and variable documentation. A chart image or three headline numbers is not a dataset. Add FAQPage only when visible questions exactly match the frontmatter records and current search-engine policy supports it.
Required metadata includes canonical URL, publication and fieldwork dates, analysis date, version, authors, reviewer, sponsor or conflicts, design, population, sample size, geography, playbook taxonomy, schema, links, and FAQs. Add the repository, licence, codebook, instrument, code, and correction contact where supported. Set screenshotsPending = true while capture comments remain.
Full example
The following skeleton is copy-pasteable. Its example uses a coded audit of public pages, so the unit is a page, not a person or company. Replace every bracketed field; do not publish brackets or infer unavailable facts.
# Do cited pages state their update date? A coded audit of [N] pages
**Headline finding:** Of [final N] public pages cited for [prompt set] during [dates], [count] displayed a visible content date. This descriptive result does not show that a date causes citation.
**Fieldwork:** [dates]
**Analysis version:** [version and date]
**Data and codebook:** [stable repository URL]
**Licence:** [licence]
## Key findings
- [Finding with numerator, denominator, unit, and scope.]
- [Subgroup finding with each subgroup base.]
- [Null or mixed finding that constrains the headline.]
## Research question
Primary question: Among pages in [sampling frame], how many displayed [operational definition] at capture? Secondary questions: Did prevalence differ by [planned groups], and did the date match [comparison source]?
## Methodology
### Design and sampling frame
We conducted a cross-sectional coded audit of [source records] collected by [process] for [market, language, and period]. The unit was one canonical page URL.
### Inclusion and exclusion rules
We included [rules] and excluded [rules with reasons]. We resolved duplicates using [rule] and logged redirects, inaccessible pages, and ambiguous records.
### Variables and coding
“Visible content date” meant [definition]. The codebook covered [categories and edge cases]. [Number] reviewers used [overlap and adjudication]. Reliability was [metric, value, eligible base].
### Analysis
We report counts and proportions with [uncertainty method]. Subgroups were [planned/exploratory]; missing values remain visible. Analysis used [software/version], with code at [URL].
## Sample flow
Starting records: [N]
Duplicates removed: [N]
Excluded after eligibility review: [N, reasons]
Unavailable at capture: [N]
Final analyzed pages: [N]
## Findings
### Finding 1: [bounded result]
[Count] of [base N] pages met the definition ([proportion and interval]). [Non-causal interpretation.]
| Group | Met definition | Eligible base | Proportion | Notes |
|---|---:|---:|---:|---|
| [Group A] | [n] | [N] | [%] | [scope] |
| [Group B] | [n] | [N] | [%] | [scope] |
### Finding 2: [bounded subgroup result]
[Result, subgroup bases, uncertainty, and whether the comparison was planned.]
### Sensitivity check
When we changed the definition from [A] to [B], [result]. This [did/did not] change the bounded conclusion because [reason].
## Limitations
This sample covers [scope], not [excluded population]. A cross-sectional audit cannot establish cause. Coding may miss [measurement limit], [confounders] were uncontrolled, and small subgroups are descriptive.
## Data and verification
Download [cleaned data], [codebook], [analysis], and [chart values]. We withheld [field] because [specific restriction]. Send questions or corrections to [contact].
## Implications
[Two or three bounded implications tied directly to findings.] The results do not justify [tempting overreach].
## Sources and disclosures
[External source records, sponsor role, author roles, conflicts, and review statement.]
## Update log
- [Date, version]: Initial publication.
- [Date, version]: [Correction or material change, affected output, whether conclusion changed.]
## FAQ
[Visible questions matching frontmatter records.]
## Cite this study
[Canonical title, authors or organization, publication year, version, canonical URL, dataset identifier.]
The skeleton records excluded and unavailable pages because the path from starting records to final sample is part of the finding.
Design gallery
Use the same study and numbers across variants. Every layout must keep methodology, limitations, and data access reachable without visual interpretation.
Do not convert responsive states into separate editorial versions. The mobile layout must preserve qualifiers beside numbers rather than pushing them into hidden tooltips.
Quality checklist
- The question predates the preferred headline and names population, construct, and time boundary.
- The design can answer the stated question; causal words appear only when the design supports causal inference.
- The unit of analysis is explicit and stays consistent across headline, chart, table, and prose.
- Collection, analysis, publication, and review dates are distinguishable.
- The sampling frame, selection or recruitment process, eligibility rules, and exclusions are reproducible.
- Survey reporting includes wording, options, order effects, mode, recruitment, fieldwork, weighting, and respondent bases.
- Every percentage has an eligible denominator; every average names its unit and population.
- Missing data, duplicate handling, outliers, transformations, subgroup bases, and exploratory analysis are disclosed.
- Charts have accessible text and expose plotted values.
- Limitations address selection, measurement, confounding, missingness, generalization, and time—not only “more research is needed.”
- The dataset, instrument, codebook, and analysis are available at the safest useful level, or each restriction is explained.
- Sponsors, roles, conflicts, privacy, consent, licences, and external sources are disclosed.
- The canonical study URL, citation format, version, correction route, and update log are present.
- At least one reviewer can reproduce a headline number from the published materials.
Common mistakes
The defining mistake is headline-first research: choosing a dramatic claim, then searching the data for a route to it. That invites flexible definitions, selective exclusions, multiple unreported comparisons, and a conclusion that disappears under a reasonable alternative analysis.
Other failures are specific to this post type:
- Calling a customer sample “the market.” Product users differ from non-users. Name the sampled population and explain selection effects.
- Treating survey answers as observed behavior. “Respondents intend to buy” is not “buyers will purchase.” Report what the instrument measured.
- Hiding the denominator. “Half prefer A” is meaningless without the eligible base, response options, and treatment of missing answers.
- Changing definitions after seeing results. If a threshold or subgroup was exploratory, label it. Do not present post-hoc choices as a planned hypothesis.
- Reporting only favorable cuts. Show null, mixed, and contrary results. Repeated subgroup testing produces chance patterns.
- Equating correlation with causation. A relationship can reflect reverse causality, shared causes, selection, or coincidence. Match the verb to the design.
- Making raw data public by default. Transparency does not override consent, privacy, security, licences, or contracts; de-identified records may still be linkable.
- Burying limitations after the CTA. Put the qualification where readers and extraction systems encounter the conclusion.
- Calling a press release the study. The canonical report must contain the method and evidence. Media copy should link to it rather than become the only accessible artifact.
- Silently overwriting a correction. Version the materials, describe what changed, and say whether the conclusion moved.
Internal linking
The study owns its question, method, dataset, findings, and limitations. Other pages may summarize a result, but should link to the canonical report rather than compete with it.
- A statistics roundup may include the result among external evidence, but it must cite the study and preserve the sample and date. It must not host a competing copy of the findings.
- A case study may apply the result to one implementation, but that outcome does not validate the population finding or replace the research method.
- A myth-busting article may use the study to test a belief, but it owns the claim-verdict-replacement sequence rather than the dataset.
- An ultimate guide may summarize the result inside broad teaching, but the study remains the destination for method, limitations, and downloads.
Link to necessary definitions, the relevant application, dataset, and measurement framework. Avoid separate summary, press-release, findings, and chart URLs unless each has a distinct job and canonical relationship.
How to measure results
Research performance is a chain: discovery, selection as a source, accurate reuse, qualified engagement, dataset use, earned coverage, and an appropriate downstream action. Follow how we measure results to define the baseline, observation window, comparison, and limits before launch.
In AmICited, use Prompt Tracking for the research question, headline finding, benchmark queries, and adjacent “what does the data show?” formulations. Use source and citation intelligence to inspect which URL AI answers cite and whether the response preserves the population, sample, dates, and non-causal boundary. In the AmICited Cockpit , compare AI visibility, cited URLs, organic discovery, and chosen on-site events over the same window.
Separate canonical-page citations from unlinked mentions and derivative coverage. Measure method reading, data and instrument access, citation-copy actions, subscriptions, and qualified follow-on visits. Log each earned reference, reused claim, destination, and correction.
Do not translate coverage volume into scientific validity or revenue attribution. Judge whether reuse is accurate, authoritative sources cite the canonical page, and the asset still answers its intended question.
FAQ
What makes a report original research?
A report is original research when the publisher creates or directly collects the evidence, documents how it was produced, analyzes it against a stated question, and makes the result verifiable. Reformatting numbers from other publishers is synthesis, not original research.
What is the difference between a data study and a survey?
A data study analyzes observed, measured, experimental, or administrative records. A survey asks sampled people to report attitudes, experiences, intentions, or characteristics. A survey can be original research, but its responses are not direct measurements of behavior unless the design separately observes that behavior.
Does an original research page have to publish its raw data?
Publish the most detailed data that consent, privacy, licensing, security, and contractual limits allow. When row-level data cannot be released, provide an aggregated dataset, data dictionary, analysis code where practical, suppression rules, and a precise explanation of what is unavailable and why.
How large should the sample be?
There is no universal minimum. Sample adequacy depends on the research question, design, expected variation, intended subgroup analysis, uncertainty target, and selection process. Justify the sample before collecting or analyzing data instead of declaring it sufficient because it produced a striking result.
Should methodology appear before the findings?
Yes. Give the headline result briefly, then put the method and sample before the detailed findings. Readers need to know what was measured, who or what was included, and where the limits are before they interpret charts and claims.
Which schema should original research use?
Use Article for the editorial report. Add Dataset only when a real downloadable or queryable dataset is published with the required metadata and a stable landing page. Use FAQPage only when visible questions match the structured records and current policy supports it.
When should an original study be updated?
Update when the underlying population, source system, method, definitions, or market has changed enough to affect interpretation. Preserve the original fieldwork dates and version history; do not silently replace an old sample with a new one.
More tutorials in this section
Ready to put it into practice?
Free check · 7-day trial · no credit card