Academy

Original Research and Data Studies: Complete Content Specification

Build an original research study with a transparent methodology, defined sample, honest limitations, verifiable data, and findings people can confidently cite.

16 min read

An original research page publishes evidence the organization created or directly collected, explains how that evidence was produced, and lets another person inspect the basis of its claims. Its contract is question, method, sample, findings, limitations, and verification. If any of those pieces is hidden, the page may still be promotion or commentary, but it is not a dependable study.

Primary evidence attracts links through usefulness and trust; linkability is not the research objective.

The format belongs in the SEO post types system when an awareness-stage audience needs new evidence rather than another opinion, definition, or compilation.

Questions it answers

The page should let a careful reader answer these questions without contacting the author:

  • What question did the research test, what was measured or asked, and during which dates?
  • Where did the records or participants come from, what was included or excluded, and how did cleaning change the sample?
  • What analysis produced each finding, how uncertain is it, and what cannot be concluded?
  • Where can readers inspect the data, instrument, codebook, or analysis?

The headline finding should be understandable on its own, but it must not outrun the method. “Pages with shorter response times were cited more often in this dataset” is a bounded association. “Speed makes AI engines cite a page” is a causal claim and needs a design capable of isolating causation.

Study vs. survey: the instrument is not the conclusion

A data study analyzes observed, measured, experimental, transactional, or administrative records. A survey collects self-reported answers from a sampled group using a questionnaire or interview. A survey is one research method, not a synonym for research, and it measures what respondents report under the survey conditions.

That distinction changes the headline. A behavioral dataset may support “42 of 60 audited pages contained X” if the audit rules are reproducible. A survey may support “42 of 60 respondents said they use X.” It cannot silently become “70% of companies use X,” because respondents may not represent companies, the sample may not represent the market, and reported behavior may differ from observed behavior.

Name what the method actually observed
A poll of intentions does not measure future behavior. A cross-sectional dataset does not prove cause. An analysis of customers does not automatically represent non-customers. Put the unit, population, and time window in the claim.
Logo

Ready to Monitor Your AI Visibility?

Track how AI chatbots mention your brand across ChatGPT, Perplexity, and other platforms.

When to use this post type

Use original research when an important question remains unanswered, the organization has lawful access to suitable evidence, and the team can publish enough method detail for scrutiny. “We queried our database” is not a method until the page defines its records, coverage, transformations, exclusions, and biases.

Choose a confusable sibling when the reader’s actual job is different:

Post typeChoose it whenEvidence ownershipBoundary from original research
Original researchA new answer to a bounded questionPublisher creates or collects itThis is the reference type
statistics roundupCurrent numbers from existing sourcesExternal studiesCite originals; do not relabel curation as new research
case studyProof from one implementationCustomer or project evidenceOne intervention is not a population estimate
myth-busting postA belief tested and replacedNew or existing evidenceA disputed claim, not a research question, organizes it
ultimate guideComprehensive topic coverageMostly synthesisA study supports a chapter; the guide must not duplicate it

Do not commission a survey merely because a percentage makes an easy headline. Use the method that matches the construct: the underlying idea being measured, such as trust, adoption, performance, or visibility.

Best for these business types

This ranking weighs distinctive data access, editorial demand, and safe publication.

  1. Media publishers and affiliates . Research can become a recurring editorial franchise and primary source. Disclose commercial relationships, selection criteria, and sponsor influence.
  2. SaaS . Product events, audits, benchmarks, and experiments can reveal category behavior. Prevent re-identification and never present customers as the whole market.
  3. B2B services . Specialists can code audits, project records, interviews, or procurement patterns. Keep the method separate from the sales conclusion.
  4. Marketplaces . Transaction, supply, availability, and search data show behavior at scale. Geography, marketplace rules, and off-platform activity constrain generalization.
  5. Ecommerce . Tests, returns, reviews, inventory, and purchases can answer practical questions. State customer mix, seasonality, and test controls.
  6. Finance, fintech, and insurance . Proprietary records and surveys can illuminate cost, access, and behavior, but privacy, regulation, selection, and advice boundaries demand stronger review.

Healthcare can also produce valuable research, but public-facing original studies involving patient data or clinical conclusions need research governance, consent, privacy protection, and specialist review beyond an ordinary content workflow.

Search intent

Search intent is the outcome a person wants from a query. Original research usually serves awareness-stage evidence seeking: a journalist needs a source, a practitioner needs a benchmark, an analyst needs a number, or a prospective buyer wants to understand a market problem before comparing solutions.

A search engine results page may mix academic papers, company reports, press releases, news coverage, and derivative roundups. Inspect whether titles lead with a question, year, population, or finding, then check whether each page supplies a method, sample, limitations, and data.

AI answers often detach a number from its population, dates, units, or uncertainty. Write every key result as a self-contained claim: finding + population + period + qualifier. Keep a method note beside charts and a stable source URL. Never leave “users” to mean respondents, customers, visitors, or all adults by inference.

Record the query, language, country, device, date, signed-in state, and interface. The capture documents an observed result shape; it does not prove stable demand or permanent ranking behavior.

Page structure

Target 2,200–3,500 words for a focused study, excluding appendices, instruments, codebooks, and machine-readable data. Complexity should expand the method and limitations, not the promotional introduction.

SectionWord rangePurposeStatus
Hero and headline finding80–140Name question, population, period, and bounded resultRequired
Key findings120–220Preview three to five qualified resultsRequired
Research question and expectations100–220Define questions, constructs, hypotheses, and planned comparisonsRequired; preregistration conditional
Methodology350–650State design, sources, dates, instrument, variables, cleaning, exclusions, and analysisRequired before detailed findings
Sample and coverage180–320Show selection, sample flow, subgroups, and coverage limitsRequired
Findings600–1,000Present results with charts or tables, uncertainty, and interpretationRequired
Robustness checks150–350Test alternative definitions, exclusions, or modelsConditional; required for high-stakes or model-dependent claims
Limitations180–350State selection, measurement, confounding, missingness, generalization, and time limitsRequired
Data and verification120–260Link data, documentation, analysis, licence, and version or explain restrictionsRequired
Implications150–300Translate evidence into bounded actionsRequired
Sources, authorship, and updates100–220Credit inputs, roles, dates, versions, corrections, and conflictsRequired
FAQ300–550Resolve method, sample, reuse, data, schema, and update questionsRequired; 5–8 questions
CTA40–90Offer data inspection, topic tracking, or citationRequired

“Methodology first” does not require hiding the headline result until halfway down the page. State the bounded finding in the hero, preview the results, then make the full method and sample available before detailed analysis. That sequence supports scanning without asking readers to trust charts they cannot yet evaluate.

Required elements

ElementStatusExact positionWhy it belongs there
Direct answer blockAlwaysFirst contentState a bounded result with population and period
Quick overview and table of contentsAlwaysAfter key findingsRoute directly to method, sample, limitations, data, and findings
Chart blockConditionalBeside its findingShow a pattern or comparison that prose cannot convey as efficiently
Stat bandConditionalAfter hero or method summaryHighlight only fully qualified results with units and sample context
Sources blockAlwaysAfter data accessSeparate external inputs from original evidence and map claims to sources
Freshness stampAlwaysNear headline and methodDistinguish fieldwork, analysis, publication, and review dates
Update logAlwaysAfter data accessPreserve corrections and material method, data, or conclusion changes
Related content blockConditionalBefore FAQRoute to definitions or applications without duplicating the report
FAQ structureAlwaysBefore CTAMatch method and reuse questions to structured records
CTA blockAlwaysFinal blockAsk readers to inspect, cite, monitor, or apply the research

Every chart needs a descriptive title, labels, units, sample base, relevant uncertainty, accessible text, source line, and exposed values. A chart screenshot is not data access.

Frontmatter

Follow the frontmatter specification and set entity = "post-type-original-research" for this specification. For a published study, use a stable research identity such as study-ai-citations-and-response-time-2026; do not use the headline claim as the entity because the conclusion may change after correction or replication.

Use schemaType = "Article" for the editorial report. Add Dataset markup only when the page exposes a genuine dataset with a stable name, description, creator, temporal and spatial coverage where relevant, licence, distribution or access route, version, and variable documentation. A chart image or three headline numbers is not a dataset. Add FAQPage only when visible questions exactly match the frontmatter records and current search-engine policy supports it.

Required metadata includes canonical URL, publication and fieldwork dates, analysis date, version, authors, reviewer, sponsor or conflicts, design, population, sample size, geography, playbook taxonomy, schema, links, and FAQs. Add the repository, licence, codebook, instrument, code, and correction contact where supported. Set screenshotsPending = true while capture comments remain.

Full example

The following skeleton is copy-pasteable. Its example uses a coded audit of public pages, so the unit is a page, not a person or company. Replace every bracketed field; do not publish brackets or infer unavailable facts.

# Do cited pages state their update date? A coded audit of [N] pages

**Headline finding:** Of [final N] public pages cited for [prompt set] during [dates], [count] displayed a visible content date. This descriptive result does not show that a date causes citation.

**Fieldwork:** [dates]  
**Analysis version:** [version and date]  
**Data and codebook:** [stable repository URL]  
**Licence:** [licence]

## Key findings

- [Finding with numerator, denominator, unit, and scope.]
- [Subgroup finding with each subgroup base.]
- [Null or mixed finding that constrains the headline.]

## Research question

Primary question: Among pages in [sampling frame], how many displayed [operational definition] at capture? Secondary questions: Did prevalence differ by [planned groups], and did the date match [comparison source]?

## Methodology

### Design and sampling frame

We conducted a cross-sectional coded audit of [source records] collected by [process] for [market, language, and period]. The unit was one canonical page URL.

### Inclusion and exclusion rules

We included [rules] and excluded [rules with reasons]. We resolved duplicates using [rule] and logged redirects, inaccessible pages, and ambiguous records.

### Variables and coding

“Visible content date” meant [definition]. The codebook covered [categories and edge cases]. [Number] reviewers used [overlap and adjudication]. Reliability was [metric, value, eligible base].

### Analysis

We report counts and proportions with [uncertainty method]. Subgroups were [planned/exploratory]; missing values remain visible. Analysis used [software/version], with code at [URL].

## Sample flow

Starting records: [N]  
Duplicates removed: [N]  
Excluded after eligibility review: [N, reasons]  
Unavailable at capture: [N]  
Final analyzed pages: [N]

## Findings

### Finding 1: [bounded result]

[Count] of [base N] pages met the definition ([proportion and interval]). [Non-causal interpretation.]

| Group | Met definition | Eligible base | Proportion | Notes |
|---|---:|---:|---:|---|
| [Group A] | [n] | [N] | [%] | [scope] |
| [Group B] | [n] | [N] | [%] | [scope] |

### Finding 2: [bounded subgroup result]

[Result, subgroup bases, uncertainty, and whether the comparison was planned.]

### Sensitivity check

When we changed the definition from [A] to [B], [result]. This [did/did not] change the bounded conclusion because [reason].

## Limitations

This sample covers [scope], not [excluded population]. A cross-sectional audit cannot establish cause. Coding may miss [measurement limit], [confounders] were uncontrolled, and small subgroups are descriptive.

## Data and verification

Download [cleaned data], [codebook], [analysis], and [chart values]. We withheld [field] because [specific restriction]. Send questions or corrections to [contact].

## Implications

[Two or three bounded implications tied directly to findings.] The results do not justify [tempting overreach].

## Sources and disclosures

[External source records, sponsor role, author roles, conflicts, and review statement.]

## Update log

- [Date, version]: Initial publication.
- [Date, version]: [Correction or material change, affected output, whether conclusion changed.]

## FAQ

[Visible questions matching frontmatter records.]

## Cite this study

[Canonical title, authors or organization, publication year, version, canonical URL, dataset identifier.]

The skeleton records excluded and unavailable pages because the path from starting records to final sample is part of the finding.

Use the same study and numbers across variants. Every layout must keep methodology, limitations, and data access reachable without visual interpretation.

Do not convert responsive states into separate editorial versions. The mobile layout must preserve qualifiers beside numbers rather than pushing them into hidden tooltips.

Quality checklist

  • The question predates the preferred headline and names population, construct, and time boundary.
  • The design can answer the stated question; causal words appear only when the design supports causal inference.
  • The unit of analysis is explicit and stays consistent across headline, chart, table, and prose.
  • Collection, analysis, publication, and review dates are distinguishable.
  • The sampling frame, selection or recruitment process, eligibility rules, and exclusions are reproducible.
  • Survey reporting includes wording, options, order effects, mode, recruitment, fieldwork, weighting, and respondent bases.
  • Every percentage has an eligible denominator; every average names its unit and population.
  • Missing data, duplicate handling, outliers, transformations, subgroup bases, and exploratory analysis are disclosed.
  • Charts have accessible text and expose plotted values.
  • Limitations address selection, measurement, confounding, missingness, generalization, and time—not only “more research is needed.”
  • The dataset, instrument, codebook, and analysis are available at the safest useful level, or each restriction is explained.
  • Sponsors, roles, conflicts, privacy, consent, licences, and external sources are disclosed.
  • The canonical study URL, citation format, version, correction route, and update log are present.
  • At least one reviewer can reproduce a headline number from the published materials.

Common mistakes

The defining mistake is headline-first research: choosing a dramatic claim, then searching the data for a route to it. That invites flexible definitions, selective exclusions, multiple unreported comparisons, and a conclusion that disappears under a reasonable alternative analysis.

Other failures are specific to this post type:

  • Calling a customer sample “the market.” Product users differ from non-users. Name the sampled population and explain selection effects.
  • Treating survey answers as observed behavior. “Respondents intend to buy” is not “buyers will purchase.” Report what the instrument measured.
  • Hiding the denominator. “Half prefer A” is meaningless without the eligible base, response options, and treatment of missing answers.
  • Changing definitions after seeing results. If a threshold or subgroup was exploratory, label it. Do not present post-hoc choices as a planned hypothesis.
  • Reporting only favorable cuts. Show null, mixed, and contrary results. Repeated subgroup testing produces chance patterns.
  • Equating correlation with causation. A relationship can reflect reverse causality, shared causes, selection, or coincidence. Match the verb to the design.
  • Making raw data public by default. Transparency does not override consent, privacy, security, licences, or contracts; de-identified records may still be linkable.
  • Burying limitations after the CTA. Put the qualification where readers and extraction systems encounter the conclusion.
  • Calling a press release the study. The canonical report must contain the method and evidence. Media copy should link to it rather than become the only accessible artifact.
  • Silently overwriting a correction. Version the materials, describe what changed, and say whether the conclusion moved.

Internal linking

The study owns its question, method, dataset, findings, and limitations. Other pages may summarize a result, but should link to the canonical report rather than compete with it.

  • A statistics roundup may include the result among external evidence, but it must cite the study and preserve the sample and date. It must not host a competing copy of the findings.
  • A case study may apply the result to one implementation, but that outcome does not validate the population finding or replace the research method.
  • A myth-busting article may use the study to test a belief, but it owns the claim-verdict-replacement sequence rather than the dataset.
  • An ultimate guide may summarize the result inside broad teaching, but the study remains the destination for method, limitations, and downloads.

Link to necessary definitions, the relevant application, dataset, and measurement framework. Avoid separate summary, press-release, findings, and chart URLs unless each has a distinct job and canonical relationship.

How to measure results

Research performance is a chain: discovery, selection as a source, accurate reuse, qualified engagement, dataset use, earned coverage, and an appropriate downstream action. Follow how we measure results to define the baseline, observation window, comparison, and limits before launch.

In AmICited, use Prompt Tracking for the research question, headline finding, benchmark queries, and adjacent “what does the data show?” formulations. Use source and citation intelligence to inspect which URL AI answers cite and whether the response preserves the population, sample, dates, and non-causal boundary. In the AmICited Cockpit , compare AI visibility, cited URLs, organic discovery, and chosen on-site events over the same window.

Separate canonical-page citations from unlinked mentions and derivative coverage. Measure method reading, data and instrument access, citation-copy actions, subscriptions, and qualified follow-on visits. Log each earned reference, reused claim, destination, and correction.

Do not translate coverage volume into scientific validity or revenue attribution. Judge whether reuse is accurate, authoritative sources cite the canonical page, and the asset still answers its intended question.

FAQ

What makes a report original research?

A report is original research when the publisher creates or directly collects the evidence, documents how it was produced, analyzes it against a stated question, and makes the result verifiable. Reformatting numbers from other publishers is synthesis, not original research.

What is the difference between a data study and a survey?

A data study analyzes observed, measured, experimental, or administrative records. A survey asks sampled people to report attitudes, experiences, intentions, or characteristics. A survey can be original research, but its responses are not direct measurements of behavior unless the design separately observes that behavior.

Does an original research page have to publish its raw data?

Publish the most detailed data that consent, privacy, licensing, security, and contractual limits allow. When row-level data cannot be released, provide an aggregated dataset, data dictionary, analysis code where practical, suppression rules, and a precise explanation of what is unavailable and why.

How large should the sample be?

There is no universal minimum. Sample adequacy depends on the research question, design, expected variation, intended subgroup analysis, uncertainty target, and selection process. Justify the sample before collecting or analyzing data instead of declaring it sufficient because it produced a striking result.

Should methodology appear before the findings?

Yes. Give the headline result briefly, then put the method and sample before the detailed findings. Readers need to know what was measured, who or what was included, and where the limits are before they interpret charts and claims.

Which schema should original research use?

Use Article for the editorial report. Add Dataset only when a real downloadable or queryable dataset is published with the required metadata and a stable landing page. Use FAQPage only when visible questions match the structured records and current policy supports it.

When should an original study be updated?

Update when the underlying population, source system, method, definitions, or market has changed enough to affect interpretation. Preserve the original fieldwork dates and version history; do not silently replace an old sample with a new one.

Measure whether your research becomes a trusted source
Track the questions your study answers, inspect which pages AI engines cite, and verify that reused findings preserve the sample, dates, and limitations.

← All Academy tutorials

Ready to put it into practice?

Free check · 7-day trial · no credit card