Research · Methodology

AI search research method: audit evidence tiers and stages

In brief. This is the six-step method used to review C-SEO Bench, What Gets Cited, and other primary AI search research. It keeps measured outcomes separate from static audit proxies and requires matching evidence tiers, pipeline stages, and runtime records.

Maintained by Jeff Patterson and Agency Enterprise · Updated August 30, 2026

Evidence standard: overview

The evidence base starts with peer-reviewed or formally accepted work from venues such as NeurIPS, ICLR, KDD, SIGIR, and ACL. Each review uses the primary paper and official proceedings. The reason is traceability: a summary cannot replace the method, data, result, and limits reported by the authors.

Marketing articles, proposed standards, unpublished benchmarks, and vendor posts supply questions. They cannot establish a causal effect or justify a scoring weight on their own.

How the paper-by-paper review works

  1. Record the exact research question, data, models, and experimental regime.
  2. Identify the measured outcome: retrieval, rank, cited words, citation order, or another metric.
  3. Capture effect direction, uncertainty, significance testing, and reported caveats.
  4. Compare the paper with existing tool claims and factors.
  5. State what the paper does not justify changing.
  6. Map any defensible change to a factor, evidence tier, and pipeline stage.

AI search research method terms

Primary source

A primary source refers to the paper or official proceedings that report the original study.

Experimental regime

An experimental regime refers to the data, models, prompts, engines, and controls used in a study.

Measured outcome

A measured outcome means that the review names the exact retrieval, rank, citation, or text metric.

Evidence tier

An evidence tier is defined as the label for the strength and scope of support behind a factor.

Pipeline stage

A pipeline stage refers to technical eligibility, retrieval alignment, citation fitness, or provenance.

Static proxy

A static proxy is a type of observable page feature used when the audit cannot reproduce a private engine measurement.

Translation into a deterministic audit

Many research methods use private engine feedback or a large language model judge. A deterministic page audit cannot reproduce those methods. When the audit uses a static proxy, it labels that proxy as heuristic or conditional. It does not borrow the paper's effect size. This means an observed page feature never inherits a causal claim that the audit did not measure.

Experts set the thresholds and category weights. The values are public and stable within a version. They are not fitted probabilities. The reason for versioned weights is repeatability. As a consequence, a major scoring change requires a new baseline.

Maintenance policy

Research is re-reviewed when a new peer-reviewed result challenges a current factor, introduces a better outcome measure, or changes how a pipeline stage should be modeled. Score changes that alter historical comparability require a migration note and new baselines. This works because a major scoring change is documented before old and new results are compared.

  • Primary-source link required
  • Experimental regime and metric recorded
  • Scope limits written next to the finding
  • Evidence map and runtime registry updated together
  • Contradicted claims removed rather than silently reworded

Key takeaways

  • The source is the original paper or official proceedings.
  • The outcome is recorded exactly as the study measured it.
  • The tier is the stated strength and scope of the evidence.
  • The proxy is labeled when the static audit cannot reproduce the research method.
  • The release is blocked when the evidence map and runtime record disagree.

Bottom line: a research claim becomes an audit factor only after its source, scope, metric, tier, stage, and implementation agree.

Research sources

According to C-SEO Bench, source position dominated its tested rewrites [1]. According toAutoGEO, useful optimization rules varied by engine and query [2]. According toWhat Gets Cited, its conclusions came from 252,000 controlled trials [3].

According to the evidence map, every scored factor records a tier and stage [4]. According to the version 2 migration guide, unsupported presence checks became unscored diagnostics [5].

  1. C-SEO Bench, NeurIPS 2025.
  2. AutoGEO, ICLR 2026.
  3. What Gets Cited, SIGIR 2026.
  4. aiseo-audit evidence map, project research record.
  5. Version 2 migration guide, project documentation.

The release check requires 100% agreement between the evidence map and runtime record for tier and stage.