This paper proposes a transparent framework for studying AI visibility. It does not present a validated index, rank organisations, predict citation probability, or guarantee commercial performance.
Abstract
AI-assisted information systems are becoming an important route through which people discover organisations, products, services, and public information. Existing discussions often combine technical accessibility, search performance, generated answers, citations, and business outcomes into one broad idea of visibility. That combination makes it difficult to identify what was observed, what was inferred, and what remains unknown.
The AI Visibility Index is proposed as a multi-dimensional research framework. It separates source access, identity, information architecture, semantic clarity, evidence, retrieval presence, representation accuracy, attribution, and maintenance. The framework is designed for reproducible observation rather than optimisation folklore. It requires declared test conditions, independent review, uncertainty reporting, and explicit separation between diagnostic measures and outcomes.
Version 0.1 defines the research questions, pilot methodology, data requirements, review process, and conditions under which the framework should be revised or abandoned. Empirical validation is a future phase. No organisation-level result is claimed in this publication.
Keywords: AI visibility, information retrieval, machine discovery, source attribution, entity clarity, evaluation, reproducibility.
Research questions and objectives
The central question is whether a transparent set of technical, semantic, retrieval, evidence, and representation measures can describe meaningful differences in how organisations are discovered and represented in AI-assisted information tasks.
- Which observable website conditions are associated with successful source retrieval?
- Which conditions are associated with accurate organisation and product representation?
- How stable are observations across systems, prompts, dates, locations, and interaction modes?
- Which measures are direct observations and which are diagnostic proxies?
- Can independent reviewers apply the classifications consistently?
- Which changes make a source clearer without encouraging manipulation?
The first objective is definition. The second is reproducibility. The third is usefulness to engineers and information owners. A composite score will not be created unless evidence shows that aggregation is interpretable, stable, and less misleading than reporting dimensions separately.
Methodology
Framework specification
Before outcome testing, the research team will define dimensions, labels, units of analysis, exclusions, query intents, collection limits, and planned analyses. Changes will be recorded in a revision log.
Declared sample
The pilot will use a controlled sample across software and technology, education, healthcare, professional services, nonprofit and community organisations, and public or civic information. Sector, region, size category, language, platform conditions, and selection source will be recorded. A convenience sample will not be described as representative.
Source audit
Researchers will collect public observations about response status, canonical identity, robots controls, rendered content, headings, navigation, entity definitions, structured data, authorship, dates, sources, and internal relationships. Tool version, crawl date, request limits, and exclusions will be recorded.
Fixed discovery tests
Each source will be evaluated against a declared set of information tasks. The record will include the exact prompt or query, system, mode, date, location where relevant, result, selected source, attribution behaviour, and representation assessment. Exploratory observations will remain separate from the fixed test set.
Independent review and replication
At least two reviewers will classify a defined subset. They will record evidence locations, confidence, uncertainty, and conflicts of interest. Disagreements will be preserved and adjudicated by the research lead. The final report will publish the method, sample logic, version information, de-identified records where permitted, and reproduction instructions.
Measurement framework
Version 0.1 proposes these dimensions for testing. They are not validated weights or a final score:
- Access and crawlability: whether the public source can be reached and processed under declared conditions.
- Canonical identity: whether the organisation, domain, products, and official relationships are distinguishable.
- Information architecture: whether important information has stable and coherent locations.
- Semantic clarity: whether sections, claims, labels, and terminology communicate consistent meaning.
- Evidence and provenance: whether important claims identify authorship, dates, sources, or supporting context.
- Freshness and maintenance: whether time-sensitive information is dated and maintained.
- Retrieval presence: whether relevant source material appears in a declared test.
- Representation accuracy: whether an observed description matches documented facts and limitations.
- Attribution: whether an interface identifies or links the source when attribution is available.
Results should first be reported as distributions, examples, missingness, and uncertainty. Any later aggregation must show the effect of alternative scoring choices and must not imply causation.
Ethics, privacy, and threats to validity
The pilot is limited to public information unless a separate agreement establishes another basis. Researchers will not request passwords or private content. Organisation names, screenshots, quotations, metrics, and organisation-level results require the publication permission specified in the participation materials.
AI systems change over time. Providers expose different modes. Prompts, locations, languages, and dates affect observations. Public pages do not reveal every retrieval process. Reviewers may interpret clarity and accuracy differently. A score may encourage gaming if its construction is simplified.
The research team must distinguish evidence from interpretation, disclose relationships with sampled organisations, minimise collected data, retain records for a declared period, and provide a correction process for factual errors. Product or commercial interests must not determine the result.
Review gate and planned outputs
This working paper is intended to produce a public annotation guide, a reproducible audit protocol, a small documented benchmark dataset, a technical report on cross-system stability, and a revision log. The first empirical report will be published only after the pilot is complete and the research lead confirms that the evidence supports the stated claims.
Before that report, the paper should receive technical review, methodology review, and ethics or privacy review where required. Reviewers should be named in the publication record. A publication date is not evidence of peer review.
What would change our mind?
We would revise or abandon the framework if independent reviewers cannot apply its definitions consistently, if repeated tests are too unstable for the intended use, if the dimensions do not distinguish meaningful conditions, or if the framework creates stronger incentives for signal manipulation than for useful public information.
Working paper 001. Pre-review. No validated index score, organisation ranking, or empirical finding is claimed.