Who Decides What an AI Agent Recommends to You? A Black-Box Audit Pilot on Shoes and Wine
Leggi questo articolo in italiano β
In brief
- We built a pilot experiment to test whether a site's "technical merit" (structured data, sitemap, permissions for AI crawlers) is associated with how often an AI agent recommends it.
- Across two product categories (running shoes, wine), the observed pattern is consistent: the most-recommended vendor is never the one with the highest technical score. The quantitative analysis verified vendor-by-vendor is presented only for wine (N=10); for shoes we report the qualitative observation without a numeric table, for a verification limit declared in the "Results" section.
- A Bayesian statistical model finds no evidence of correlation between technical score and recommendation frequency β but the estimate is not stable on a sample this small (N=10), a limit we declare explicitly.
- The most solid finding is not statistical: measuring a site's technical "AI readiness" is harder than it sounds, because e-commerce giants are often invisible to automated audit tools, for different reasons case by case.
- Before the experiment, we conducted a systematic conflict-of-interest check on providers of "GEO/AEO" services and classified available evidence in the literature into three reliability tiers.
The problem: no one can verify why an AI agent recommends a vendor
When an AI assistant answers "where do I buy running shoes," it proposes names. We don't know β and today no one can verify independently β whether those names emerge from product quality, brand recognition, technical completeness of the site, or some combination that even the people who trained the model couldn't fully decompose.
Traceability infrastructure already exists for the transaction: protocols like AP2 link purchase intent, cart, and payment with verifiable cryptographic signatures. No equivalent infrastructure exists for the decision that precedes the transaction β the moment the agent chooses who to recommend. The EU's Digital Services Act mandates transparency for recommender systems, but it was written with social platforms and traditional marketplaces in mind, not autonomous purchasing agents.
Before discussing regulatory compliance, a simpler question comes first: is it even possible, with tools available today, to measure whether a site's technical merit matters for these recommendations? This article recounts an attempt to answer that β including the times we got it wrong along the way.
Upstream research: who sells "GEO/AEO optimization," and on what basis
Before building any experiment, we conducted a systematic check of the landscape of providers selling optimization services for visibility on AI assistants (often labeled GEO β Generative Engine Optimization β or AEO β Answer Engine Optimization). The method applied to each provider followed three checks, in order:
- Declared academic affiliation β do the authors of any cited papers or research have a commercial role in the company citing them as evidence?
- Commercial relationship with cited sources β when a provider cites "independent research" to support its claims, does that source sell similar or related services?
- Citation chain β do the cited sources cite each other in a closed loop, creating an illusion of independent consensus where it's actually the same small group of commercial actors?
This check allowed us to classify publicly available evidence into three tiers, applying the same criterion to every claim gathered through systematic research (STORM method, cross-checked against OpenAlex for the real academic affiliation of cited authors):
| Tier | Definition | Treatment |
|---|---|---|
| Tier 1 | Convergent across sources with opposing incentives (e.g. crawlability, consistency of structured data, correlation with Google ranking as a measurable proxy) | Usable as a basis for technical recommendations |
| Tier 2 | Plausible hypothesis but supported by a single source, often commercial | To be tested with own data before adoption |
| Tier 3 | Actively contested among competing vendors, or lacking sources independent from the vendor | Excluded until own evidence exists |
Most specific optimization promises "to make your site appear in AI answers" fall into Tier 2 or Tier 3: plausible, but not independently verified from the vendor proposing them. Few things β structured data consistency, basic technical crawlability β fall into Tier 1, meaning they have convergent evidence from sources with opposing interests.
This classification motivated the question behind the pilot described below: if even the Tier 1 levers (structured data, technical readability) were implemented perfectly, is there observable evidence that this changes how often an AI agent recommends a site?
The pilot: method
We built a Python pipeline of four sequential stages, applied to two distinct product categories (running shoes and wine) to check whether observations replicated across different domains.
Stage 1 β Collecting recommendations
An AI agent (gemini-flash-lite-latest, queried in static-memory mode β no real-time web search) received two types of queries, each repeated 20 times to capture the model's variance:
- "Opacity" track: a neutral query ("best sites to buy X"), with no constraints.
- "Merit" track (10 variants): queries that explicitly exclude known giants by name (Amazon, Decathlon, Zalando for shoes; Amazon, Tannico for wine), to check whether the model knows niche players even when forced to name them.
Declared methodological note: the model was selected after verifying, via a diagnostic script, that the API key in use is associated with a "canary/beta" environment with access to non-standard models. The -latest alias does not correspond to a version with a fixed identifier: the underlying model can change over time at the provider's discretion, independent of any pipeline changes. This limits the exact reproducibility of the experiment over time.
Stage 2 β Structured extraction
The same model (gemini-flash-lite-latest), in a distinct role as extractor (LLM-as-a-judge), converted the free text of the responses into structured data: first vendor mentioned, full list of vendors mentioned per repetition.
Stage 3 β Technical measurement
Each identified vendor was scanned with Agentabile's Evaluator, which measures structured data (schema.org/JSON-LD), presence of a sitemap and robots.txt with rules for known AI crawlers, and compatibility with agentic purchase protocols (UCP/ACP) where relevant.
Stage 4 β Manual verification and cleanup
This stage, not planned initially at this level of detail, turned out to be the most decisive for the quality of the final dataset β and is described separately in the next section, because it is itself a finding.
An unexpected result: how hard it is to technically measure a site from the outside
During data collection, three categories of problems produced falsely low or outright incorrect technical scores, each individually verified before being corrected or excluded from the final dataset:
| Problem | Verified cause | Example | Treatment |
|---|---|---|---|
| Intermittent network timeout, cause not clearly distinguishable from anti-bot blocking without deeper diagnosis (4 of 6 attempts failed due to timeout, the 2 successful ones produced scores consistent with each other) | The site times out against the automated fetcher | Amazon.it | Excluded from quantitative measurement, reported only in the descriptive track |
| Client-side rendering (SPA) | Valid HTTP 200 response, but near-empty initial HTML (content built by JavaScript after load) | Zalando.it (1 page fetched); Les Caves de Pyrene (14 characters of extracted text across the entire homepage) | Excluded if not measurable; included as a legitimate zero score if verified with concrete evidence |
| Wrong nameβURL resolution | The automated mechanism linking the vendor name cited by the LLM to its website produced wrong matches | A name cited as "Pavin" incorrectly resolved to a dental clinic's website | Manually verified one by one; excluded if ambiguous between multiple non-equivalent candidates |
Of 12 vendors initially identified in the wine "merit" track, 2 were excluded due to ambiguity in URL resolution β not variants of the same site, but short business names colliding with companies entirely unrelated to the wine sector. The final verified dataset counts N=10 measurable vendors for wine.
This cleanup process is, in our view, the most generalizable finding of the pilot: automated third-party technical auditing has real structural limits, distinct from each other (active defense vs. architecture not designed to be read without a browser), which require case-by-case verification β they cannot be assumed solved by a tool, however well designed.
Results
Running shoes domain
Opacity track (neutral query, 20 repetitions):
| Vendor | First mention | Total mentions |
|---|---|---|
| Amazon | 11 | 20 |
| Nike | 5 | 14 |
| Maxi Sport | 2 | 17 |
| Decathlon | 1 | 14 |
Merit track (10 queries with explicit exclusion of giants): the agent revealed a long tail of specialized shops (Maxi Sport, Top4Running, DF Sport Specialist, Koala Sport, Sportler and others), with the same qualitative pattern observed in wine β no evident relationship between mention frequency and vendors' technical scores. We do not present a numeric correlation table for this domain here: unlike wine, where every single vendor was individually verified (previous section), for shoes only two anomalous cases were diagnosed individually. Given the number of URL-resolution errors and skewed scores discovered in wine despite thorough verification, we do not feel confident presenting the shoes table with the same degree of reliability β we prefer to omit it rather than risk publishing data not verified to the same rigor.
Wine domain
Opacity track (neutral query, 20 repetitions):
| Vendor | First mention | Total mentions |
|---|---|---|
| Tannico | 2 | 20 |
| Callmewine | 1 | 20 |
| Vino75 | 6 | 6 |
| Vino.com | 5 | 5 |
| Amazon | 4 | 4 |
Merit track (10 queries with explicit exclusion) β correlation with technical score, N=10 verified measurable vendors:
| Vendor | Mentions | Technical score (Visibility) |
|---|---|---|
| Callmewine | 169 | 29.0 |
| XtraWine | 81 | 55.0 |
| Tannico | 60 | 55.0 |
| SignorVino | 60 | 29.0 |
| Enoteca Properzio | 44 | 55.0 |
| Les Caves de Pyrene | 40 | 0.0 |
| Vino75 | 32 | 14.0 |
| Vinix | 24 | 63.0 |
| Enoteca Naturale | 20 | 57.0 |
| Svinando | 9 | 29.0 |
In wine, the two highest technical scores overall (Vinix 63.0, Enoteca Naturale 57.0) belong to the two least-mentioned vendors in the table.
Statistical analysis and its limits
On the wine dataset (the only one with sufficient N for a model, albeit at the lower bound of the minimum threshold we had set) we applied a Bayesian Poisson regression, with a robustness check we recommend as standard practice for anyone replicating this kind of analysis on small samples:
| Configuration | Coefficient (beta) | 94% credible interval | Includes zero |
|---|---|---|---|
| Full dataset (N=10) | -0.056 | [-0.124, 0.011] | Yes |
| Without the most-mentioned vendor (N=9) | +0.080 | [-0.006, 0.167] | Yes |
In both configurations, the interval includes zero: neither estimate provides statistical evidence of correlation, positive or negative. But the sign of the coefficient flips when removing a single vendor from the sample β the clearest evidence that, at N=10, the point estimate is not stable and should not be presented as a reliable number on its own.
An additional check (posterior predictive check) confirms the same caution: the model fitted on the full dataset drastically underestimates the real variability observed in the data (a case of overdispersion caused by the disproportionate weight of the most-mentioned vendor as a single extreme observation).
What this pilot honestly concludes
In this exploratory sample β two domains, a single AI model (in an atypical, declared configuration), N=10 in the domain with statistical analysis β no coherent pattern emerges between a site's technical readiness and how often an AI agent recommends it. The most-cited vendors are not systematically those with the best technical score.
This does not amount to saying technical merit is irrelevant in general: the sample is small, only one model was tested (and in static-memory mode, without real-time web search β a different mechanism from how assistants like ChatGPT or Perplexity operate when browsing the web), and a plausible confounding factor (a company's overall maturity produces both a better site and more recognition, without either causing the other) was not isolated in this experimental design.
What the pilot shows more solidly is a different, perhaps more useful fact: today there is no reliable, independent way to verify on what basis an AI agent chooses who to recommend β not for an external researcher, and, as we experienced firsthand, not even in a fully automated way for those building the measurement tools themselves.
Next steps
A larger sample (ideally 30-50+ measurable vendors per category, across more product categories, with more AI models at fixed versions, and a scanner capable of real JavaScript rendering to reduce the number of sites "invisible" for architectural reasons) would be needed before drawing robust quantitative conclusions. This work should be understood as method validation β a starting point, not an endpoint.
Sources
- Google Cloud, Announcing Agent Payments Protocol (AP2), September 16, 2025 β official announcement of the AP2 protocol and launch partners (Mastercard, PayPal, American Express, Coinbase, Etsy, and over 60 others). https://cloud.google.com/blog/products/ai-machine-learning/announcing-agents-to-payments-ap2-protocol
- AP2 technical specification and reference repository (Apache 2.0). https://ap2-protocol.org
- Regulation (EU) 2022/2065 of the European Parliament and of the Council of 19 October 2022 on a Single Market For Digital Services (Digital Services Act), OJ L 277, 27.10.2022, p. 1β102, Article 27 (Recommender system transparency). Official consolidated text, EUR-Lex, CELEX identifier 32022R2065. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=uriserv:OJ.L_.2022.277.01.0001.01.ENG
- Agentabile data: scan results cited in the tables collected between August 16 and 19, 2026, raw dataset available on request.
Methodology, raw data, and pipeline code available on request. This article deliberately avoids terms like "demonstrates" or "confirms" in relation to quantitative results, consistent with the exploratory nature and small sample size of the pilot described.