πŸŽ‰ Plugin WordPress per WooCommerce disponibile β€” un feed ACP per il tuo catalogo. Scopri di piΓΉ β†’
AgentabileAGENTABILE

2026-08-20

Who Decides What an AI Agent Recommends to You? A Black-Box Audit Pilot on Shoes and Wine

Leggi questo articolo in italiano β†’

In brief

The problem: no one can verify why an AI agent recommends a vendor

When an AI assistant answers "where do I buy running shoes," it proposes names. We don't know β€” and today no one can verify independently β€” whether those names emerge from product quality, brand recognition, technical completeness of the site, or some combination that even the people who trained the model couldn't fully decompose.

Traceability infrastructure already exists for the transaction: protocols like AP2 link purchase intent, cart, and payment with verifiable cryptographic signatures. No equivalent infrastructure exists for the decision that precedes the transaction β€” the moment the agent chooses who to recommend. The EU's Digital Services Act mandates transparency for recommender systems, but it was written with social platforms and traditional marketplaces in mind, not autonomous purchasing agents.

Before discussing regulatory compliance, a simpler question comes first: is it even possible, with tools available today, to measure whether a site's technical merit matters for these recommendations? This article recounts an attempt to answer that β€” including the times we got it wrong along the way.

Upstream research: who sells "GEO/AEO optimization," and on what basis

Before building any experiment, we conducted a systematic check of the landscape of providers selling optimization services for visibility on AI assistants (often labeled GEO β€” Generative Engine Optimization β€” or AEO β€” Answer Engine Optimization). The method applied to each provider followed three checks, in order:

  1. Declared academic affiliation β€” do the authors of any cited papers or research have a commercial role in the company citing them as evidence?
  2. Commercial relationship with cited sources β€” when a provider cites "independent research" to support its claims, does that source sell similar or related services?
  3. Citation chain β€” do the cited sources cite each other in a closed loop, creating an illusion of independent consensus where it's actually the same small group of commercial actors?

This check allowed us to classify publicly available evidence into three tiers, applying the same criterion to every claim gathered through systematic research (STORM method, cross-checked against OpenAlex for the real academic affiliation of cited authors):

TierDefinitionTreatment
Tier 1Convergent across sources with opposing incentives (e.g. crawlability, consistency of structured data, correlation with Google ranking as a measurable proxy)Usable as a basis for technical recommendations
Tier 2Plausible hypothesis but supported by a single source, often commercialTo be tested with own data before adoption
Tier 3Actively contested among competing vendors, or lacking sources independent from the vendorExcluded until own evidence exists

Most specific optimization promises "to make your site appear in AI answers" fall into Tier 2 or Tier 3: plausible, but not independently verified from the vendor proposing them. Few things β€” structured data consistency, basic technical crawlability β€” fall into Tier 1, meaning they have convergent evidence from sources with opposing interests.

This classification motivated the question behind the pilot described below: if even the Tier 1 levers (structured data, technical readability) were implemented perfectly, is there observable evidence that this changes how often an AI agent recommends a site?

The pilot: method

We built a Python pipeline of four sequential stages, applied to two distinct product categories (running shoes and wine) to check whether observations replicated across different domains.

Stage 1 β€” Collecting recommendations

An AI agent (gemini-flash-lite-latest, queried in static-memory mode β€” no real-time web search) received two types of queries, each repeated 20 times to capture the model's variance:

Declared methodological note: the model was selected after verifying, via a diagnostic script, that the API key in use is associated with a "canary/beta" environment with access to non-standard models. The -latest alias does not correspond to a version with a fixed identifier: the underlying model can change over time at the provider's discretion, independent of any pipeline changes. This limits the exact reproducibility of the experiment over time.

Stage 2 β€” Structured extraction

The same model (gemini-flash-lite-latest), in a distinct role as extractor (LLM-as-a-judge), converted the free text of the responses into structured data: first vendor mentioned, full list of vendors mentioned per repetition.

Stage 3 β€” Technical measurement

Each identified vendor was scanned with Agentabile's Evaluator, which measures structured data (schema.org/JSON-LD), presence of a sitemap and robots.txt with rules for known AI crawlers, and compatibility with agentic purchase protocols (UCP/ACP) where relevant.

Stage 4 β€” Manual verification and cleanup

This stage, not planned initially at this level of detail, turned out to be the most decisive for the quality of the final dataset β€” and is described separately in the next section, because it is itself a finding.

An unexpected result: how hard it is to technically measure a site from the outside

During data collection, three categories of problems produced falsely low or outright incorrect technical scores, each individually verified before being corrected or excluded from the final dataset:

ProblemVerified causeExampleTreatment
Intermittent network timeout, cause not clearly distinguishable from anti-bot blocking without deeper diagnosis (4 of 6 attempts failed due to timeout, the 2 successful ones produced scores consistent with each other)The site times out against the automated fetcherAmazon.itExcluded from quantitative measurement, reported only in the descriptive track
Client-side rendering (SPA)Valid HTTP 200 response, but near-empty initial HTML (content built by JavaScript after load)Zalando.it (1 page fetched); Les Caves de Pyrene (14 characters of extracted text across the entire homepage)Excluded if not measurable; included as a legitimate zero score if verified with concrete evidence
Wrong name→URL resolutionThe automated mechanism linking the vendor name cited by the LLM to its website produced wrong matchesA name cited as "Pavin" incorrectly resolved to a dental clinic's websiteManually verified one by one; excluded if ambiguous between multiple non-equivalent candidates

Of 12 vendors initially identified in the wine "merit" track, 2 were excluded due to ambiguity in URL resolution β€” not variants of the same site, but short business names colliding with companies entirely unrelated to the wine sector. The final verified dataset counts N=10 measurable vendors for wine.

This cleanup process is, in our view, the most generalizable finding of the pilot: automated third-party technical auditing has real structural limits, distinct from each other (active defense vs. architecture not designed to be read without a browser), which require case-by-case verification β€” they cannot be assumed solved by a tool, however well designed.

Results

Running shoes domain

Opacity track (neutral query, 20 repetitions):

VendorFirst mentionTotal mentions
Amazon1120
Nike514
Maxi Sport217
Decathlon114

Merit track (10 queries with explicit exclusion of giants): the agent revealed a long tail of specialized shops (Maxi Sport, Top4Running, DF Sport Specialist, Koala Sport, Sportler and others), with the same qualitative pattern observed in wine β€” no evident relationship between mention frequency and vendors' technical scores. We do not present a numeric correlation table for this domain here: unlike wine, where every single vendor was individually verified (previous section), for shoes only two anomalous cases were diagnosed individually. Given the number of URL-resolution errors and skewed scores discovered in wine despite thorough verification, we do not feel confident presenting the shoes table with the same degree of reliability β€” we prefer to omit it rather than risk publishing data not verified to the same rigor.

Wine domain

Opacity track (neutral query, 20 repetitions):

VendorFirst mentionTotal mentions
Tannico220
Callmewine120
Vino7566
Vino.com55
Amazon44

Merit track (10 queries with explicit exclusion) β€” correlation with technical score, N=10 verified measurable vendors:

VendorMentionsTechnical score (Visibility)
Callmewine16929.0
XtraWine8155.0
Tannico6055.0
SignorVino6029.0
Enoteca Properzio4455.0
Les Caves de Pyrene400.0
Vino753214.0
Vinix2463.0
Enoteca Naturale2057.0
Svinando929.0

In wine, the two highest technical scores overall (Vinix 63.0, Enoteca Naturale 57.0) belong to the two least-mentioned vendors in the table.

Statistical analysis and its limits

On the wine dataset (the only one with sufficient N for a model, albeit at the lower bound of the minimum threshold we had set) we applied a Bayesian Poisson regression, with a robustness check we recommend as standard practice for anyone replicating this kind of analysis on small samples:

ConfigurationCoefficient (beta)94% credible intervalIncludes zero
Full dataset (N=10)-0.056[-0.124, 0.011]Yes
Without the most-mentioned vendor (N=9)+0.080[-0.006, 0.167]Yes

In both configurations, the interval includes zero: neither estimate provides statistical evidence of correlation, positive or negative. But the sign of the coefficient flips when removing a single vendor from the sample β€” the clearest evidence that, at N=10, the point estimate is not stable and should not be presented as a reliable number on its own.

An additional check (posterior predictive check) confirms the same caution: the model fitted on the full dataset drastically underestimates the real variability observed in the data (a case of overdispersion caused by the disproportionate weight of the most-mentioned vendor as a single extreme observation).

What this pilot honestly concludes

In this exploratory sample β€” two domains, a single AI model (in an atypical, declared configuration), N=10 in the domain with statistical analysis β€” no coherent pattern emerges between a site's technical readiness and how often an AI agent recommends it. The most-cited vendors are not systematically those with the best technical score.

This does not amount to saying technical merit is irrelevant in general: the sample is small, only one model was tested (and in static-memory mode, without real-time web search β€” a different mechanism from how assistants like ChatGPT or Perplexity operate when browsing the web), and a plausible confounding factor (a company's overall maturity produces both a better site and more recognition, without either causing the other) was not isolated in this experimental design.

What the pilot shows more solidly is a different, perhaps more useful fact: today there is no reliable, independent way to verify on what basis an AI agent chooses who to recommend β€” not for an external researcher, and, as we experienced firsthand, not even in a fully automated way for those building the measurement tools themselves.

Next steps

A larger sample (ideally 30-50+ measurable vendors per category, across more product categories, with more AI models at fixed versions, and a scanner capable of real JavaScript rendering to reduce the number of sites "invisible" for architectural reasons) would be needed before drawing robust quantitative conclusions. This work should be understood as method validation β€” a starting point, not an endpoint.


Sources

  1. Google Cloud, Announcing Agent Payments Protocol (AP2), September 16, 2025 β€” official announcement of the AP2 protocol and launch partners (Mastercard, PayPal, American Express, Coinbase, Etsy, and over 60 others). https://cloud.google.com/blog/products/ai-machine-learning/announcing-agents-to-payments-ap2-protocol
  2. AP2 technical specification and reference repository (Apache 2.0). https://ap2-protocol.org
  3. Regulation (EU) 2022/2065 of the European Parliament and of the Council of 19 October 2022 on a Single Market For Digital Services (Digital Services Act), OJ L 277, 27.10.2022, p. 1–102, Article 27 (Recommender system transparency). Official consolidated text, EUR-Lex, CELEX identifier 32022R2065. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=uriserv:OJ.L_.2022.277.01.0001.01.ENG
  4. Agentabile data: scan results cited in the tables collected between August 16 and 19, 2026, raw dataset available on request.

Methodology, raw data, and pipeline code available on request. This article deliberately avoids terms like "demonstrates" or "confirms" in relation to quantitative results, consistent with the exploratory nature and small sample size of the pilot described.