Vet an AI Search Agency: 100-Point RFP Scorecard
Choosing an AI search agency is a procurement decision, not a screenshot competition. Whether you are buying Answer Engine Optimisation (AEO), Generative Engine Optimisation (GEO), or what AEO and GEO actually involve as a combined programme, the principle holds: select on evidence, controllable work and commercial fit, not on a screenshot showing that a brand appeared in one ChatGPT answer.
This guide is written for CMOs, heads of search, growth leaders and procurement teams who need a defensible way to compare bidders and avoid paying for activity that never moves the business.
TL;DR: Pick an AI search agency on evidence, controllable work and commercial fit using an eight-part RFP and a weighted 100-point scorecard, not on one screenshot of a brand appearing in a ChatGPT answer.
Key Takeaways
- Selection should rest on evidence, controllable work and commercial fit, never on a screenshot showing a brand appeared in one ChatGPT answer.
- A defensible RFP has eight parts: fixed scope, a repeatable baseline, raw evidence, separation of monitoring/analysis/implementation, a prioritised work plan, measurement that splits presence from citation from impact, governance, and a bounded pilot.
- Score bidders out of 100 across eight weighted categories: baseline design (20), execution capability (20), diagnosis and prioritisation (15), measurement and commercial linkage (15), evidence and transparency (10), governance and risk (10), team and operating fit (5), commercial clarity (5).
- Disqualify any agency that guarantees placement or citations in independent AI systems, hides its prompts and sampling method, or proposes fake reviews, undisclosed paid endorsements or spam links.
- No public source provides a universal AI-visibility score, guaranteed citation method or fair retainer for every organisation, so validate current features and contract terms during procurement.
What should an AI search agency RFP actually ask for?
A defensible RFP has eight parts, and a bidder who cannot satisfy them is not ready for your budget. Require each one explicitly:

- Fixed market, audience, country, product, and buyer intent definitions
- Repeatable baseline across named search and AI surfaces
- Raw evidence behind mentions and citations
- Clear separation between monitoring, analysis, and implementation
- Prioritised work plan covering technical SEO, content, entity evidence, and off-site corroboration
- Measurement distinguishing presence, citation, and commercial impact
- Governance for claims, access, data, approvals, and subcontractors
- Bounded pilot with acceptance criteria and exit route
Why do so many AI search proposals mislead buyers?
Most weak proposals quietly swap one thing for another, and naming the swaps is the fastest way to see through them. Guard against three substitutions:
- Presence for authority: brand mentions do not equal citations.
- Citation for demand: cited pages may generate no qualified visits.
- Monitoring for improvement: dashboards reveal gaps without closing them.
Do you need to monitor, improve, or prove?
Decide which of the three jobs you are buying before you read a single proposal, because each has a different output and a different owner. Map your need against this scope:
| Need | Output | Typical Owner |
|---|---|---|
| Monitor | Repeated observations across defined prompt and platform panel | Tool, analyst, or agency |
| Improve | Shipped technical fixes, stronger pages, entity clarification | Agency plus internal specialists |
| Prove | Change logs, raw observations, first-party data, qualified outcomes | Shared agency/client measurement team |
What information should every bidder receive before quoting?
Give every bidder the same written brief, or their quotes will be guesses dressed as plans. A complete pre-RFP brief includes:
- Countries, languages, and customer segments in scope
- Products or services generating qualified revenue
- Priority commercial, comparison, problem, and brand prompts
- Known regulated or high-risk claims
- CMS, analytics, Search Console, Bing Webmaster Tools, and CRM environment details
- Existing SEO, content, PR, and development partners
- Pages and markets that cannot be changed
- Approval owners and expected review times
- Internal engineering, editorial, and subject-matter capacity
- Current organic-search performance and technical constraints
- Business outcome the programme should influence
- Budget shape, pilot window, and procurement requirements
What evidence should an AI search agency be able to show?
An agency worth hiring can hand you an auditable evidence pack, anonymised or permissioned, rather than a proprietary score. Insist on all seven elements:

- Baseline definition with engines, interfaces, dates, locations, logged-in state, prompts, repetitions, and sampling rules
- Raw observations (saved answers or exports with citations)
- Intervention log showing URLs changed, technical tickets, evidence created, outreach performed, and dates
- Control boundary clarifying agency vs. client vs. external changes
- Outcome chain linking presence, citation, first-party search data, referral behaviour, and qualified pipeline
- Uncertainty documentation including volatility, missing data, and attribution limits
- Decision record demonstrating actionable changes resulting from evidence
How do you score competing bidders objectively?
Score every bidder out of 100 across eight weighted categories, so commercial polish never outweighs delivery capability. Use this scorecard:
| Category | Weight | Full-Mark Criteria |
|---|---|---|
| Baseline and research design | 20 | Fixed scope, representative buyer intents, repeat runs, platform/location controls, raw evidence, explicit limitations |
| Diagnosis and prioritisation | 15 | Separates technical, content, entity, off-site, conversion causes; ranks work by business value and evidence |
| Execution capability | 20 | Named people shipping technical fixes, expert-led content, structured data, evidence assets, legitimate authority work |
| Measurement and commercial linkage | 15 | Separates presence, citations, engagement, qualified outcomes; uses first-party data |
| Evidence and transparency | 10 | Auditable examples, intervention logs, methodology, uncertainty, no unverifiable guarantees |
| Governance and risk | 10 | Clear access, privacy, claim review, approvals, subcontracting, ownership, security, escalation controls |
| Team and operating fit | 5 | Named accountable lead, realistic client dependencies, cadence, collaboration model |
| Commercial clarity | 5 | Monitoring, implementation, pass-through costs, assumptions, change control, exit terms separated |
Which bidders should be disqualified on sight?
Some behaviours end the conversation, whatever the rest of the proposal claims. Mark a bidder ineligible if they:
- Guarantee placement, citations, or recommendations in independent AI systems
- Refuse disclosing prompts and sampling methodology
- Propose fabricated reviews, undisclosed paid endorsements, spam links, or mass query-variant pages
- Cannot identify system access owners for sensitive analytics, CRM, or publishing systems
- Use client data for unrelated model training without acceptable terms
- Cannot separate their work from platform volatility
- Make regulated claims without competent review processes
What questions expose a weak AI search agency fast?
Six questions separate operators from resellers, because each rewards method and punishes hand-waving. Put them to every shortlisted bidder:
"Show us one prompt from baseline to decision." Strong responses walk through raw answers, source URLs, repeated runs, diagnosis, intervention, rechecks, and business interpretation. Weak responses redirect to proprietary scores.
"What would make you recommend doing nothing?" Credible answers cite noise, weak commercial relevance, no credible intervention, insufficient sample, or stronger competing priorities. Agencies that always recommend monthly activity are describing a billing model, not a decision model.
"Which deliverables do not improve visibility by themselves?" Legitimate answers admit that monitoring, crawler access rules, schema, and dashboards are useful but do not guarantee recommendations.

"How do you distinguish a mention, citation and recommendation?" Require a written taxonomy. Linked sources, unlinked brand mentions, and ranked shortlist positions must not be combined as identical wins.
"How will you use our first-party platform data?" Agencies should understand what Google and Bing first-party observations can and cannot prove, preserve source data, and avoid blending incompatible metrics.
"Who performs the work?" Ask for named roles, allocation percentages, and subcontractor identification.
What does a good AI search pilot look like?
A useful pilot tests method and working fit, not just output, and it accepts on three fronts.
Delivery acceptance:
- Agreed technical and editorial changes shipped
- Evidence and approvals recorded
- Raw observations available
- Tracking and first-party data connected
- No unresolved high-risk claim or access issues
Learning acceptance:
- Team can explain which sources and page types engines selected
- Findings produce a ranked next-action list
- Variability and missing data are visible
- Buyer can reproduce the reporting logic
Outcome observation:

- Changes in search eligibility, answer presence, or citations are recorded
- Referral and conversion behaviour reported where reliable
- No causal claims exceed the test design
What contract controls protect the buyer?
The statement of work, not the pitch deck, is where risk is actually managed. Define each of these controls explicitly:
- Content, research, prompt sets, dashboards, and raw export ownership
- Access levels and least-privilege expectations
- Personal, customer, and confidential data treatment
- Approval rules for claims, pages, outreach, and public statements
- Subcontractor or AI tool data processing disclosure
- Permitted and prohibited link, review, and mention-acquisition methods
- Change control procedures for platform, interface, or coverage modifications
- Historic baseline preservation when prompts evolve
- Pass-through tool, media, PR, and production costs
- Incident, correction, and escalation procedures
- Termination, export, and credential-revocation steps
- Prohibition against representing observations as guaranteed rankings or causal proof
What are the biggest red flags in an AI search proposal?
Ten patterns should trigger rejection or hard investigation. Watch for:
- The magic file: one crawler or
llms.txtchange sold as strategy - The schema fantasy: markup proposed without checking visible facts
- The single-answer case study: one favourable output treated as stable ranking
- The platform soup: Google AI features, ChatGPT, Perplexity, Copilot observations combined without explanation
- The content quota: success defined as articles produced rather than market evidence created
- The citation guarantee: agency promises outputs controlled by external models
- The anonymous authority plan: paid listicle placements, fake reviews, or undisclosed endorsements replacing genuine expertise
- The dashboard retainer: most budget observing problems while implementation stays out of scope
- The attribution miracle: every direct or branded visit credited to AI search
- The dependency omission: engineering, expert review, and PR assumed but unresourced
How should you make the final decision?
Run the shortlisted bidders through a live page and one buyer prompt, then compare their diagnosis with the SAGEO competitor-analysis method. The strongest team distinguishes evidence from hypothesis and prioritises the smallest defensible intervention.
Strong proposals are specific enough to audit and modest enough to trust. They document sample boundaries, agency-controllable work, client dependencies that might block delivery, and evidence that survives after the presentation materials expire. They do not predict every model change; they show operating methods that survive change through sound technical SEO, useful expert-led information, unambiguous entities, visible evidence, legitimate third-party corroboration, repeatable observation, and commercial judgement. This is why AI-search visibility matters enough to buy properly rather than quickly.
What sources back this guidance?
This guide draws on published platform and market documentation, used directionally rather than as universal benchmarks:

- Google's official guidance on generative AI features and on evaluating third-party AEO/GEO advice
- Microsoft's Bing Webmaster Tools documentation on citations, cited pages, grounding queries, and observational limitations
- Microsoft's official explanation of the Citation Share metric as observational, not a ranking system or a traffic-share measure
- OpenAI's crawler documentation regarding platform-specific controls
- Commercial RFP template sources
- Vendor-published market pricing ranges (used directionally, not as universal benchmarks)
One caveat to keep in view: no public source provides a universal AI-visibility score, guaranteed citation method or fair retainer for every organisation. Platform interfaces, coverage and reports change, and commercial pricing pages reflect vendors' own offers. Validate current features and contract terms during procurement.
Next step
Before you issue the RFP, run the SAGEO 50-point audit to establish which technical, content, entity, and measurement gaps genuinely require agency help. If your shortlist is impossible to compare, start with the SAGEO framework and commission a decision-grade baseline before committing to any long AI-search retainer.
© 2026 SAGEO. Search, Answer, and Generative Engine Optimisation.
Sources
- https://developers.google.com/search/docs/fundamentals/ai-optimization-guide
- https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview
- https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare
- https://developers.openai.com/api/docs/bots
- https://higoodie.com/blog/ai-search-rfp-template/
- https://humanswith.ai/blog/what-aeo-and-geo-actually-cost-in-2026/
- https://thedigitalelevator.com/blog/aeo-and-geo-pricing-guide/