How Watermelon Ghost measures AI visibility for luxury hospitality — the full technical methodology behind every GEO Visibility Audit. Document version 2.0 · June 2026.
The GEO Machine answers a single question: when a high-intent traveler asks an AI platform for a hotel recommendation, does your brand surface, and in what light?
As travelers increasingly use ChatGPT, Claude, Perplexity, and Gemini instead of Google for hotel discovery, visibility in these platforms directly impacts booking decisions. Traditional SEO audits measure search engine rankings — a GEO audit measures presence in the conversational AI responses where travelers ask natural-language questions about where to stay.
Each audit runs 60 standardized queries across 4 AI platforms, producing 240 total responses. Brand and competitor mentions are detected with word-boundary precision. Sentiment is classified. Competitor displacement is measured. The output is a scored audit with a prioritized 90-day action plan grounded in specific query-level findings.
Travelers don't ask one kind of question. They move through four distinct decision moments — each with different intent, different query language, and different weight in how AI platforms form recommendations.
| Moment | Weight | Queries | Definition |
|---|---|---|---|
| Recommendation | 45% | 27 generic | How AI positions the brand when users ask for category-level recommendations |
| Discovery | 20% | 6 generic | How AI surfaces the brand in exploratory, category-browsing searches |
| Comparison | 20% | 3 branded + 12 generic | How the brand appears when travelers compare specific options |
| Trust | 15% | 9 branded + 3 generic | Signals of authority, credibility, and safety in AI responses |
Recommendation carries the highest weight (45%) because it represents the moment where booking decisions are most directly influenced — the traveler is actively choosing, not just browsing. Weights are fixed industry constants derived from hospitality-sector travel intent modeling, ensuring consistent cross-brand comparison.
| Query Type | Count | Share | What It Measures |
|---|---|---|---|
| Branded | 12 | 20% | What AI platforms say about you — sentiment, recommendation strength, accuracy |
| Generic | 48 | 80% | Whether AI platforms mention you at all — organic visibility without naming your brand |
Queries are not hand-written per brand. The system uses a library of 60+ query patterns with placeholders filled from each property's intake data. This ensures every brand audit is methodologically identical while using brand-relevant specifics.
| Placeholder | Source | Example |
|---|---|---|
| {property} | Brand name | Bardo Savannah |
| {region} | Geographic data | Savannah Historic District |
| {competitor} | Discovered competitors | The Dewberry |
| {persona} | Traveler archetypes | couples, solo traveler |
| {experience} | Amenity categories | spa, fine dining |
| {season} | Temporal context | spring, winter holidays |
| {occasion} | Trip purpose | anniversary, remote work retreat |
Query generation runs through Claude (Sonnet 4) with guardrails that enforce correct branded/unbranded distribution, prevent duplicate queries, and reject off-template output. All placeholders must be resolved — no unsubstituted templates reach execution.
Four AI platforms are queried, reflecting the fragmented landscape travelers actually use. Running all four prevents single-platform bias — a brand visible on ChatGPT but invisible on Perplexity has a real visibility problem.
| Platform | Model | Search Capability |
|---|---|---|
| ChatGPT | GPT-4o | Web search enabled |
| Claude | Claude (Anthropic) | Brave Search integration |
| Perplexity | Sonar | Search-native architecture |
| Gemini | Gemini (Google) | Google Search grounding |
For queries that include the brand name, scoring measures how the brand is presented, not just whether it appears. Claude classifies each response for sentiment:
| Sentiment | Score | Example |
|---|---|---|
| Positive | +1 | "Bardo Savannah is one of the finest boutique hotels in the South" |
| Neutral | 0 | "Bardo Savannah is located in the Historic District" |
| Negative | −1 | "Bardo Savannah has received mixed reviews for service" |
Branded scores are weighted and normalized to a 0–100 scale per decision moment.
For queries without any brand name, scoring measures pure presence: did the AI mention you at all? Mentioned = 1, Not mentioned = 0. Normalized to 0–100 as a percentage of generic queries where the brand appeared.
Each decision moment score combines its branded and generic components, then the overall GEO Score is the weighted sum:
GEO Score = (Recommendation × 0.45) + (Discovery × 0.20) + (Comparison × 0.20) + (Trust × 0.15)
| Tier | Range | Meaning |
|---|---|---|
| Leader | 70–100 | Dominant presence across AI platforms |
| Considered Option | 40–69 | Appears but not consistently positioned as a top choice |
| Afterthought | 10–39 | Mostly invisible except in direct name searches |
| Not Mentioned | 0–9 | Effectively no AI presence |
Scoring and gap analysis use deterministic algorithms with no AI component. This is deliberate: the numbers must be reproducible. Run the same responses through the scoring engine twice and you get the same result. AI is used where judgment is required — sentiment, classification, copy generation — not where arithmetic is required.
Claude analyzes every branded query response across all platforms for:
This granularity matters because "mentioned" isn't enough. A brand can be mentioned negatively, mentioned as an afterthought, or positioned as the definitive recommendation — and those outcomes demand radically different strategic responses.
Share of Voice measures competitive position in the generic query landscape — the queries where no brand is named and the AI chooses who to surface.
Methodology: Generic queries only. Branded queries are excluded to avoid inflating the metric — of course "Bardo Savannah" appears when you ask about "Bardo Savannah." Share of Voice is a measure of competitive visibility on a level playing field.
The system doesn't assume it knows your competitors. During the Discovery stage, it scans all AI responses for competitor names, surfaces the top 5 real competitors by mention frequency, and writes them back into the analysis. This means the competitive comparison reflects who the AI platforms actually name alongside you, not a static list.
Every competitor and brand mention is classified into one of four types:
| Mention Type | Meaning |
|---|---|
| Direct Recommendation | The AI explicitly recommends this property |
| List Inclusion | The property appears in a list of options |
| Contrast Mention | Mentioned in contrast to another property |
| Passing Reference | Mentioned in passing without endorsement |
Displacement occurs when a competitor is mentioned but your brand is either absent (not mentioned at all) or outranked (both mentioned, competitor positioned ahead). Displacements are aggregated into threat levels:
| Threat Level | Displacements | Strategic Meaning |
|---|---|---|
| Primary | ≥ 10 | Systematically displacing you across multiple territories |
| Significant | ≥ 5 | Regularly outranks you in specific contexts |
| Moderate | ≥ 2 | Appears ahead of you in specific query clusters |
| Minimal | 1 | Isolated displacement, likely query-specific |
When the brand receives zero mentions on a query, that query is a gap. Gaps are grouped by territory — topic domains that represent clusters of traveler intent (e.g., "luxury spa experiences," "historic district hotels," "beachfront dining").
Each territory is classified by competitor fill rate and displacement pattern:
| Gap Type | Fill Rate | Meaning |
|---|---|---|
| Unowned | < 30% | No brand owns this territory — open opportunity |
| Contested | 30–70% | Multiple brands compete for visibility |
| Owned by Competitor | ≥ 70% | A specific competitor dominates this territory |
| Authority Gap | Majority outranked | Your site exists but AI doesn't treat it as authoritative |
| Page Exists, Not Indexed | N/A | Content exists on your site but AI platforms don't reference it |
When a site inventory is available, gap analysis is cross-referenced against the brand's actual website content to determine whether a gap is a content gap (the site doesn't address this topic) or an indexing gap (the content exists but AI platforms don't surface it). This produces specific content action types:
A single audit is a snapshot. Two or more audits, run over time, reveal trajectory. The Process phase aggregates data from multiple collection runs and calculates:
The action plan translates analysis into a sequenced, scored 90-day execution plan with three tiers:
| Tier | Timeframe | Scope | Examples |
|---|---|---|---|
| Quick Wins | This week | Low-effort, immediate-impact fixes | Fix broken schema, claim unverified GBP listing, add missing page titles |
| Month 1–3 | 90 days | Strategic content and authority work | Create content for unowned territories, build backlinks, platform-specific optimization |
| Bonus | Opportunistic | Additional high-potential actions | Advanced schema strategies, competitive differentiation plays |
Actions are ranked by composite scoring of GEO Score Impact (how much this action could move the overall score), Commercial Impact (business value beyond the score), and Implementation Effort (resource requirements). The system enforces diversity across tiers, validates minimum requirements, and applies a hard stop if Tier 1+2 actions don't collectively close enough gaps to be needle-moving.
Before running the full 60-query budget, the system runs 2 canary queries across all platforms to detect brand detection regex mismatches, platform search capability issues, API authentication problems, and configuration errors. If preflight fails, the full run is blocked — preventing wasted API spend on a broken configuration.
Generated queries are validated against correct branded/unbranded counts per decision moment, minimum branded queries per moment, duplicate detection, and placeholder resolution. Queries that violate guardrails are regenerated.
Before the action plan is finalized, factual claims are extracted and presented for human review. Claims can be approved, rejected, or modified — ensuring the final output is editorially sound before it reaches the client.
| Component | Technology |
|---|---|
| Runtime | Bun (TypeScript) |
| Query Generation | Claude Sonnet 4 (Anthropic API) |
| Query Execution | ChatGPT, Claude, Perplexity, Gemini APIs |
| Sentiment Analysis | Claude (Anthropic API) |
| Mention Classification | Claude (Anthropic API) |
| Scoring & Gap Analysis | Pure computation (deterministic, no AI) |
| Competitor Discovery | Claude (Anthropic API) + regex |
| Site Inventory | Puppeteer (headless Chrome) + Claude |
| Report Generation | HTML → Chrome headless → PDF |
| Report Copy | Claude (Anthropic API), 2-stage generation + validation |