b084aeff6e
Sitemap filter was excluding cookies/ and privacidad/ but not aviso-legal/, even though that page is marked noindex,follow -- a noindexed page has no business being listed for crawl discovery. Also adds docs/seo-audit/mnqcatering.com-audit/ -- the full multi-agent SEO audit run this session against the corrected live domain (www.mnqeventos.es), including the FULL-AUDIT-REPORT.md, ACTION-PLAN.md, and per-category findings files.
162 lines
19 KiB
Markdown
162 lines
19 KiB
Markdown
# GEO / AI Search Readiness Audit — mnqeventos.es (corrected)
|
||
|
||
**Audit date:** 2026-09-15
|
||
**Live target audited:** https://www.mnqeventos.es
|
||
**Method:** Live DNS resolution, direct HTTP fetch (curl), SSR-aware rendering via `render_page.py --mode auto`, repository source inspection for the pending (undeployed) fix.
|
||
|
||
> **Correction notice:** An earlier pass of this audit targeted `mnqcatering.com`, which is not a registered domain (WHOIS: no match; DNS: NXDOMAIN) — that was a mistake in the audit target, not a finding about the real site, and produced a meaningless 0/100 score. The real production domain is **`https://www.mnqeventos.es`**, confirmed live and resolving (`169.58.117.246`, HTTP/2 200 on `/`, `/robots.txt`, `/llms.txt`). This document replaces that earlier pass entirely.
|
||
|
||
## Summary
|
||
|
||
The live site is technically reachable and reasonably crawlable (SSR HTML, permissive `robots.txt`, working `llms.txt`), but it is **actively telling every crawler — including every AI crawler in scope — the wrong canonical domain**. The canonical tag, Open Graph tags, JSON-LD entity `url`, the `robots.txt` `Sitemap:` directive, and the visible footer link all point to `mnqcatering.com`, a domain that is unregistered and does not resolve. This is a live, current-production defect (not a legacy leftover in docs) with direct GEO consequences: AI crawlers that respect canonicalization will attribute this content to a dead URL, entity graphs built from the `CateringService` JSON-LD will resolve to a broken `url` field, and `og:image`/`og:url` previews (used by some AI answer engines for citation cards) will fail to load.
|
||
|
||
The good news: the root cause is already fixed in the repository (commit `a6692b7 fix(seo): correct canonical domain to www.mnqeventos.es`, plus `astro.config.mjs`, `MainLayout.astro`, `SiteFooter.astro`, `contacto.astro`, `public/robots.txt`, `public/llms.txt` all now reference `mnqeventos.es`) — it just **has not been deployed to production yet**. The live `llms.txt` served today is also a stale, thinner version missing the address and area-served lines that exist in the repo's current `llms.txt`. Once deployed, several of the findings below resolve automatically; I've flagged which ones.
|
||
|
||
## Findings
|
||
|
||
### 1. Canonical URL, Open Graph, and JSON-LD entity `url` all point to a dead domain — HIGH (live defect, fix pending deploy)
|
||
- **Severity:** High
|
||
- **Evidence (live, fetched 2026-09-15):**
|
||
- `<link rel="canonical" href="https://www.mnqcatering.com/">` on `/`, and `https://www.mnqcatering.com/contacto`, `https://www.mnqcatering.com/nosotros` on those pages respectively.
|
||
- `<meta property="og:url" content="https://www.mnqcatering.com/">`, `<meta property="og:image" content="https://www.mnqcatering.com/images/foto-bodas.jpg">` — the image URL will fail to load for any consumer (AI citation card, social preview) since `mnqcatering.com` does not resolve.
|
||
- JSON-LD `CateringService` schema on `/`, `/contacto`, `/nosotros`: `"url":"https://www.mnqcatering.com"`.
|
||
- Footer visible link: `<a href="https://www.mnqcatering.com">www.mnqcatering.com</a>` on every page — a dead outbound link on a live site.
|
||
- `umami` analytics script tag: `data-domains="www.mnqcatering.com"` (low GEO impact, but confirms the whole template still ships the old domain).
|
||
- **Impact:** AI/search crawlers that respect `rel=canonical` will index/attribute content to a non-existent URL, which can suppress citation entirely or cause AI engines to silently drop the page from consideration when the canonical target 404s/fails to resolve. Structured-data entity resolution (how Google/Bing/AI systems build a knowledge-graph node for "MNQ Catering y Evento") is anchored to a broken `url` field.
|
||
- **Status:** Already fixed in repo (`astro.config.mjs: site: 'https://www.mnqeventos.es'`, `MainLayout.astro` line 33/98, `SiteFooter.astro` line 22, `contacto.astro` line 108) but **not yet deployed**.
|
||
- **Recommendation:** Deploy the pending fix immediately. This is the single highest-impact, already-solved change — it just needs to ship. Effort: none (code complete), deploy only.
|
||
|
||
### 2. `robots.txt` Sitemap directive points to a domain that does not exist — HIGH (live defect, fix pending deploy)
|
||
- **Severity:** High
|
||
- **Evidence (live):**
|
||
```
|
||
User-agent: *
|
||
Allow: /
|
||
|
||
Sitemap: https://www.mnqcatering.com/sitemap-index.xml
|
||
```
|
||
Verified: `curl -o /dev/null -w '%{http_code}' https://www.mnqcatering.com/sitemap-index.xml` → connection failure / `000` (domain unregistered). The correct, working sitemap is live and reachable at `https://www.mnqeventos.es/sitemap-index.xml` → HTTP 200, but nothing in the live `robots.txt` points there.
|
||
- **Impact:** Crawlers that discover the sitemap only via `robots.txt` (a common pattern for AI/search crawlers doing full-site discovery) get a dead pointer and may fail to enumerate all indexable pages, slowing or preventing discovery of `/bodas-reales`, `/corporativo-business`, `/reuniones-familiares`, etc.
|
||
- **Status:** Already fixed in repo `public/robots.txt` (now points to `https://www.mnqeventos.es/sitemap-index.xml`) — pending deploy.
|
||
- **Recommendation:** Ships with fix #1's deploy. Effort: none, deploy only.
|
||
|
||
### 3. `robots.txt` is permissive for AI crawlers — GOOD (no action needed)
|
||
- **Severity:** Positive finding
|
||
- **Evidence:** `User-agent: * / Allow: /` — a single wildcard rule with no disallows. This implicitly allows GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, CCBot, anthropic-ai, and cohere-ai; nothing is blocked.
|
||
- **Recommendation:** No change needed. Optionally add explicit named-agent blocks for CCBot/anthropic-ai/cohere-ai only if the business wants to opt out of training-only crawling while keeping AI *search* visibility (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot) — this is optional per the audit brief, not a defect.
|
||
|
||
### 4. `llms.txt` is present and live, but the deployed version is thinner and has the wrong canonical domain than the repo's current version — MEDIUM (live defect, fix pending deploy)
|
||
- **Severity:** Medium
|
||
- **Evidence — live, fetched 2026-09-15 (`https://www.mnqeventos.es/llms.txt`, HTTP 200):**
|
||
```
|
||
MNQ Catering y Evento
|
||
|
||
- Canonical domain: https://www.mnqcatering.com
|
||
- Business type: Premium catering service
|
||
- Services: Bodas, eventos corporativos, reuniones familiares
|
||
- Primary contact channels: WhatsApp (+34 678 17 15 13), phone (+34 678 17 15 13)
|
||
- Key pages: /, /nosotros, /bodas-reales, /corporativo-business, /reuniones-familiares, /contacto
|
||
```
|
||
- **Repo's current `public/llms.txt` (committed, not yet deployed)** adds two lines the live version is missing:
|
||
```
|
||
- Address: Calle Camino Vivero, 6, 29014, Málaga, España
|
||
- Area served: Málaga (city) and Provincia de Málaga
|
||
```
|
||
and correctly says `Canonical domain: https://www.mnqeventos.es`.
|
||
- **Format note (applies to both live and repo versions):** Neither follows the llms.txt spec convention closely — no `# Title` H1, no `>` summary blockquote, no markdown-style `[Key Page](full-url): description` links (currently plain relative paths in a flat bullet list). This is a minor quality gap versus the informal but now-common llms.txt convention; it doesn't block parsing but is a missed opportunity for AI agents that expect the structured format.
|
||
- **Recommendation:** Deploy the repo's current version immediately (fixes domain + adds address/area-served — real local-entity signals). Separately, reformat to standard llms.txt structure (H1 title, one-line summary blockquote, markdown links per key page with a short description each) for better parseability by LLM-based agents. Effort: Low (content already exists, just needs reformatting).
|
||
|
||
### 5. No FAQ-style content on the homepage; direct-answer content is confined to `/contacto` — MEDIUM
|
||
- **Severity:** Medium
|
||
- **Evidence:** `/contacto` carries a genuine `FAQPage` JSON-LD block with 3 Q&A pairs (e.g. "¿Qué datos conviene enviar para pedir una propuesta?" → "Fecha estimada, lugar del evento, número aproximado de asistentes y el tipo de servicio que necesita.") and matching visible question-format subheadings ("Qué necesitamos para orientarle", "Cómo planteamos el presupuesto"). This is a genuine AEO positive — but it exists only on one page. The homepage, `/nosotros`, and the three service pages (not fetched this pass, but referenced identically in nav) use marketing prose headings ("Servicio planificado según ubicación, formato y ritmo del evento") rather than question-based H2/H3s, and no passage on the homepage is a self-contained 134–167-word answer block — most are 1–3 short sentences of brand copy.
|
||
- **Recommendation:** Extend the `FAQPage` + question-based-heading pattern already proven on `/contacto` to the service pages (bodas, corporativo, reuniones-familiares) and `/nosotros`, each with 3–5 self-contained Q&A pairs in the 40–80 word range (FAQ answers don't need to hit the full 134–167-word "article passage" target, but should stay self-contained and specific). Effort: Medium — content writing + schema markup, reusing the existing `/contacto` pattern as a template.
|
||
|
||
### 6. No brand/entity linking signals (no `sameAs`, no social profile links anywhere on-site) — MEDIUM-HIGH
|
||
- **Severity:** Medium-High
|
||
- **Evidence:** Checked `/`, `/contacto`, `/nosotros` for LinkedIn, Instagram, Facebook, X/Twitter, YouTube, TikTok links — zero matches. The `CateringService` JSON-LD has no `sameAs` array. No mentions of Wikipedia, Reddit, or YouTube presence found on-site (external presence not independently verified this pass — no search/social-listening tool was used).
|
||
- **Impact:** Per the brand-mention correlation data, YouTube mentions (~0.737) and Reddit presence are the strongest predictors of AI citation; Wikipedia entity presence is also high-value. Zero on-site `sameAs`/social signals means AI systems building an entity profile for "MNQ Catering y Evento" have nothing to cross-reference beyond the domain itself.
|
||
- **Recommendation:** (a) Add a `sameAs` array to the `CateringService` JSON-LD linking to whatever active social profiles exist (Instagram/Facebook are typical for catering businesses even without LinkedIn/YouTube). (b) If no social profiles exist yet, creating even a modest Instagram/YouTube presence (venue photos, event highlight reels) would meaningfully move the needle given YouTube's outsized correlation with AI citation. Effort: Low for `sameAs` markup (if profiles exist), Medium-High for building new social presence from scratch.
|
||
|
||
### 7. No address / local-entity structured data live (fix pending deploy) — MEDIUM
|
||
- **Severity:** Medium
|
||
- **Evidence:** Live JSON-LD `CateringService` on `/`, `/contacto`, `/nosotros` has no `address`, `areaServed`, or `geo` fields — only `name`, `url` (wrong domain), `description`, `telephone`, `serviceType`, `contactPoint`. No address text appears anywhere in the visible page copy either. The repo's current `llms.txt` (undeployed) already has the address/area-served text; it's unclear from this pass whether the `CateringService` JSON-LD itself is being updated with structured `address`/`areaServed`/`geo` fields, or whether the fix is llms.txt-only.
|
||
- **Recommendation:** Once the domain fix deploys, separately confirm (or add) structured `address` and `areaServed` fields inside the `CateringService` JSON-LD itself (not just llms.txt prose) — this is what powers Google's local entity panel and is a stronger machine-readable signal than plain text. Effort: Low, since the address text is already sourced (`Calle Camino Vivero, 6, 29014, Málaga, España`).
|
||
|
||
### 8. Server-side rendering confirmed — GOOD (no action needed)
|
||
- **Severity:** Positive finding
|
||
- **Evidence:** `render_page.py --mode auto` on `https://www.mnqeventos.es/` returned `"is_spa": false`, `"mode_used": "raw"` — the raw HTTP response already contains the full rendered page (hero copy, service cards, testimonials, footer, all JSON-LD) with no client-side hydration gap. This matches the Astro `output: 'server'` architecture confirmed in `astro.config.mjs`.
|
||
- **Recommendation:** No change needed. This is the single strongest technical-accessibility asset the site has — every AI crawler sees the same content a browser does, with zero JS-execution risk.
|
||
|
||
### 9. Descriptive image alt text present — GOOD (partial credit)
|
||
- **Severity:** Positive finding, with a gap
|
||
- **Evidence:** Homepage images carry specific, descriptive `alt` text (e.g. `"MNQ Catering — salón de bodas premium"`, `"Vieiras con azafrán"`, `"Bartender de coctelería"`) rather than generic/empty alts. However, there is no video content anywhere observed, and no image captions or surrounding text that would let an AI engine extract a citable fact from an image (e.g. no recipe steps, no menu item descriptions tied to photos).
|
||
- **Recommendation:** Keep the alt-text discipline for new images. Consider short video content (event highlight reels) for the Multi-Modal dimension, given video/YouTube's strong correlation with AI citation.
|
||
|
||
## GEO Health Score
|
||
|
||
**Overall: 46 / 100**
|
||
|
||
| Dimension | Weight | Score (0–100) | Weighted | Rationale |
|
||
|---|---|---|---|---|
|
||
| Citability | 25% | 50 | 12.5 | Solid FAQ schema + Q&A headings on `/contacto`; homepage/other pages are marketing prose, not self-contained answer blocks |
|
||
| Structural Readability | 20% | 55 | 11.0 | Clean heading hierarchy; question-format H2/H3 only on `/contacto`, not site-wide |
|
||
| Multi-Modal Content | 15% | 35 | 5.25 | Good descriptive alt text; no video, no YouTube presence, no image-adjacent citable facts |
|
||
| Authority & Brand Signals | 20% | 30 | 6.0 | No `sameAs`/social links, no visible address/local-entity data live, wrong canonical/entity `url` currently live |
|
||
| Technical Accessibility | 20% | 55 | 11.0 | SSR confirmed, permissive robots.txt, fast HTTP/2 — undercut by broken canonical, broken sitemap pointer, broken og:image |
|
||
|
||
**If the already-committed domain fix (commits `a6692b7` and related) is deployed as-is**, Technical Accessibility and Authority & Brand Signals would both improve materially (est. +15–20 points combined toward Technical, +5–10 toward Authority once address data ships in structured form), pushing the overall score into the high-50s/low-60s without any new work — that deploy is by far the highest-leverage action available right now.
|
||
|
||
## AI Crawler Access Status (robots.txt, live)
|
||
|
||
| Crawler | Status |
|
||
|---|---|
|
||
| GPTBot | Allowed (wildcard `Allow: /`) |
|
||
| OAI-SearchBot | Allowed (wildcard `Allow: /`) |
|
||
| ClaudeBot | Allowed (wildcard `Allow: /`) |
|
||
| PerplexityBot | Allowed (wildcard `Allow: /`) |
|
||
| CCBot | Allowed (wildcard `Allow: /`) — no opt-out currently configured |
|
||
| anthropic-ai | Allowed (wildcard `Allow: /`) — no opt-out currently configured |
|
||
| cohere-ai | Allowed (wildcard `Allow: /`) — no opt-out currently configured |
|
||
|
||
Note: `robots.txt` itself is reachable, but its `Sitemap:` directive is broken (see Finding #2), which can hinder full-site discovery regardless of the permissive `Allow` rule.
|
||
|
||
## llms.txt Status
|
||
|
||
**Present, HTTP 200, but stale/incomplete relative to the repo's current (undeployed) version.**
|
||
- Live: 5 bullet lines, wrong canonical domain (`mnqcatering.com`), no address/area-served.
|
||
- Repo (pending deploy): 7 bullet lines, correct canonical domain, includes address and area served.
|
||
- Neither version follows the full llms.txt markdown convention (H1 + summary + linked key pages) — see Finding #4.
|
||
|
||
## Brand Mention Analysis
|
||
|
||
Not independently verified via external search/social-listening tools this pass (none available in this environment/session). On-site evidence only: **zero** `sameAs` links, **zero** social profile links (Instagram/Facebook/LinkedIn/YouTube/TikTok/X) found in the HTML of `/`, `/contacto`, `/nosotros`. This is a gap regardless of what external presence may exist, since AI entity-resolution systems weight *on-site* declared links (`sameAs`) heavily as a disambiguation signal. Recommend a follow-up pass with live search/social tools (or DataForSEO's `ai_opt_llm_ment_search`, if that MCP integration is enabled for this environment — it was not available this session) to check actual Wikipedia/Reddit/YouTube/LinkedIn mention volume before investing in new social presence.
|
||
|
||
## Platform-Specific Scores (qualitative estimate — no live scraping tool available this session)
|
||
|
||
| Platform | Estimated readiness | Basis |
|
||
|---|---|---|
|
||
| Google AI Overviews | Low-Medium | FAQPage schema present (favorable), but broken canonical/entity URL currently live undercuts trust signals |
|
||
| ChatGPT (browsing/search) | Low | No `llms.txt` linking convention, weak brand-mention surface, broken og:image would fail citation-card rendering |
|
||
| Perplexity | Low-Medium | Permissive robots.txt and clean SSR content favor crawlability; same entity/URL trust gap as above |
|
||
| Bing Copilot | Low-Medium | Same reasoning as Google AIO (shares underlying Bing index signals) |
|
||
|
||
These are directional estimates based on on-site signals only, not live AI-response testing. If DataForSEO's `ai_optimization_chat_gpt_scraper` becomes available in a future session, re-run for actual observed visibility rather than inferred readiness.
|
||
|
||
## Top 5 Highest-Impact Changes
|
||
|
||
1. **Deploy the already-committed domain fix to production** — Effort: None (code complete) / Impact: High. Fixes canonical tags, `og:url`/`og:image`, JSON-LD `url`, footer link, and the broken `robots.txt` sitemap pointer in one deploy. This is the single highest-leverage action and requires zero new development.
|
||
2. **Confirm the deployed `llms.txt` matches the repo's current version (with address + area-served) after deploy**, and reformat it to standard llms.txt markdown convention (H1, summary blockquote, linked key pages with descriptions) — Effort: Low.
|
||
3. **Add structured `address`/`areaServed`/`geo` fields to the `CateringService` JSON-LD** (not just llms.txt prose) on `/`, `/contacto`, `/nosotros` — Effort: Low, address text already sourced.
|
||
4. **Extend the `/contacto` FAQPage + question-heading pattern to the three service pages and `/nosotros`** — Effort: Medium, reuses an existing, working template.
|
||
5. **Add a `sameAs` array to the JSON-LD entity linking real social profiles** (or build a minimal video/YouTube presence if none exist) — Effort: Low if profiles exist, Medium-High if starting from zero; highest-correlation lever per the brand-mention data once other fixes ship.
|
||
|
||
## What Already Works Well
|
||
|
||
- **Genuine SSR, no hydration gap.** `render_page.py` confirms the raw HTTP response is the full page — every AI crawler sees complete content with zero JS-execution dependency.
|
||
- **Permissive `robots.txt`.** No AI crawler (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot) is blocked; a single wildcard `Allow: /` covers everything.
|
||
- **`llms.txt` exists and is live** (HTTP 200), and the repo's pending version is well-formed content-wise (business type, services, contact channels, key pages, address, area served) even if the markdown formatting could be tightened.
|
||
- **Real `FAQPage` schema + question-based subheadings on `/contacto`**, with concise, accurate, self-contained answers — this is exactly the AEO pattern that should be replicated site-wide.
|
||
- **Descriptive, specific image alt text** throughout the homepage gallery and service cards.
|
||
- **Fast, modern hosting**: HTTP/2, `alt-svc: h3`, clean response headers, no redirect chains observed on the pages checked.
|
||
- **The hard part (root-cause domain fix) is already done in code** — the team caught and fixed the canonical-domain issue proactively; it's purely a deployment gap away from resolving several of the findings above.
|