I built a website that wants to be quoted by AI assistants and, instead of guessing, I put three widely-recommended signals into practice: a robots.txt entry that lets GPT-style crawlers in, an llms.txt file that lists every page, and JSON-LD schema that describes the content. After testing each one, I found only one can fail without a warning, and the others don’t do what most people claim.
Why the three signals matter
Webmasters have been told that AI-focused crawlers need explicit permission, that a plain-text “llms.txt” index helps large language models (LLMs) locate relevant passages, and that JSON-LD structured data boosts the odds of being cited. The promise is simple: add these files, and your site will start showing up in AI-generated answers, driving traffic and brand visibility.
1. Allowing AI crawlers in robots.txt
The robots.txt line that mentions GPTBot (or any other AI bot) is only a request. If a content delivery network (CDN) or firewall blocks the bot, the request never reaches the site, and the crawler can’t read the file at all. The only reliable way to know whether the bot can access you is to check the HTTP status code for each major bot. A 403 (forbidden) or 503 (service unavailable) response means the rest of the plumbing is moot—no bot, no indexing, no citation.
2. The llms.txt file
An llms.txt file is just a markdown list of URLs, meant to give LLMs a clean map of a site so they don’t have to infer navigation from HTML alone. In practice, Google’s documentation says it ignores the file, and no major search engine has pledged to use it for citation purposes. Keeping it around costs almost nothing, and it can be handy for custom AI agents, but it should not be advertised as a shortcut to more AI citations.
3. JSON-LD schema
JSON-LD embeds structured data—author names, publication dates, product IDs—directly in a page. Many marketers assume that adding schema will make their pages appear more often in AI answers, but a study of 1,885 pages found no measurable lift in citations from schema markup alone. The real value of JSON-LD lies in entity resolution: it tells a machine that “Jane Doe” on one page is the same “Jane Doe” on another, preventing confusion between similarly named people or products.
The real bottleneck: indexing
All three signals are plumbing; they only work if the page is indexed in the first place. A page that never appears in an index can’t be retrieved, no matter how many crawler allowances or schema tags you add. The most reliable indicator of indexing health is the coverage report in Google Search Console (or the equivalent tool for other engines). If a page shows up as “not indexed,” you need to fix that before worrying about any AI-specific signals.
A practical checklist
- Verify crawler access – Test the HTTP response for GPTBot, ClaudeBot, or any other AI crawler you care about. This step can fail silently; a block won’t generate a warning in your analytics.
- Confirm index coverage – Use Search Console to see which pages are indexed, which are excluded, and why. Resolve any “crawl errors” or “noindex” directives first.
- Add the plumbing – Once access and indexing are solid, drop in llms.txt and JSON-LD. Treat them as low-maintenance helpers rather than citation guarantees.
Counter-point
Some vendors still market llms.txt generators and schema plugins as SEO boosters for AI. The data I collected, plus the official stance from major search engines, suggests those claims are overstated. The tools aren’t harmful, but they shouldn’t be the centerpiece of an AI-citation strategy.
What to watch next
Monitor crawler logs, keep an eye on Search Console alerts, and periodically test your JSON-LD with validation tools to keep the pipeline flowing.
Takeaway: If you want your site quoted by AI, stop treating llms.txt and schema as magic tickets. Make sure the bots can actually reach you and that your pages are indexed; only then will the extra files serve their modest, supportive role.
Source: the original post on dev.to
