How to get your website cited by ChatGPT, Perplexity and Google AI Overviews
A practical checklist for making your business easy for AI assistants to find, understand and recommend, without abandoning the SEO basics that still matter.
Pro Indies Team · · 8 min read
More buyers now start with a question to an AI assistant instead of a search box. "Which agency builds Shopify stores in Jaipur?" or "What's a good tool for lead scoring?" gets a short answer with a handful of sources. If your business isn't one of them, you're invisible in that moment.
Getting cited by AI assistants rests on the same foundations as good SEO. It just rewards clarity and consistency even more. Below is the checklist we use for our own site and for clients, followed by exactly what we did on proindies.com, the mistakes that most often undo the work, and a 30-day plan you can follow.
1. Let the right crawlers in
AI products read the web through their own crawlers, and many sites block them by accident with an old robots.txt or an aggressive firewall rule. Check that you're not blocking the agents you want to be visible in, such as OAI-SearchBot and ChatGPT-User (ChatGPT), Claude-SearchBot and Claude-User (Claude), and PerplexityBot and Perplexity-User (Perplexity).
It helps to know there are two kinds of AI crawler. Search and user-triggered agents, like the ones above, fetch pages so an assistant can answer a question and cite you. Training crawlers, such as GPTBot, ClaudeBot and CCBot, collect content for building future models. You can allow one kind and not the other.
Google's AI Overviews are built on regular Google Search, so they depend on Googlebot. The separate Google-Extended token controls whether your content is used for Gemini models, not whether you appear in Search.
Decide deliberately which of these you allow, write it down in robots.txt, and make sure your CDN or security plugin isn't quietly overriding it.
2. Serve real HTML
Many AI crawlers fetch a page and read the HTML without running JavaScript. If your content only appears after a client-side app loads, some of them see an empty page.
Static generation or server-side rendering fixes this. As a quick test, view the page source (not the inspector) and confirm your headings, service descriptions and prices are actually in it.
3. Answer questions directly
AI assistants quote passages that answer a question cleanly. Structure important pages so they can be quoted:
- Use descriptive headings that match how people ask ("How long does a Shopify build take?").
- Put the answer in the first sentence under the heading, then add detail.
- Add a short FAQ to service pages covering cost, timelines, process and who it's for.
- Keep one topic per page, with a clear title and meta description.
This also happens to be what makes pages convert better for humans.
4. Add structured data
Schema.org markup tells machines exactly what a page is about: your organization, the services you offer, your location, articles and FAQs. Use Organization (or ProfessionalService), Service, Article and BreadcrumbList where they apply.
Google now shows FAQ rich results only for a small set of authoritative sites, but the markup still helps any system that reads it understand your content unambiguously. Give your organization a stable @id and reference it from every other block, so all your markup describes one entity.
5. Keep your facts consistent everywhere
Assistants cross-check. If your website says you're in Jaipur, your Google Business Profile says Sikar and a directory lists an old phone number, you look less trustworthy. Align your business name, locations, contact details, services and founding year across your site, Google Business Profile, social profiles and major directories.
6. Consider an llms.txt file
llms.txt is a proposed standard: a plain Markdown file at the root of your site that summarizes who you are and links to your most important pages. Support from AI providers is still uneven, so treat it as a low-cost extra rather than a ranking lever. It takes an hour to write and forces you to state your positioning clearly, which is useful on its own.
7. Earn mentions beyond your own site
Assistants weigh what others say about you. Reviews, client case studies published by partners, podcast appearances, community answers and listings in credible directories all add up. One genuinely useful guide that others link to beats ten thin blog posts.
8. Measure it
- In GA4, look at referral traffic from sources like
chatgpt.com,perplexity.ai,gemini.google.comandcopilot.microsoft.com. - Once a month, ask the major assistants the questions your buyers ask and note whether, and how, you're mentioned.
- Track branded search volume in Search Console; AI mentions often show up as more people searching your name.
- Verify your site in Bing Webmaster Tools as well. Microsoft Copilot draws on Bing's index, and Bing is easy to forget about.
What we did on our own site
We apply this checklist to proindies.com, so here's the concrete version.
Our robots.txt names the AI crawlers individually and allows each one: OpenAI's GPTBot, OAI-SearchBot and ChatGPT-User; Anthropic's ClaudeBot, Claude-SearchBot and Claude-User; PerplexityBot and Perplexity-User; Google-Extended; Applebot-Extended; Amazonbot; DuckAssistBot; CCBot and several others. We allow training crawlers as well as search agents, because our goal is to be findable and citable in as many assistants as possible. The file also points to our sitemap and our llms.txt.
Our llms.txt is a plain Markdown summary of the business: where we're based, each service with a link to its page, the numbers we stand behind (150+ projects in 40+ countries since 2016, 82% client retention), selected work with published outcomes, common questions with direct answers, and our contact details. It's written for a reader who knows nothing about us, which is a useful discipline for any business.
The site is a static HTML export. Every page, including this one, is prebuilt HTML, so a crawler that never runs JavaScript still gets the full text.
Every page carries ProfessionalService markup with one shared identifier, https://proindies.com/#organization, along with our address, contact points, service area and social profiles. Each service page adds Service markup that names that organization as the provider. Pages with a visible FAQ add FAQPage, inner pages carry BreadcrumbList, and posts like this one use BlogPosting.
When we first added these files, we also went through every page and fixed the drift: the contact email was normalized to one address, the phone number was written the same way in all the structured data, and a duplicate <title> on the services page was removed. Small things, but they're exactly what assistants cross-check.
Finally, the sitemap is generated from the site's content at build time, so a new service, case study or post is listed without anyone having to remember.
Common mistakes
- Allowing a bot in
robots.txtwhile the CDN or firewall blocks it. Check your bot-protection settings and your server logs, not just the file. - Assuming a named group inherits the
*rules. It doesn't. A crawler follows the most specific group that matches it, so if you add a section forGPTBot, repeat anyDisallowlines you still want it to respect. - Blocking
Google-Extendedto stay out of AI Overviews. It doesn't do that. AI Overviews come from Search, whichGooglebotcontrols. - Markup that doesn't match the page. FAQ schema for questions nobody can see, or ratings that aren't shown, goes against Google's structured data guidelines and undermines trust in the rest.
- Key facts that only live in images, PDFs or scripts. Prices in a banner image or services inside a widget that loads later are hard for any crawler to read.
- Treating
llms.txtas a shortcut. It summarizes good pages. It can't replace them. - Stale listings. An old address or discontinued service on a directory keeps getting repeated until you fix it at the source.
- Writing for machines. Pages stuffed with phrases like "best AI-optimized agency" read badly and give an assistant nothing worth quoting.
A simple 30-day plan
Week 1: access. Review robots.txt and your CDN bot settings. Run the view-source test on your ten most important pages. Verify the site in Google Search Console and Bing Webmaster Tools if you haven't already.
Week 2: answers. Pick the five pages that matter most to sales. Rewrite each opening so it answers the main question in the first sentence, and add a short FAQ built from questions customers really ask.
Week 3: facts and markup. Add or fix structured data on those pages. Then align your name, address, phone number and services across your site, Google Business Profile, social profiles and the directories you're listed in.
Week 4: extras and a baseline. Write your llms.txt. Set up referral tracking for AI assistants in GA4. Write down ten questions your buyers ask, put each one to the main assistants, and record what they say today, so next month's check has something to compare against.
FAQ
How long before AI assistants start citing us?
There's no fixed timeline. Assistants that search the web live can pick up a page soon after it's crawled. What a model learned during training only changes when that model is updated. Track it monthly rather than daily.
Will allowing AI crawlers hurt our Google rankings?
No. Rules for AI crawlers don't change how Googlebot crawls or ranks your site. The main cost is server load, which is rarely a concern for a static or well-cached site.
Should we allow training crawlers or only search agents?
That's a business decision. Search and user-triggered agents are what make you citable in live answers. Training crawlers affect what future models know about you. We allow both, but a publisher whose content is the product might reasonably block training.
Do we need an llms.txt file?
No. It's optional and support is uneven. It's cheap to write, though, and the exercise of summarizing your business clearly is worth doing anyway.
The short version
Make your site easy to crawl, write pages that answer real questions, describe your business the same way everywhere, and give others reasons to mention you. None of it is a trick, and all of it compounds.
If you'd like a second pair of eyes, our SEO & AI search visibility team runs a focused audit and fixes the highest-impact issues first.
