AI answers

I scanned five Dubai law firm sites. Crawlers read all of them.

Five Dubai law firm homepages, scanned September 2026. Every one is readable by AI crawlers. Not one is written to be quoted. The gap is shape, not access.

Dominik · Published · Updated

1 heading phrased as a question across five Dubai law firm homepages. Our own scan of five homepages, 23 September 2026
Source: Our own scan of five homepages, 23 September 2026

TL;DR

Five Dubai law firm homepages were scanned on 23 September 2026. All five serve between 751 and 1,461 readable words with JavaScript switched off and none block AI crawlers. The access problem everyone talks about does not exist here. What is missing is shape: across all five sites there is exactly one heading phrased as a question. Assistants quote answers and nobody had written any.

Ask ChatGPT to recommend a corporate lawyer in Dubai. Count how many firms it names. Now count how many exist.

I expected the reason to be technical. It is not and the number that proves it surprised me.

Why doesn’t ChatGPT mention my law firm?

Most likely not because it cannot read your website. On 23 September 2026 I scanned five Dubai law firm homepages the way an AI crawler sees them: raw HTML, JavaScript switched off. Every one was readable. The problem is further along than access and almost nobody is looking there.

Here is what came back.

What was measuredResult across five firms
Readable words with JavaScript off751 to 1,461
Sites blocking AI crawlers0 of 5
Sites with business schema3 of 5
Sites serving an llms.txt0 of 5
Headings phrased as a question1, across all five sites combined

The firms are not named. This is a category finding, not a report card on anyone in particular and publishing a named firm’s weak result would be a strange thing to do to people I might one day work with.

Isn’t the problem that AI can’t read my site?

That is the common diagnosis and here it was wrong. All five sites served between 751 and 1,461 readable words without JavaScript, which is a normal amount of text for a homepage. None of them disallowed GPTBot, ClaudeBot or PerplexityBot. On the two checks everyone worries about, all five passed.

This matters because the advice in circulation is mostly about access. Unblock the crawlers. Server-render your pages. Good advice and these firms had already taken it, mostly by accident: four of the five run on ordinary content management systems that render on the server by default.

They did the hard part without trying. Then they stopped.

So what is actually missing?

Something plainer. Across five entire websites, I found one heading phrased as a question. Not one per site. One in total.

An assistant answering “do I need a lawyer to set up in DIFC” is looking for text that already reads like a reply to that question. What these sites offer instead is what every professional services site offers: Our Practice Areas. Corporate and Commercial. About the Firm. Categories, not answers.

A crawler can read all of it. There is simply nothing in it to quote.

That is the difference between being indexed and being cited and it is the whole of the problem for firms like these. Your site is in the library. Nothing in it answers the question being asked.

Does this actually cost anything?

Fewer people reach the page in the first place and the independent evidence for that is thinner than the marketing suggests.

The one study worth citing is from the Pew Research Center, published 22 July 2025, which tracked 68,879 queries from 900 US adults in March 2025. When an AI summary appeared, users clicked a traditional search result on 8% of visits, against 15% when no summary appeared. Only 1% clicked a link inside the summary itself.

Google disputes the methodology and that objection deserves stating rather than burying. The sample is American and the data is from early 2025. I would not build a business case on it alone.

What I would say is narrower and harder to argue with: if a buyer asks an assistant and your firm is not named, you were not in that conversation. Whether that conversation replaces a Google search or merely precedes one, you were absent from it.

What would you change first?

Take the ten questions a client asks in a first meeting and answer them on the site, one heading each, in the client’s words.

Not “Corporate Structuring Advisory”. The actual sentence a founder types: do free zone companies pay corporate tax in the UAE. Then answer it in about fifty words, directly, before any preamble about the firm’s heritage.

That is unglamorous and it is most of the work. The schema, the llms.txt, the technical layer, all of it helps at the margin. None of it creates something worth quoting where nothing exists.

If you want to know what your own site serves before you change anything, the free AI check runs the same scan I ran here and takes about ten seconds.

The honest caveat

A scan tells you what a crawler receives. It does not tell you whether an assistant will name you, because nobody outside those companies knows how the ranking works and anyone who claims otherwise is selling something.

What a scan does tell you is whether you have given them anything to work with. Five firms here, all readable, all open and between them one question anybody bothered to answer.

That is a fixable problem, which is the only good news in this piece.

Sources: Direct scan of five Dubai law firm homepages, 23 September 2026, requested with an AI crawler user agent and JavaScript disabled; Pew Research Center, 'Google users are less likely to click on links when an AI summary appears in the results', 22 July 2025

FAQ

Why doesn't ChatGPT mention my law firm?

Usually not because it cannot read your site. In a scan of five Dubai law firms on 23 September 2026, all five were readable and none blocked AI crawlers. The problem is that nothing on the pages is written as an answer to a question a client would ask, so there is nothing quotable.

Do AI crawlers read JavaScript?

Mostly no. GPTBot, ClaudeBot and PerplexityBot largely do not execute JavaScript, so they see the raw HTML a server sends. Google does render JavaScript, which is why a site can rank normally in search and still be close to empty in the answer layer.

What is an llms.txt file?

A plain text file at the root of a domain that tells AI systems which pages matter and how to describe the business. It is a young convention with uneven support. None of the five firms scanned on 23 September 2026 served one, so it is currently a differentiator rather than a requirement.