Ask an assistant which agency to use for a villa sale in Dubai and it has to get its facts from somewhere. I wanted to know what five real Dubai agency homepages actually hand it, so on 2 October 2026 I scanned them the way a crawler sees them: raw HTML, JavaScript switched off, no login, no interaction.
I expected a story about being blocked. The real story was about how much of the site survives contact with a crawler at all.
How do real estate agents get clients when buyers now ask an assistant first?
Mostly the way they always have: referrals, a feed that shows finished work and an agent whose name a past client remembers. An AI assistant’s answer is a newer layer on top of that, not a replacement for it. On 2 October 2026 I scanned five Dubai real estate agency homepages to see what that newer layer actually has to work with.
Here is what came back.
| What was measured | Result across five agencies |
|---|---|
| Readable words with JavaScript off | 338 to 4,750 |
| Sites blocking AI crawlers | 0 of 5 |
| Sites with real estate agent or organisation schema | 5 of 5 |
| Sites with FAQ (question and answer) schema | 2 of 5 |
| Sites serving an llms.txt | 2 of 5 |
| Headings phrased as a question | 7, across all five sites combined |
No agency is named. This is a category finding about Dubai real estate marketing, not a report on any one firm. Naming a prospect’s weak result in public marketing would be a strange way to treat people Nomo Creative might work with next year.
Do real estate agency websites block AI crawlers?
Not these five. None of the five homepages disallowed GPTBot, ClaudeBot or PerplexityBot in robots.txt. Every request returned a normal page, not an error or a block page. On the access question everyone worries about, all five passed.
That matches what the law firm scan found in September: access is rarely the problem anymore. Most sites that look modern already render on the server or allow the crawler through, often without anyone deciding to.
So what is the actual gap?
Volume, not access. The richest of the five homepages served 4,750 readable words with JavaScript off, a normal amount of homepage text. The thinnest served 338. Both sites looked complete in a browser. One of them, read the way an assistant reads it, was close to a blank page with a logo on it.
That gap usually comes from how the site is built, not what the agency chose to say. A homepage built as a single page application loads most of its real content after the page arrives, in JavaScript a browser runs and most AI crawlers do not. The agency’s listings, team bios and area guides exist. The crawler never sees them.
This is also where the social half connects to the search half. An agency posting listings and sold boards to Instagram and LinkedIn every week is generating exactly the fresh, specific, dated material an assistant would want to quote: this postcode, this price, this month. If the website serving that content to a crawler is mostly empty, the posting work on social never reaches the answer layer either. Same agency, same effort, two different outcomes depending on one technical choice nobody reviewed.
Does schema and an llms.txt actually help?
They are the easy part and these agencies had mostly done it. All five homepages carried real estate agent or organisation schema, two carried FAQ schema with real question and answer pairs marked up for a machine to read. Two served an llms.txt file pointing a crawler straight at listings and service pages. Seven headings across five sites were phrased as an actual buyer question rather than a service category, a better showing than the five Dubai law firms scanned in September managed between them.
None of that fixes a page where the words themselves never arrive. Schema describes content. It does not create content that was not rendered.
Does any of this change whether a buyer clicks through?
Fewer buyers are landing on any individual result page when an AI summary sits above the normal list. The independent evidence for that is thinner than most marketing blogs suggest. The one study worth citing is Pew Research Center, published 22 July 2025, tracking 68,879 queries from 900 US adults in March 2025: with an AI summary present, users clicked a traditional result on 8% of visits against 15% without one. Only 1% clicked a link inside the summary itself. Google disputes the methodology. That objection is worth stating rather than skipped past. The sample is American and the data is six months old by the time this runs.
What survives that caveat is narrower: if a buyer asks an assistant which agency to use and yours does not come up, you were not in that conversation, whatever your Google ranking says.
What would a Dubai agency actually check first?
Load the homepage with JavaScript switched off and read what is left. A browser extension or a plain curl request does it in under a minute. If most of the listings, the team and the area guides vanish, that is the finding. It has nothing to do with content strategy yet.
Past that, the same connection applies that this scan kept surfacing: a feed that has not posted in weeks and a website an assistant reads as nearly empty are often the same neglect showing up twice. The free AI check runs a version of this scan on your own site in about ten seconds, so you can see your own numbers in that table rather than someone else’s.
The honest caveat
A scan like this tells you what a crawler receives. It does not tell you whether an assistant will actually name your agency over the one down the street, because nobody outside those companies publishes the ranking logic. Anyone who claims to have reverse engineered it is selling something.
What it does tell you is whether you have given an assistant anything to find. Five agencies here, none blocked, all carrying schema. Between 338 and 4,750 words of difference in what the same homepage actually hands over. That is a fixable gap worth checking before assuming the problem is content when it might be architecture.