Which Oregon senior living communities does AI actually recommend?

An audit of 1,329 sources behind 40 AI answers about senior living in six Oregon markets.


Summary

We asked one AI engine — Anthropic’s Claude, with live web search — the questions Oregon families actually ask when a parent needs care — what are the best assisted living facilities in Eugene, what does it cost per month in Medford, how do I choose a community for my mom in Portland — and then, instead of reading the answers, we read the sources.

Forty answers, on two dates, in six markets. 1,329 individual source records. 178 distinct domains. Every one recorded, classified, and joined to Oregon’s own licensed-facility register.

Read the numbers below as one engine, two dates, six Oregon markets, five questions each. Not “AI”. Not ChatGPT, which was not tested. And nothing here — nothing anywhere in this project — connects an AI citation to a single enquiry, a single tour, or a single admission.

The short version of what we found:

QuestionWhat the scan showsn
Whose pages does the assistant read?Aggregator and directory sites are 60.9% of all cited sources. On the questions that actually name communities to a family, 95.4%.842 citations / 239 on naming questions
Do operators’ own websites get read?Rarely. 3.9% of citations overall, 1.3% on naming questions. Of 62 communities the answers named, exactly 1 had its own website cited anywhere in its market.842 / 183 communities
What is the biggest single variable?The type of question. A “who is best” question draws on 39 domains and is 95.4% aggregator. A “how do I choose” question draws on 113 domains, a third of them government.239 / 491
Is the cited set closed to newcomers?No. 178 domains, 100 of them cited only once or twice. Fifteen genuinely new domains — one registered a month before the scan — drew 64 citations, and 75% of answers cite at least one.1,329
What predicts a community being named?Two things survive controls: licensed beds (OR 1.16 per 10 beds) and presence on the A Place for Mom city roster (OR 5.67). Both are associations, not levers. The direction of the arrow is not established, and whether claiming a listing changes anything is completely untested.183
What does not predict it?On-site schema markup, technical site score, star rating, photo count, published pricing, review recency, and operator footprint. All null. Schema and technical score returned a null on two separate tests — and the schema coefficient is negative, not zero.26–183

The one-line version: in this corpus the assistant does not recommend senior living communities so much as it recommends the directories that rank them — and the operator’s own website is almost never in the room.

One thing we are not publishing. We know exactly which communities were absent from every answer in their market. That list stays private. Publishing it would be a public failure report on named businesses that had no notice and no right of reply, in a category where reputation is a safety signal to families. Aggregate patterns are the finding; the per-facility version belongs in a private conversation with the operator, if anywhere.


What was measured, and what a “citation” is and is not here

The corpus is 40 saved Anthropic transcripts:

  • 30 answers — six Oregon markets (Bend, Eugene, Gresham, Medford, Portland, Salem), five questions each, scanned 17 August 2026. Two questions per market that name communities (“what are the best assisted living facilities in X”, “which memory care communities in X are highly rated”), one on monthly cost, two general-advice questions.
  • 10 answers — a hyper-local scan around one Portland address, scanned 18 August 2026, adding Medicaid and neighbourhood-level questions.

The unit of analysis is one retrieved source — one entry in the tool’s returned source list. Domains are reduced to registrable form. The six-market subset alone is 842 retrieved sources; the full corpus is 1,329 across 178 domains.

This is a retrieval measurement, not an attribution measurement, and the difference matters. The API distinguishes a web_search_result — a page returned to the model — from a web_search_result_location, the per-claim attribution the model actually makes. These transcripts retain the first and not the second. Every figure below therefore describes what the engine was shown, not what it demonstrably used. Where this article says “cited”, read “retrieved and placed in front of the model.” Two consequences follow. Repeat searches inside a single answer inflate counts: in the Bend “best facilities” answer the model ran the same query twice and the same seven URLs entered the corpus twice. And source position means index in the concatenated result list for that whole answer, not rank inside any one search.

Half the corpus is Portland, and that is a limitation, not a design choice. The Portland hyper-local scan contributes 487 retrieved sources and the Portland market scan 190 — 677 of 1,329, or 50.9%. The other five markets supply 652 between them. Four of the ten hyper-local questions are also near-restatements of market-scan questions (“how do I choose an assisted living community for my mom in Portland?” appears in both, returning 74 sources and 62). Every full-corpus figure in this article — the 178 domains, the concentration percentiles, the age table, the government share — is weighted towards one city and, in part, towards one neighbourhood inside it. The 40 answers are not 40 independent observations.

Every domain in the six-market subset was classified by hand from its recorded title and URL, with nothing left in a default bucket. Domain ages come from RDAP registration records retrieved 23 August 2026, resolved for 172 of 178 domains, with every apparently-young domain checked against the Wayback Machine and excluded if its archive history predated its registration date by more than a year — five domains were dropped that way, including the most-cited candidate.

For the facility-level analysis, the population is Oregon’s own register: ODHS licensed facilities with status = Open in the six scanned cities, collapsing assisted-living and residential-care licences at the same address under the same brand into one community. That gives 183 communities, of which 62 (33.9%) were named in at least one answer, across 89 naming events. That population is the point — it includes the 121 communities no answer mentioned, which is what makes any of the correlations below meaningful rather than circular.

Nothing here was estimated. Where a figure could not be observed on a page, it is recorded as unverified and left out.


Finding 1 — the substrate is directories, not operators

Across the six-market scan, here is who the assistant read:

Source classCitationsShare
Aggregator / directory51360.9%
Government / regulator16619.7%
Advocacy / nonprofit / ombudsman839.9%
Operator’s own website333.9%
Legal reference333.9%
Media / academic91.1%
Social50.6%

The most-cited domains were oregon.gov (92), aplaceformom.com (59), usnews.com (47), seniorly.com (35), caring.com (29), payingforseniorcare.com (29), and state.or.us and seniorhomes.com tied on 27. Two of the top eight are Oregon government hosts.

That aggregator category is also more consolidated than it looks. Caring.com acquired SeniorHomes.com in 2015, PayingForSeniorCare.com in 2018 and SeniorAdvice.com in 2020 — so three of the seven most-cited domains, 85 citations between them, belong to one company. SeniorAdvisor.com belongs to A Place for Mom. The cited substrate in Oregon senior living is, functionally, a small number of roll-ups plus a state government register.

The operator finding is the sharper one. Of the 33 citations that did land on an operator’s own website, 25 fell on the two general-advice questions — a corporate blog post titled something like “Questions to ask when touring an assisted living facility,” quoted as generic guidance. Exactly one operator citation landed on a “what are the best facilities” question.

And of the 62 communities the answers actually named, exactly one had its own website cited as a source anywhere in its market. The correlation between “own site cited” and how often a community is named is rho = −0.037, p = 0.62, n = 183 — nothing at all.

Being named and being cited are, in this data, nearly disjoint events.


Finding 2 — question type is the real variable

Split the same corpus by what the family is asking and the picture changes completely.

Question typeAnswersCitationsDistinct domainsDomains cited once
Naming — “best…”, “highly rated…”122393916
Cost — “what does it cost per month”61123313
Advice — “how do I choose”, “what to ask”1249111350

An advice question draws from a domain pool nearly three times wider than a naming question, from roughly twice the citations.

The class mix moves with it (six-market subset, 842 sources):

ClassNaming (239)Cost (112)Advice (491)
Aggregator / directory95.4%95.5%36.3%
Government2.9%0%32.4%
Advocacy / nonprofit0.4%16.7%
Legal reference6.7%
Operator’s own website1.3%4.5%5.1%
Media / academic1.8%
Social1.0%

Each column sums to 100.0%. Nothing is dropped.

The pattern is consistent: the more a question asks for a judgment about named businesses, the harder it collapses onto a small set of directory domains. The more it asks about process, regulation or procedure, the wider and more open the source pool.

This matters because it is the one structural fact in the study that anyone can act on without believing anything about mechanism. If your content answers “how do I choose,” you are competing in a 113-domain pool. If it answers “who is best,” you are competing in a 39-domain pool that is 95% owned by directories.


Finding 3 — regulatory authority does not transfer to “who is best”

Government sources are 233 of 1,329 citations (17.5%), across 14 distinct regulator domains. (Two further public-sector domains — a county library and a health-department software vendor, four citations between them — are public but not regulators, and are outside this count.) Oregon’s licensing register is directly and repeatedly read: ltclicensing.oregon.gov was cited 23 times across 14 of the 40 answers, in all six markets, including individual facility-detail pages rather than just the search homepage. The ODHS long-term-care quality page drew 18 citations and the Oregon Consumer Guide PDF drew 10.

So the register is machine-readable and being read. But look at where those citations land:

Question type (full corpus, 1,329)Government citationsShare of that question type
Advice21933.6%
Medicaid (10-question scan only)425.0%
Hyper-local23.4%
Naming82.2%

Government sources are a third of the substrate on “how do I check a facility’s record” and essentially absent from “who is the best.”

There is a near-natural experiment in the corpus for this. thecareaudit.com is a directory built directly on the same ODHS inspection data, and it was retrieved 14 times across 10 of the 40 answers, in every market. All 14 are on advice questions. Zero on naming questions.

One caution on the coverage comparison, because it is easy to overstate. The site advertises “571 licensed assisted living facilities in Oregon.” Our own ODHS pull holds 572 open facilities — but that is 240 assisted-living licences plus 332 residential-care licences, and “assisted living” alone is 240. The two totals match almost exactly only if the Care Audit’s label covers both licence types, which is likely but which we have not confirmed with them. Treat 571 and 572 as suggestive of the same underlying register, not as a verified like-for-like count.

Licensing-derived authority transfers. It transfers to the questions that do not name anybody.


Finding 4 — the cited set is not closed, and new domains enter it fast

The 60.9% aggregator headline suggests a sealed oligopoly. The domain-level distribution does not agree.

Across the full corpus, the top domain holds 9.3% of citations, the top three hold 20.6%, the top ten hold 41.8%, and the top fifty hold 82.4%. One hundred of the 178 domains — 56% — were cited only once or twice. Both facts are true at once: the aggregator category is a fortress, and the roster of domains inside it turns over.

Domain age, for the 172 domains with a resolved registration date (median age 19.3 years). Those 172 domains account for 1,265 of the 1,329 retrieved sources, and the share column below is computed on that 1,265, not on the full corpus — the six unresolved domains and their 64 sources sit outside it:

Age at scanDomainsSourcesShare of the 1,265
Under 1 year9403.2%
1–2 years8463.6%
2–3 years5151.2%
3–5 years6282.2%
5–10 years2415612.3%
10–15 years1512810.1%
15+ years10585267.4%

Old domains dominate, exactly as an incumbency argument predicts. But the under-two-year band is not empty, and after the re-registration control it still holds 11 domains and 55 citations.

Restricting to domains registered within 2.6 years of the scan and with no earlier archive history: 15 domains, 64 citations, 4.8% of the corpus. Three details make that more than a curiosity.

  1. grouphomepath.com was registered 17 July 2026 and cited on 17 August 2026 — a domain one month old, cited four times, in both scans.
  2. 30 of the 40 answers (75%) cite at least one genuinely new domain. This is not one fluke in one market; it is a routine feature of the source list.
  3. New domains are not buried at the bottom. One of them first appears at source positions 4 through 8 in each of the nine answers that retrieved it. It also recurs lower down in six of those answers, as far down as position 23.

And one of them reached the commercially valuable lane. embraceageprepared.com, registered 5 September 2024, is a programmatic listicle site — its sitemaps enumerated 147 URLs when we checked them on 24 August 2026: 101 blog posts, 33 pages, plus category, author and staff URLs, running city-by-care-type “best of” pages across more than a dozen Oregon cities, with titles like “Best memory care in Eugene, Oregon” and “Best assisted living communities in Portland, OR.” Seven of those URLs were retrieved in this corpus, 15 times between them, 12 on naming questions, which places the domain eighth on the naming-question table with 12, level with Yelp and SeniorAdvisor.com and ahead of PayingForSeniorCare. It appeared in 7 of the 14 naming answers, spanning four of the six markets. Note what it is not: a two-year-old site with a hundred programmatic city pages, not a seven-page boutique.

That is n = 1. One site under two years old cleared the lane. A single case is a proof of possibility, not a base rate — but it does falsify the stronger claim that naming questions are sealed to anything without a decade of accumulated authority.

The pattern shared by the new sites that got cited

Offered as description, not as a formula:

  • They answer the question in the question’s own words. “Best memory care in Eugene, Oregon” against “which memory care communities in Eugene, OR are highly rated.”
  • They order rather than list. The licensing-derived directory surfaces documented violation counts per facility and builds ordered tables on them — while explicitly disclaiming that it “does not conduct inspections, assign ratings, or make recommendations.” The listicle site orders by “best.” Neither publishes a neutral alphabetical register. An answer to “who is best” needs an ordering.
  • They are year-stamped. Among the 64 citations to genuinely new domains, 21 (33%) had a “2025” or “2026” year stamp in the page title. Among the 980 citations to domains ten years or older, 59 (6%) did. Fisher exact OR = 7.62.

That last one carries three caveats and all three matter. The unit is retrieved sources, not independent observations, so the clustering inflates any nominal significance — which is why no p-value is quoted for it. This is a title-text pattern and it is not a freshness test. The API returns a page_age field on every result; it is present on all 1,329 records in this corpus and populated on 412. We tested it. Year-stamped pages are nominally fresher — median 5.7 months against 12.4 — but most of that gap is domain age, because a young domain cannot host an old page, and once domain age is stratified out the difference weakens sharply and is not significant among domains under ten years old (p = 0.059). The populated subsample is also non-random: page_age is present on 41.5% of aplaceformom.com citations and 3.3% of oregon.gov citations, so it describes the commercial aggregator layer rather than the corpus, and the 81 year-stamped citations come from just 20 domains with three supplying 44 of them. We therefore report the year stamp as a title-text pattern only. Whether real page freshness predicts retrieval remains untestable here, because page_age exists only for pages that were retrieved and this corpus contains no sample of pages that were not. Finally, the pattern is entirely consistent with a duller explanation: new sites in this space simply adopt the “2026 guide” title format, and something else about them drives retrieval. It is not a mechanism and should not be sold as one.


Finding 5 — what predicts a community being named

This is the part operators care about, and it is the part where most of the answers are null.

Working population: 183 licensed communities across the six market cities, outcome = how many of that market’s five answers named it (0–5).

PredictornrhopMarket-blockedVerdict
On the A Place for Mom city roster183+0.2530.00056+0.249, p = 0.0013Survives everything
Licensed beds (size)183+0.2360.0013+0.229, p = 0.0018Real, modest, survives
Review volume on that roster145+0.2370.0042+0.225, p = 0.0076Collapses under rank — see below
Licences at address (campus)183+0.1840.013Weak, collinear with beds
Operator campus count183+0.1350.068+0.019, p = 0.80Dies under blocking
Has own website183+0.1250.092Not established
Accepts Medicaid183−0.1420.055Flagged, not claimed
Star rating74+0.1130.34Null
Technical site score /10052+0.1040.46+0.096, p = 0.51Null
Photo count on listing109+0.0820.40Null
Pricing shown on listing145+0.0460.58Null
Own website cited as a source183−0.0370.62Null
Days since most recent review82−0.0440.69Null
Pages with schema markup52−0.2770.047Negative, not positive

Fourteen predictors appear in the table above, and more than twenty were tested across the two underlying analyses. At α = 0.05 about one spurious result is expected, and a Bonferroni-style threshold on twenty tests is 0.0025. Only roster presence and licensed beds clear it — review volume, at p = 0.0042, does not. Everything else on that table is exploratory, including the negative signs.

Size is real and it is not a lever

Logistic regression with market fixed effects gives OR 1.158 per 10 licensed beds, p = 0.0038. Compounded across the 55-bed gap between a 30-bed and an 85-bed building, that is roughly 2.2 times the odds of being named, holding market constant. Odds are not likelihood: at the observed 33.9% base rate, 2.2 times the odds is about 1.6 times the probability. Nobody buys beds to get cited. The practical use of this finding is as a control: any future claim that some intervention moved citation has to beat building size first.

Note what dies: operator footprint. The pooled correlation of +0.135 was Portland having both more chains and more competition. Once market is held constant it is +0.019, p = 0.80 — no association at all.

Roster presence is the strongest association — and it is an association

Being listed on the A Place for Mom city roster at all: Fisher exact OR = 5.67, p = 0.00044, n = 183, surviving both a control for licensed beds (+0.204, p = 0.0056) and market blocking (+0.249, p = 0.0013). Caring.com’s own listing variable agrees directionally (OR 2.44, p = 0.0079) though it weakens under market blocking.

This must be read as an association and nothing more, and the arrow is genuinely not established.

Start with the mechanism this study’s own Finding 1 implies. Naming answers are 95.4% aggregator sources, and A Place for Mom is the single largest naming domain. The engine largely cannot name a community it was never shown, and what it was shown was the aggregator’s city page. Roster presence is therefore closer to a precondition for being nameable in this corpus than to an independent predictor discovered inside it. That is a weaker claim than a causal one and a more useful one than an unexplained correlation.

Two further reasons not to treat 5.67 as a lever. Aggregators decide who to list partly on size, partner status and market, so a large, well-known, long-established building is both more listable and more nameable — the same underlying property could be producing both facts. And a closely related variable was deliberately kept out of the table above: “a page specific to this facility was retrieved in this market” scores rho +0.336, Fisher OR 18.99, p 0.000002 — the largest effect anywhere in this dataset, and entirely tautological, because the engine fetched that page while composing the answer that names the facility. It is excluded because it measures co-occurrence inside a single generation. It is disclosed here because the presentable number should not be the only one a reader gets to see.

Whether claiming an unclaimed profile changes anything is completely untested — no before-and-after data exists anywhere in this project, and a cross-section cannot answer it. Anyone quoting 5.67 as evidence that claiming a listing produces citations is going beyond what this measured.

Reviews: a real-looking effect that does not survive the obvious alternative

Among the 145 communities with an observed review count, review volume correlates with citation at rho = +0.237 (p = 0.0042), holding at +0.202 with beds controlled. That looks like a finding until you control for rank on the category page, at which point it goes to r = +0.080, p = 0.34.

Reviews and roster rank correlate at rho = −0.702. A Place for Mom ranks its city pages partly on review score, so “has many reviews” and “sits near the top of the page” are close to the same fact about a facility. Control either and the other goes null. This corpus cannot tell them apart. The association also fails to replicate on Caring.com (rho +0.099, p = 0.45, n = 61, on a sample truncated to that site’s top 20 per market), and it does not clear multiple-comparison correction.

The counter-examples are numerous enough to matter on both sides. Fourteen facilities with 40 or more reviews were never named in any answer in their market. Meanwhile 20 of the 58 cited-and-listed facilities have zero reviews on the aggregator page the engine read, including two Bend communities each named in three of that market’s five answers with no reviews at all. Zero-review share was 34% among cited facilities and 59% among uncited — a real gap, and nowhere near a rule.

Review recency, previously untestable, now has data: null, n = 82.

Schema markup: two nulls, on two different units

This one is worth stating flatly because it is the single most commonly sold “AI visibility” intervention.

  • First test, 26 operator domains. Mean citations for sites with schema: 1.47 (n = 19). Without schema: 1.43 (n = 7). Indistinguishable.
  • Second test, 52 communities at facility level. Mean citations with schema 0.44 (n = 41) versus 1.00 without (n = 11); pages_with_schema rho = −0.277, p = 0.047.

These are two related nulls rather than a replication in the strict sense — different units (operator domains against communities), different outcome variables, different scans. What they share is the direction: nothing.

Do not read the second one as “schema hurts.” With 11 no-schema communities it is noise, and it is confounded in ways the obvious candidates do not explain. The defensible statement, holding on both tests, is narrower and sufficient: no positive association between on-site structured data and AI citation exists in this data. Schema keeps its value for rich results and entity consistency. It has no measured value here.

Google’s own documentation is consistent with that, for whatever it covers. Under a heading warning against overfocusing on structured data, Google writes: “Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add. However, it’s a good idea to continue using it as part of your overall SEO strategy, as it helps with being eligible for rich results on Google Search.” The second sentence is part of the quote and part of our position.


Finding 6 — the surface moves, and nobody documents how it works

Two things bound how much weight any of the above can carry.

The provider does not publish a selection mechanism. Anthropic’s documentation for the web search tool describes the process in three steps — Claude decides when to search, the API runs the searches and returns results, Claude answers with citations. That is the entire published account of how sources are chosen: it names no index, no ranking method and no relevance criterion.

It does document control surfaces, and honesty requires listing them. Six parameters are documented — max_uses, allowed_domains, blocked_domains, user_location, allowed_callers and response_inclusion — of which user_location is the one that matters here, because it “allows you to localize search results.” Caller-side domain filtering is real and documented, as is an organisation-level domain restriction in the Console. So is dynamic filtering: from web_search_20260209 onward, “Claude instead writes and runs code that filters the results first, so only relevant content reaches the context window.” Our own scans ran with that filtering active — the transcripts contain 150 completed code-execution result blocks alongside 152 direct search-result blocks — so a filtering pass we cannot inspect sits between the search engine and the source list we counted. Separately, page_age (“when the site was last updated”) is a field on each returned result, not a parameter; it is populated on 412 of the 1,329 records and we report what testing it showed in Finding 4.

Domain authority, backlinks, schema and site age are not mentioned anywhere in that documentation — which is not evidence they play no role, but does mean no published provider source supports the incumbency argument either.

Independent work finds the AI surface much less stable than organic search, and — depending on the engine — somewhat less head-concentrated. Grossman et al. (How Generative AI Disrupts Search, SIGIR 2026, 11,500 queries) report that Google Search draws 37.8% of its sources from the top-1,000 domains, 1.3 points higher than AI Overviews (36.5%) and 9.9 points higher than Gemini (27.9%). Read that carefully: AI Overviews are within a point and a half of classic Google. Only Gemini diverges meaningfully, so this paper does not support a blanket “AI is less concentrated” claim. Their sharper figure is coverage rather than share — top-1,000 domains appear for 52.7% of traditional SERP queries against 40.0% for AI Overviews and 32.6% for Gemini.

A separate preprint (Zhang, Ye, Peng, Garimella and Tyson, arXiv 2512.09483 — not peer-reviewed) analysed 55,936 queries across six LLM search engines and found greater domain diversity than traditional engines, with 37% of cited domains unique to LLM search.

On instability, Kirsten et al. (Characterizing Web Search in the Age of Generative AI, Findings of ACL 2026) found AI Overview retrieved-page overlap across a two-month gap was 18%, against 45% for organic Google. They also measured answer polarity flips on yes/no/mixed decision questions repeated five minutes apart: 9–27% across five generative engines at temperature zero, with AI Overviews at 17% — and note that temperature is not even controllable on AI Overviews or GPT-Search, so the band is not one system’s variance. Semrush — a vendor, flagged as such — tracked 230,000 prompts over 13 weeks (snapshots 14 July to 12 October 2025) and reported ChatGPT’s Reddit citation rate falling from close to 60% of responses in early August 2025 to around 10% by mid-September 2025.

Any single-snapshot AI visibility measurement, including this one, is partly measuring noise. That is a reason to date every figure and re-measure, not a reason to ignore the numbers.

One further caution on the optimisation literature. The widely-cited GEO paper (KDD 2024) reports content tactics improving visibility by up to 40%, but its experimental setup places the source document inside a five-document context and measures which of the five gets quoted more. It says nothing about how a document enters the retrieval set. Its own limitations section concedes it did not evaluate effects on search rankings. An independent peer-reviewed follow-up from a different group (C-SEO Bench, NeurIPS 2025) adds the multi-actor case — what happens when competitors adopt the same tactics — and finds most such methods “not only largely ineffective but also frequently have a negative impact on document ranking.”


What this study does not establish

Stated plainly, because this audience will check.

  1. One engine. Every measurement is Anthropic’s web search. Nothing here speaks to ChatGPT, Perplexity, Gemini or Google AI Overviews, which may use entirely different retrieval. Whether these findings replicate on another engine is the single largest open question and it has not been tested.
  2. One state, one vertical, two dates. Oregon senior living, scanned 17 and 18 August 2026. Forty answers.
  3. No denominator on the new-entrant finding. The 15 new domains that got cited are survivors. Nothing in this corpus reveals how many comparable sites launched in the same window and were never cited at all. The rate could be one in five or one in five hundred. Treat the observed timelines as “how fast it can happen when it works,” never as “how fast it will happen.”
  4. No causal mechanism, anywhere. Why any particular site was cited is unknown. Every correlation reported here is descriptive.
  5. Citation is not traffic, and traffic is not revenue. Nothing in this corpus, or anywhere in our data, connects an AI citation to a single enquiry, tour, admission or dollar. That chain is the actual business question and it is completely unmeasured. Independent context suggests caution: as of May 2026, roughly 6.8% of US ChatGPT prompts carried any citation at all, and under 4% in professional services (Similarweb, vendor-reported).
  6. The reviews question is unresolved, not answered. At n = 145 the smallest correlation detectable with 80% power is 0.231, and the observed beds-controlled figure is 0.202 — below that line. This sample can rule out a large review effect. It cannot establish or refute an effect of the size actually observed.
  7. Source counts are unstable and were not used as a metric. The five Gresham questions were asked twice on 17 August 2026. Three came back near-identical (20 vs 20, 25 vs 25, 35 vs 34 sources); two swung hard — “how do I choose…” returned 38 sources on one run and 84 on the other, and “what are the best…” returned 15 and 7. The two near-duplicate Portland pairs behave the same way (74 vs 62, 64 vs 79). Instability is real, it is not uniform, and it is why only the named communities and the classified domains are treated as signal.
  8. The corpus is Portland-weighted and not fully independent. 50.9% of the 1,329 sources come from two Portland scans, four of the ten hyper-local questions restate market-scan questions, and the 40 answers are treated as independent in the year-stamp test when they are not entirely. Every full-corpus figure carries this.
  9. The retrieval/attribution gap. These are sources returned to the model, not per-claim citations the model made. A domain’s count partly measures how often a query was re-run inside one answer. Treat every count as “how often the engine was shown this,” not “how often the engine relied on it.”

How to check this yourself

You do not need our tooling to test the central claim. Ask an AI engine with live web search “what are the best assisted living facilities in [your city],” then open the source list rather than reading the answer. Count how many of those sources are directories, how many are government, and how many are the websites of the communities being recommended. Then ask “how do I choose an assisted living community in [your city]” and count again.

In this corpus, that second list was roughly three times wider and a third governmental. If you get the same shape, the finding travels. If you do not, we would genuinely like to know — the one-engine, one-state limitation above is the honest reason to expect variation.

Every number in this article traces to a saved transcript, Oregon’s public licensed-facility register, a public RDAP or Wayback lookup, or a cited public page, all dated. Each of the 40 transcripts carries an API-returned container timestamp that fixes its scan date — 30 on 17 August 2026 for the six market scans, 10 on 18 August 2026 for the hyper-local Portland scan — independently of any local file metadata. Those 40 answers are the whole corpus. Scans this practice ran later, in other Oregon markets and against individual operators, are not part of it and contribute no figure in this article. Where something could not be verified, it is absent rather than estimated.


Who made this, and why

Direct-Me LLC is a one-person SEO and AI-visibility practice in Oregon working with senior living operators. We built the scanning tools for client work, found we had a dataset nobody else seemed to have, and published the aggregate patterns because the alternative — a market where everyone asserts how AI picks sources and nobody shows the source list — is worse for the operators we work with than an honest set of nulls.

Two things we will not do. We will not publish which communities were invisible; that conversation happens privately with the operator or not at all. And we will not promise an AI citation. Nobody controls model output, this study measures associations on one engine on two dates, and anyone telling you otherwise is selling you something.

Questions, corrections, or a replication on another engine: Direct-Me LLC, direct-me.org.

Scans dated 17 and 18 August 2026. Domain registration data retrieved 23 August 2026; external sources and quotations re-verified 24 August 2026. Figures are accurate as of those dates and will drift.

Scroll to Top