A diagnostic Menlo & Oak runs for your company. We ask AI engines questions written the way your customers write them, and show you, with counts and verbatim quotes, whether they name your brand, whether they recommend it and whether they cite your site.
In “Your Brand Is Not in the Answer” we argued that AI search no longer returns a list of links: it writes an answer and names a handful of brands. That paper's ninety-day plan starts with a baseline: knowing where you stand before you touch the site.
This diagnostic is that baseline, done by us. We run it with Olson, a tool Sebastián Gebhardt built at Yáneken to measure exactly this, and which Menlo & Oak offers to its clients. It isn't software we install for you, or a subscription: it's work we do for you, and it ends in a report.
In the same paper we said most tools in this space are mediocre, and that checking a fixed set of questions by hand every quarter is more honest than most dashboards. We built the diagnostic on that idea: fixed questions, counts in plain view, and the verbatim sentence behind every judgement.
We don't calculate a visibility score. We measure three different things, separately, for each brand:
A brand can be named without anyone recommending it, or have its site cited without the answer naming it. A single number would blend the three and hide which one is missing.
Every rate comes with its counts (for example, “7 of 40”, not just “18%”), so you know how much evidence sits behind each figure. And the same three figures are measured for your competitors, on the same answers.
We agree with you which brands to measure, against which competitors and on which engines.
We draft the question bank the way your customers would ask, without looking at your site. You review it before the first measurement; from then on it stays fixed until the second, because changing it would break the comparison.
We run the questions and deliver the report.
Your team works through the focus plan.
We run the same bank again, on the same engines and under the same conditions, and compare question by question. With a bank of at least ten unbranded questions, and if your brand appears in at least three of them, the comparison comes with a range showing whether the difference goes beyond what the choice of questions alone would produce.
No. There's no score, index or grade, by design. There are three separate figures, each with its counts.
No. We query the engines through their APIs. What an app shows each person depends on their location, their history and their memory, so we don't present these figures as what any particular customer sees.
No. Nobody controls what an engine answers. The second measurement records what changed between one run and the next; it doesn't prove the change is down to the work done.
When you decide. Each measurement is run on request, because each one has a cost. It isn't ongoing tracking.
We report it as inconclusive. The site reading shows what our reader received, not what any particular AI crawler sees.