Most visibility vendors sell a score and hide the math. We publish the method, because the method is the product: if you cannot see how a number was made, you cannot trust a decision built on it. This page is the whole discipline, including the mistake that taught us its most important rule.
A keyword is what someone types into a search box. A buyer prompt is what someone asks an assistant when they are ready to act: "Suno alternative that doesn't take rights to my songs" carries a budget and a decision in one sentence. We write each prompt set from the category, the competitors, and the one line you tell us a buyer would say, then rank every prompt by buying intent so the report reads highest-stakes first.
Prompt sets are locked per client. The same prompts run every time, because a changed question makes every past answer incomparable.
One prompt run once proves nothing: assistants are non-deterministic, and answers move between runs, engines, and days. So every prompt runs multiple times per engine, every answer is timestamped at capture, and the full raw text of every answer is kept, not just the verdict. In a paid Snapshot that is 5 prompts, 3 engines, 2 runs each: 30 captures with the run count printed on the report.
The dataset behind our own published numbers is append-only. Nothing is ever rewritten, because the trend across days is worth more than any single day, and a trend you can edit is not a trend.
Whether a brand was named is decided by deterministic text extraction against a declared brand list, not by an AI grading an AI. The same answer always scores the same way, and anyone holding the raw text can check any row.
Our first self-scan baseline said one engine never named us. It was wrong. A harness defect was returning empty answers, and an empty answer looks exactly like an engine declining to name you. We found it, threw out the baseline publicly, fixed the harness to fail loudly, and re-ran everything. The rule it left behind is permanent:
An empty answer is an engine failure, never a miss. A scored miss must contain a real answer that named someone else.
That distinction is the difference between measurement and theater, and most dashboards do not make it.
Assistants personalize. A signed-in session can favor the account holder, which is why every self-referential number in our own teardown is labeled an upper bound. Client captures come from sessions with no connection to your company, which is the cleaner read.
Engine coverage is disclosed per report: if one engine failed a run or hit a limit, the report says so and shows like-for-like tables instead of blending unequal samples into one flattering rate.
Nobody can make an assistant say anything, and we put that in writing on every page. What can be engineered is the evidence an engine finds when it looks: valid schema and JSON-LD, a coherent entity story across your pages, an llms.txt, citable pages that answer the buyer prompt directly, and comparison surfaces where the category decides. Those are durable inputs. Answers move; the inputs stay improved.
Every fix ships as something you can inspect: paste-ready patches or a pull request, with the before-evidence attached and a re-run scheduled to measure the delta. Delivery, spec, and refunds are published on the front page and in the terms.
The method above produced the self-scan teardown, losses included. Your category gets the same discipline: captures, scoring, prioritized fixes, refund-backed.