Dimension scores rebuilt onto anchored, evidence-cited rubrics: computed before narrated, band-led, with derived confidence. Prompted by a numerate founder asking how a score is derived.
The entire input is a brand name. What comes back is a six-dimension read of the brand, scored against written rubrics, run through a stack of checks built to disprove it. And it's honest about what it can't see from the outside.
This is the proof of work. Maybe you're the kind of operator who wants to know how a score got to be a score. Maybe you've seen enough AI-generated "brand audits" to be suspicious on sight. Both instincts are right. This page is built for them.
No questionnaire, no data room, no homework. A brand name is enough to start. A URL helps, if you have one.
Everything below is that arrow, slowed down. None of it is "ask a model what it thinks and format the answer." Each stage has a defined input, a defined output, and a defined way to be wrong on purpose. Most AI tools skip that last part. It's how the wrong thing gets caught before it reaches you.
The real pipeline, not an idealized diagram. Research and drafting are the bulk of the work; judgment stays with people, and with gates that exist only because something once slipped past.
Two feeds. The first is the brand's public signals: homepage, packaging, social, reviews, category forums, retailer listings, search presence. Read directly, never guessed from a footer. The second is the Design-Strategy OS: 276 pages of concepts, methods and strategy, wired together by more than 6,400 internal links, built from 75 sources and the field's best thinkers (Rumelt, Dunford, Christensen, Neumeier, Wheeler, Sharp, Aaker, Binet & Field), sharpened by 17 years building brands. It's what tells the engine what "good" looks like.
Three checks, running on different schedules. First, the build itself. Around 290 automated checks run every time a deck is assembled. If a number can't be traced to where it came from, if a quote isn't actually found at the source it's credited to, or if a caveat gets dropped between the appendix and the slide, the build stops and the file never gets written. Second, the library. A regression test of 67 fixed questions runs against the knowledge base after every ingest, so new reading can't quietly break what was already there. Third, the scoring. Brands get re-scored blind and the results compared, which is how the bands stay honest. Six brands the last time. That one runs at intervals, not continuously.
The brand is scored on Clarity, Consistency, Differentiation, Resonance, Availability, and Credibility. Each dimension is broken into written criteria scored against anchored rubrics. The score is computed from cited evidence first, then explained. Not the other way around.
Before anything reaches a slide it runs a stack of passes built to disprove the read. An assumption audit flags what public signals can't see, and what could already have killed the idea inside the company. A disconfirming search checks any fact an insight leans on. A senior-strategist critique. A tournament that makes the competing insights fight until one survives, so the headline finding is the one that beat the others rather than the one written first. A fact-check against primary sources. A claim-source-scope trace. And a two-way regulatory sweep, which checks both that a recommendation doesn't break a rule and that any rule we assert is real. Then, once the deck is actually built, it goes to a red-team of independent models from different labs, and anything two of them agree on comes back to be answered before it ships.
The read becomes a deck. Competitors run through three labeled lenses: the shelf set you're actually chosen against, the size-and-structure peers you resemble (so we benchmark what's achievable, not what a giant does), and a scan of anyone running your same strategic move. Every comparison says which lens it's using. If you're sending the read to advisors or referrers rather than acting on it yourself, an eyes-on version is available on request: same analysis, no pitch at the end. Just the work.
Brady walks the low-confidence findings one by one: keep, kill, or reshape. He does the final read before anything goes out. The engine does the reading and the arithmetic. The judgment about what's worth saying, and what's too thin to stand behind, is his.
A deck where every number points to the criterion behind it and every fact points to its source. If it can't show its work, it doesn't ship.
The engine reads the outside of a brand well. It's deliberately unwilling to pretend the outside is the whole story. Where it can't settle a question, it says so. And says what would.
A read that can't see internal data and pretends it can is the tell of a slop audit. Naming the limit, and what would resolve it, is the difference between a diagnosis and a horoscope.
This section exists because a numerate founder asked us, plainly: how does evidence become 55? What defines 100%? What are the weights? Fair question. The honest first answer was that the old number couldn't fully show its work. So we rebuilt it. Here's the version that can.
| Swap / onliness test | 40 · verified |
| Defensible asset vs. adjective | 70 · verified |
| Distinctiveness (recognisable as itself) | 75 · verified |
| Point-of-parity vs. point-of-difference | 45 · inferred |
| Competitive-frame correctness | 55 · verified |
The number comes from the rows. The story is written to explain it. Never a story with a number bolted on to look precise. Every point on the radar chart traces back to a table like this one.
A strategic read is a judgment; we mark it as one, with a confidence level. But every external fact (a competitor's rating, a distribution number, a quoted line of copy) lands in a dedicated appendix with its source and date. The deck is the argument; the appendix is the evidence you can click through and check.
This is the single thing that separates the read from a confident-sounding guess. Before a claim reaches a slide it carries a machine-checkable triple: the claim, its source, and its scope. An adversarial pass runs after the deck is built to catch anything that drifted from its source while being written into a headline. A number that can't name where it came from doesn't ship as a fact. It gets softened to a clearly-marked opinion, or cut.
Scope is the quiet one that matters most: a stat that's true nationally but presented as local, or true in 2024 but presented as current, is the kind of error that survives every check except this one. So it gets its own.
Every gate in the pipeline is a scar. It's there because a specific thing once went wrong. Rather than hope it wouldn't happen again, we built a pass that makes it fail loudly. A few of the real ones:
An early build pulled the top stock image for a keyword sight-unseen and put a bowl of chips on a chocolate brand's cover. Every mechanical check passed. It existed, it rendered, the aspect ratio was right. A machine that never looks at the picture can't know it's wrong. Now a vision pass looks at every image before it ships.
The engine once suggested "harmonizing" a product descriptor for consistency. The variance was legally required. A compound-coating product can't be labelled "chocolate" under Canadian rules. Recommending that to a 17-year operator would have been worse than saying nothing. There's now a two-way regulatory sweep: don't advise a violation, and don't invent a rule that isn't real.
A rigorous-looking matrix compared a founder-led brand to giants that shared its aisle. The retailer stocks both and private-labels neither. So it doesn't treat them as substitutes. A perfect analysis inside a wrong frame is still wrong. That's why competitors now run through three labeled lenses instead of one.
The scoring section above is the receipt for this one. A fair question we couldn't fully answer became the rebuild that means we can.
A tool that admits its wrong customer is worth more to its right one. This is a diagnostic read of a brand from the outside. That makes it powerful for some jobs and wrong for others.
A real system changes when the field tells it to. Most of these came from an operator or a designer catching something, and us fixing the engine, not the deck.
This page is the walkthrough. These two go further, for anyone who wants to check the rigour rather than take it on faith.
One diagram of both systems: the methodology library that holds the discipline, the engine that applies it, and the single seam between them. What crosses, what is forbidden from crossing, and which steps have a human in them.
Open the system map →The full written walkthrough, about twenty minutes. How the library maintains itself, how it reaches the engine, how scoring and the competitive lenses work, where the human sits, and the parts that are still rough.
Read the long version →Everything you just read is what stands behind the number. If that's the kind of rigor you want pointed at your brand, or you'd like to see it run on one you already know well, that's the offer.
Get your Brand Size-Up · $499This page describes the live engine. It's a moving system. The version above is where it stands today.