Get your Brand Size-Up $499
The Brand Size-Up · How it works

One word in.
A whole system after.

The entire input is a brand name. What comes back is a six-dimension read of the brand, scored against written rubrics, run through a stack of checks built to disprove it. And it's honest about what it can't see from the outside.

This is the proof of work. Maybe you're the kind of operator who wants to know how a score got to be a score. Maybe you've seen enough AI-generated "brand audits" to be suspicious on sight. Both instincts are right. This page is built for them.

The one-word-in idea

You give it a name.
That's the whole ask.

No questionnaire, no data room, no homework. A brand name is enough to start. A URL helps, if you have one.

"Acme Foods" read the public brand · score six dimensions · hunt the counter-case · compose the deck · a person signs off · delivered

Everything below is that arrow, slowed down. None of it is "ask a model what it thinks and format the answer." Each stage has a defined input, a defined output, and a defined way to be wrong on purpose. Most AI tools skip that last part. It's how the wrong thing gets caught before it reaches you.

The mental model

What actually happens

The real pipeline, not an idealized diagram. Research and drafting are the bulk of the work; judgment stays with people, and with gates that exist only because something once slipped past.

1 Inputs

Two feeds. The first is the brand's public signals: homepage, packaging, social, reviews, category forums, retailer listings, search presence. Read directly, never guessed from a footer. The second is the Design-Strategy OS: 276 pages of concepts, methods and strategy, wired together by more than 6,400 internal links, built from 75 sources and the field's best thinkers (Rumelt, Dunford, Christensen, Neumeier, Wheeler, Sharp, Aaker, Binet & Field), sharpened by 17 years building brands. It's what tells the engine what "good" looks like.

public signals only, at intake276 pages · 6,400+ connections

How the engine is checked

Three checks, running on different schedules. First, the build itself. Around 290 automated checks run every time a deck is assembled. If a number can't be traced to where it came from, if a quote isn't actually found at the source it's credited to, or if a caveat gets dropped between the appendix and the slide, the build stops and the file never gets written. Second, the library. A regression test of 67 fixed questions runs against the knowledge base after every ingest, so new reading can't quietly break what was already there. Third, the scoring. Brands get re-scored blind and the results compared, which is how the bands stay honest. Six brands the last time. That one runs at intervals, not continuously.

2 The engine: the six-dimension read

The brand is scored on Clarity, Consistency, Differentiation, Resonance, Availability, and Credibility. Each dimension is broken into written criteria scored against anchored rubrics. The score is computed from cited evidence first, then explained. Not the other way around.

anchored rubricsaudience & jobs-to-be-done passproduct-experience read

3 The adversarial gates: hunt the counter-case, name the blind spots

Before anything reaches a slide it runs a stack of passes built to disprove the read. An assumption audit flags what public signals can't see, and what could already have killed the idea inside the company. A disconfirming search checks any fact an insight leans on. A senior-strategist critique. A tournament that makes the competing insights fight until one survives, so the headline finding is the one that beat the others rather than the one written first. A fact-check against primary sources. A claim-source-scope trace. And a two-way regulatory sweep, which checks both that a recommendation doesn't break a rule and that any rule we assert is real. Then, once the deck is actually built, it goes to a red-team of independent models from different labs, and anything two of them agree on comes back to be answered before it ships.

assumption auditdisconfirming searchinsight tournamentcross-model red-teamfact-checkclaim · source · scoperegulatory sweep

4 The deck composer

The read becomes a deck. Competitors run through three labeled lenses: the shelf set you're actually chosen against, the size-and-structure peers you resemble (so we benchmark what's achievable, not what a giant does), and a scan of anyone running your same strategic move. Every comparison says which lens it's using. If you're sending the read to advisors or referrers rather than acting on it yourself, an eyes-on version is available on request: same analysis, no pitch at the end. Just the work.

three-lens competitive seteyes-on version on request

5 The editorial pass: a person, named

Brady walks the low-confidence findings one by one: keep, kill, or reshape. He does the final read before anything goes out. The engine does the reading and the arithmetic. The judgment about what's worth saying, and what's too thin to stand behind, is his.

keep / kill / reshape walkfinal human sign-off

6 Delivered

A deck where every number points to the criterion behind it and every fact points to its source. If it can't show its work, it doesn't ship.

Honest limits

What it reads, and what it won't conclude

The engine reads the outside of a brand well. It's deliberately unwilling to pretend the outside is the whole story. Where it can't settle a question, it says so. And says what would.

What it reads

  • The homepage, product and pricing pages, the five-second first-read test
  • Packaging, on-pack claims, certifications, label copy
  • Social presence and cadence, verified follower counts, the gap between what a brand says and what it posts
  • Reviews and category forums: the actual experience record, not the star average
  • Retailer listings, distribution footprint, search presence, third-party validation
  • The competitive shelf, the structural peers, and who else is running the same play

What it won't conclude from the outside alone

  • Why sales actually move.
    Would take: internal sell-through, velocity, and margin data.
  • Whether a positioning wins in-market.
    Would take: a regional A/B or a real market test. Not a confident guess.
  • What the brand has already tried.
    Would take: prior research and founder context. The audit flags the assumption instead of burying it.
  • How it compares in one specific region.
    Would take: market-matched comparisons at the same scope, not a national stat applied locally.
  • Contracts, mandates, and constraints inside the company.
    Would take: a conversation. These are exactly what the assumption audit surfaces as questions.

A read that can't see internal data and pretends it can is the tell of a slop audit. Naming the limit, and what would resolve it, is the difference between a diagnosis and a horoscope.

The scoring rubric

How a score becomes a score

This section exists because a numerate founder asked us, plainly: how does evidence become 55? What defines 100%? What are the weights? Fair question. The honest first answer was that the old number couldn't fully show its work. So we rebuilt it. Here's the version that can.

ClarityDoes the brand agree with itself about what it is and who it's for?
ConsistencyDoes it look, sound, and behave like itself everywhere?
DifferentiationCould its positioning be swapped with a rival's and no one notice?
ResonanceDoes it speak in a voice the audience recognizes as theirs?
AvailabilityEasy to bring to mind at the moment of choice, and easy to act on?
CredibilityDoes its proof earn the trust it's asking for?
  1. Each dimension is broken into four or five written criteria, each grounded in named brand-strategy canon, not invented on the spot.
  2. Every criterion is scored 0–100 against a shared anchor scale (100 fully leveraged, 75 mostly, 50 partial, 25 weak, 0 absent), against cited evidence, tagged for whether that evidence was verified, inferred, or absent.
  3. The dimension score is the average of its criteria. Plain arithmetic, equal weight. No hidden knob to nudge a number.
  4. A band leads the result: Superb, Excellent, Strong, Solid, or Room for growth. The percentage gets demoted to supporting detail, with a confidence grade attached.
  5. Confidence is derived, not asserted: a function of how much of the score rests on cited evidence, with one line on what would raise it.
Superb · 68–100Excellent · 56–67Strong · 42–55Solid · 28–41Room for growth · <28

One dimension, all the way through (illustrative)

Swap / onliness test40 · verified
Defensible asset vs. adjective70 · verified
Distinctiveness (recognisable as itself)75 · verified
Point-of-parity vs. point-of-difference45 · inferred
Competitive-frame correctness55 · verified
Differentiation: Excellent · 57% · mean of 5 criteria · confidence HIGH (4/5 evidence-cited) · raise by: a competitor search confirming no rival owns the same format.

The number comes from the rows. The story is written to explain it. Never a story with a number bolted on to look precise. Every point on the radar chart traces back to a table like this one.

Traceability

Every fact carries a receipt

A strategic read is a judgment; we mark it as one, with a confidence level. But every external fact (a competitor's rating, a distribution number, a quoted line of copy) lands in a dedicated appendix with its source and date. The deck is the argument; the appendix is the evidence you can click through and check.

This is the single thing that separates the read from a confident-sounding guess. Before a claim reaches a slide it carries a machine-checkable triple: the claim, its source, and its scope. An adversarial pass runs after the deck is built to catch anything that drifted from its source while being written into a headline. A number that can't name where it came from doesn't ship as a fact. It gets softened to a clearly-marked opinion, or cut.

Scope is the quiet one that matters most: a stat that's true nationally but presented as local, or true in 2024 but presented as current, is the kind of error that survives every check except this one. So it gets its own.

Appendix: how a claim appears
"Rated 4.0/5 across 50 reviews"
amazon.com · 2026-07
verified
Distribution: 6 US states + national club
company post · 2026-05
verified
Category shifting toward reduced-sugar
trade report · 2026-Q1
directional
"First to invent the category"
founder-stated · unverified
directional
Why the gates exist

The system catching itself

Every gate in the pipeline is a scar. It's there because a specific thing once went wrong. Rather than hope it wouldn't happen again, we built a pass that makes it fail loudly. A few of the real ones:

The potato chips on the chocolate deck.

An early build pulled the top stock image for a keyword sight-unseen and put a bowl of chips on a chocolate brand's cover. Every mechanical check passed. It existed, it rendered, the aspect ratio was right. A machine that never looks at the picture can't know it's wrong. Now a vision pass looks at every image before it ships.

The regulatory near-miss.

The engine once suggested "harmonizing" a product descriptor for consistency. The variance was legally required. A compound-coating product can't be labelled "chocolate" under Canadian rules. Recommending that to a 17-year operator would have been worse than saying nothing. There's now a two-way regulatory sweep: don't advise a violation, and don't invent a rule that isn't real.

The wrong competitors.

A rigorous-looking matrix compared a founder-led brand to giants that shared its aisle. The retailer stocks both and private-labels neither. So it doesn't treat them as substitutes. A perfect analysis inside a wrong frame is still wrong. That's why competitors now run through three labeled lenses instead of one.

The number that couldn't show its work.

The scoring section above is the receipt for this one. A fair question we couldn't fully answer became the rebuild that means we can.

Fit

Who it's not for

A tool that admits its wrong customer is worth more to its right one. This is a diagnostic read of a brand from the outside. That makes it powerful for some jobs and wrong for others.

Not the right call if…

  • You need the answer to turn on internal data: sell-through, margin, CAC, cohort behavior. The read can't see it, and won't pretend to.
  • You want a logo or a visual identity produced. This diagnoses the brand; it doesn't design one.
  • You want a number that flatters. The bands lead, and "Room for growth" means there's real work to do.
  • You want it to replace the strategist. The engine does the reading; a person still decides what's worth acting on.

Built for…

  • An operator who senses the gap but can't name it from the inside, and wants it named with evidence.
  • A skeptic who's been burned by generic AI output and wants to see the work before they trust the read.
  • Anyone who'd rather be told "here's what we can't see, and here's what would settle it" than be sold a confident guess.
  • A first, cheap, fast look that tells you whether the deeper work is worth doing.
Engine version & changelog

It has a version number
for a reason

A real system changes when the field tells it to. Most of these came from an operator or a designer catching something, and us fixing the engine, not the deck.

v7.4scoring integrity
Field feedback
Dimension scores rebuilt onto anchored, evidence-cited rubrics: computed before narrated, band-led, with derived confidence. Prompted by a numerate founder asking how a score is derived.
competitive v2.2reference-class rule
Field feedback
Three labeled competitor lenses: shelf set, size/structure peers, same-play scan. Replacing a single set. Prompted by a founder who wanted to be benchmarked against real peers, not giants.
v7.0–7.3tournament & guards
An insight tournament that weighs central equity against novelty; an attribution test; single-case causal humility; and a guard so a deliberate channel bet changes the framing of a score, never the score itself.
image gatethe potato-chip fix
Field feedback
A vision pass that looks at every shipped image against the brand's product, category, and channel. Mechanical checks can't see a picture.
regulatory sweeptwo-way
Field feedback
Never recommend an action offside the client's regulator, and never assert a rule that isn't real. Checked against primary sources.
v6.1brand-type read
Availability split into mental availability (every brand) and a getting-found read tuned by type. So a software brand is never marked down for lacking retail shelf space.
If you want the long version

Two ways to go deeper

This page is the walkthrough. These two go further, for anyone who wants to check the rigour rather than take it on faith.

The system map

One diagram of both systems: the methodology library that holds the discipline, the engine that applies it, and the single seam between them. What crosses, what is forbidden from crossing, and which steps have a human in them.

Open the system map →

Inside the engine

The full written walkthrough, about twenty minutes. How the library maintains itself, how it reaches the engine, how scoring and the competitive lenses work, where the human sits, and the parts that are still rough.

Read the long version →

The whole input was a name.

Everything you just read is what stands behind the number. If that's the kind of rigor you want pointed at your brand, or you'd like to see it run on one you already know well, that's the offer.

Get your Brand Size-Up · $499

This page describes the live engine. It's a moving system. The version above is where it stands today.