The Bullshit Index is a 0 to 100 score of the gap between what a product claims and what the evidence supports. It is not a startup viability score: a dumb business can be honestly described, and a great product can be marketed with absurd claims. Higher means more bullshit.
The Build Signal is a separate 0 to 100 score for whether there is enough real signal (problem, product, differentiation, path to market) to keep building. It is not a probability of success. The two scores are never averaged, because your code can be strong while your story is bullshit, and your story can be great while there is nothing under it.
We built this pipeline in 2026 and ran it on our own product before charging anyone: the meter scored bullshitmeter.dev at Bullshit Index 41, Build Signal 72, verdict FIX THIS SHIT FIRST, and that card is printed on the homepage exactly as it came out.
A full report runs these stages in order. It takes a few minutes, not seconds, because gathering real external evidence is actual work.
These tables are rendered from the same constants the scorer runs (rubric version 2026-07-v1). A score you cannot audit is a vibe with a decimal point.
| Dimension | Weight | The question it answers |
|---|---|---|
| Claim inflation | 25% | Are the words materially bigger than the demonstrated reality? |
| Evidence deficit | 25% | How many important claims rely on assertion rather than proof? |
| Differentiation inflation | 20% | Is the claimed moat actually different? |
| Market assumption load | 15% | How much of the business case depends on unverified beliefs? |
| Execution gap | 15% | How far is the current artifact from the product being described? |
| Dimension | Weight | The question it answers |
|---|---|---|
| Problem reality | 18% | Is there credible evidence the problem exists and matters? |
| Pain and urgency | 15% | Does the user have a reason to solve it now? |
| Product usefulness | 17% | Does the artifact appear capable of producing a meaningful outcome? |
| Differentiation | 12% | Is there a meaningful reason to choose this over the status quo? |
| Distribution plausibility | 15% | Can this builder plausibly reach the intended user? |
| Monetization fit | 8% | Is there a believable path to capturing value if this is intended as a business? |
| Focus | 10% | Is the product understandable and narrow enough to execute? |
| Evidence quality | 5% | How much of the assessment rests on observable evidence? |
| 0 to 19 | Mostly grounded |
| 20 to 39 | Some marketing calories |
| 40 to 59 | Noticeable bullshit |
| 60 to 79 | The story is outrunning the product |
| 80 to 100 | We have entered the fog machine |
| 0 to 19 | Weak signal |
| 20 to 39 | Interesting but not enough |
| 40 to 59 | Worth a targeted test |
| 60 to 79 | Strong enough to keep building |
| 80 to 100 | Strong signal, execution now matters |
| KEEP BUILDING | The evidence supports continuing. Execution is the constraint now. |
| FIX THIS SHIT FIRST | Something specific and fixable is undermining an otherwise real project. |
| NARROW IT | The product is trying to be too many things; the evidence supports one of them. |
| PIVOT AROUND THE GOOD PART | The strongest signal is not where the pitch points. Follow the signal. |
| COOL PROJECT, BAD BUSINESS | Real craft, no believable path to capturing value. Ship it as what it is. |
| KILL IT | Weak signal, high confidence, no hidden gold, and a structural problem more coding will not solve. Used only when all of those hold. |
| NOT ENOUGH EVIDENCE | The honest verdict when the evidence cannot justify a strong recommendation. |
Every report also carries a Confidence rating (High, Medium, Low) based on how much real evidence the audit had to work with, not on model confidence theater, plus exactly three Monday actions, one to three kill criteria, and an honest list of what it could not verify.
It can inspect what you provide and research what is public. It cannot interview your customers, see private usage data, or predict the future. Findings it cannot support get killed by the critic; questions it cannot answer are listed as unverified instead of guessed at.