Bullshit Meter
Account
Methodology

How the scoring works.

The Bullshit Index is a 0 to 100 score of the gap between what a product claims and what the evidence supports. It is not a startup viability score: a dumb business can be honestly described, and a great product can be marketed with absurd claims. Higher means more bullshit.

The Build Signal is a separate 0 to 100 score for whether there is enough real signal (problem, product, differentiation, path to market) to keep building. It is not a probability of success. The two scores are never averaged, because your code can be strong while your story is bullshit, and your story can be great while there is nothing under it.

We built this pipeline in 2026 and ran it on our own product before charging anyone: the meter scored bullshitmeter.dev at Bullshit Index 41, Build Signal 72, verdict FIX THIS SHIT FIRST, and that card is printed on the homepage exactly as it came out.

The eight-stage pipeline

A full report runs these stages in order. It takes a few minutes, not seconds, because gathering real external evidence is actual work.

  1. 01Normalize. Every source you provide (repo files, pasted pitch, docs, site) is read and normalized into one evidence base. Secrets are redacted before anything reaches a model.
  2. 02Claim extraction. Every claim you make, out loud or by implication, goes into a claim ledger. Each one will be checked or explicitly marked unverifiable.
  3. 03Evidence planning. The auditor decides what would prove or disprove each claim and where to look: your artifact, the live market, or both.
  4. 04Artifact evidence. It opens the files. Up to 25 real files from your repo are read against the ledger, so the code either backs the story or it does not.
  5. 05External research. Live market checks with mandatory counter-searches built to prove you wrong, across up to 15 external sources. Enthusiasm is not evidence.
  6. 06Rubric scoring. A model scores each rubric dimension against the evidence, with the scale anchors pinned. The weighted sums are computed in code, never by the model, and the rubric version is stored with every report.
  7. 07Adversarial critic. A separate critic hunts for findings that are unfair, unsupported, or overreaching, and kills them before the verdict lands.
  8. 08Synthesis and verdict. A verdict arbiter selects one of seven verdicts from the scores, claim statuses, and surviving findings. The report writer must present that judgment and is not allowed to change it.

The exact rubric, weights included

These tables are rendered from the same constants the scorer runs (rubric version 2026-07-v1). A score you cannot audit is a vibe with a decimal point.

Bullshit Index dimensions

DimensionWeightThe question it answers
Claim inflation25%Are the words materially bigger than the demonstrated reality?
Evidence deficit25%How many important claims rely on assertion rather than proof?
Differentiation inflation20%Is the claimed moat actually different?
Market assumption load15%How much of the business case depends on unverified beliefs?
Execution gap15%How far is the current artifact from the product being described?

Build Signal dimensions

DimensionWeightThe question it answers
Problem reality18%Is there credible evidence the problem exists and matters?
Pain and urgency15%Does the user have a reason to solve it now?
Product usefulness17%Does the artifact appear capable of producing a meaningful outcome?
Differentiation12%Is there a meaningful reason to choose this over the status quo?
Distribution plausibility15%Can this builder plausibly reach the intended user?
Monetization fit8%Is there a believable path to capturing value if this is intended as a business?
Focus10%Is the product understandable and narrow enough to execute?
Evidence quality5%How much of the assessment rests on observable evidence?

Reading the scores

Bullshit Index bands

0 to 19Mostly grounded
20 to 39Some marketing calories
40 to 59Noticeable bullshit
60 to 79The story is outrunning the product
80 to 100We have entered the fog machine

Build Signal bands

0 to 19Weak signal
20 to 39Interesting but not enough
40 to 59Worth a targeted test
60 to 79Strong enough to keep building
80 to 100Strong signal, execution now matters

The seven verdicts

KEEP BUILDINGThe evidence supports continuing. Execution is the constraint now.
FIX THIS SHIT FIRSTSomething specific and fixable is undermining an otherwise real project.
NARROW ITThe product is trying to be too many things; the evidence supports one of them.
PIVOT AROUND THE GOOD PARTThe strongest signal is not where the pitch points. Follow the signal.
COOL PROJECT, BAD BUSINESSReal craft, no believable path to capturing value. Ship it as what it is.
KILL ITWeak signal, high confidence, no hidden gold, and a structural problem more coding will not solve. Used only when all of those hold.
NOT ENOUGH EVIDENCEThe honest verdict when the evidence cannot justify a strong recommendation.

Every report also carries a Confidence rating (High, Medium, Low) based on how much real evidence the audit had to work with, not on model confidence theater, plus exactly three Monday actions, one to three kill criteria, and an honest list of what it could not verify.

What the meter cannot know

It can inspect what you provide and research what is public. It cannot interview your customers, see private usage data, or predict the future. Findings it cannot support get killed by the critic; questions it cannot answer are listed as unverified instead of guessed at.

Run the free Sniff TestGlossaryFAQ