Quick Answer
Commercial judgment is the founding skill of deciding what is worth building, who it is for, and what they will pay — judgment about the market rather than about the code. YC cohort data suggests teams carrying that judgment from formation fail less often. If your team is already formed, you build commercial judgment through practice: buyer calls, recorded objections, price tests and published writing.
You can write the system. You cannot always tell which system anyone will pay for. That gap is the expensive one, and it does not close itself as the code gets better.
It has also become harder to ignore. When building was the scarce thing, a hard engineering problem bought a team time to work the market out later. That time is shorter now, and the excuse it paid for has gone with it.
Most advice here stops at hire a commercial cofounder, which is not available to a team that formed two years ago. This post covers what the Y Combinator survival numbers actually support, the five habits that build commercial judgment inside an existing technical team, and the cases where the bottleneck is somewhere else.
What Commercial Judgment Actually Decides in a Startup
Commercial judgment answers the questions engineering cannot: which problem you take on, who you take it on for, what it is worth to them, and in what words you describe it. Each one is a bet placed before the evidence arrives.
Four decisions sit inside it, and a founding team makes all four in its first year:
- Which problem is worth solving, out of the several you could solve well.
- Which buyer feels that problem badly enough to move budget this quarter.
- What the fix is worth to that buyer, in money rather than in praise.
- Which words make that buyer recognize their own situation on your page.
Steve Blank makes the same point about the first two. His version of the mistake is a founder assuming the product solves someone's problem without ever understanding what problems the customers really had. His remedy, customer development, starts outside the office rather than in a planning document.
The fourth decision is the one technical teams underrate most. A product that solves a real problem in vocabulary nobody uses is indistinguishable, from the buyer's side, from a product that solves nothing.
How to Build Commercial Judgment Without a Commercial Cofounder
You build commercial judgment the way you built engineering judgment: by accumulating exposure, keeping records, and being wrong in public often enough to get calibrated. The five habits below are the practice version of a commercial cofounder, and every one of them is available to a team that has already formed.
None of them requires a budget. All of them require hours the founders currently spend on the product.
1. Log Hours on Calls You Did Not Run
Commercial judgment is accumulated exposure, so the first habit is attendance. Sit on every sales and support call you can for a quarter, including the ones another founder is running, and take notes with a point of view rather than a transcript.
Paul Graham puts the reason plainly: “The most common unscalable thing founders have to do at the start is to recruit users manually.” The founder who recruited the user hears why they said yes.
A working target is ten buyer conversations a month in one document, each with the date, the role, and the sentence that changed the temperature of the call. You know it is working when you can predict the objection before it arrives and you are wrong less often each month.
The common mistake is attending as an observer. Silence produces a transcript, and a point of view produces judgment. Founder-led sales is the training for this, not a stopgap until you can hire.
2. Write Down the Words Buyers Actually Use
Most early teams have a vocabulary problem before they have a positioning problem. In our work with technical founders, the gap between what the founder calls the product and what the buyer calls it is usually the first thing worth fixing, and it is free to fix.
If prospects call it a reporting tool and you call it an operating layer, the buyer is not going to change vocabulary to meet you. Keep a two-column file: their words on the left, yours on the right, dated so you can see the drift.
Test it by reading your homepage headline to a buyer and asking them to say back what the product does. You know it is working when their answer and your headline use the same nouns.
The mistake is filing this under copywriting. Positioning is a hypothesis, and the vocabulary is how you run the test.
3. Test the Price Before the Product Feels Done
Willingness to pay is the fastest read on whether the problem is real, and it is available months before the product feels finished. Ask what budget the problem already sits in, what the buyer spends on the current workaround, and what a month of failure costs them.
Then name a number out loud and watch the pause. A price that produces no friction at all is usually a price nobody was going to pay either way.
You know the test worked when two unrelated buyers push back on the same line item, because that is a signal about the packaging rather than about the number. The mistake we see most often is setting price by intuition, then leaving it untouched until late-stage churn explains it for you.
4. Publish the Argument and Watch What Lands
Publishing is the cheapest laboratory an early team has. Write the argument you would make on a call, put it where your buyers already are, and read what comes back: silence, agreement, or a specific objection you had not heard before.
Google's own guidance for creators asks a question a founder can answer and an outsourced writer usually cannot: does the content “clearly demonstrate first-hand expertise and a depth of knowledge”? Most content marketing for tech companies starts here, because the words you publish can only be as good as the words you collected.
Better Marketing builds founder brands on LinkedIn out of that raw material rather than a content calendar, and an AI SEO service has nothing to make citable until a founder has written something only they could write.
You know it is working when replies quote a specific line back at you. The mistake is publishing conclusions with the evidence stripped out, which reads as confident and teaches you nothing.
5. Review the Losses, Not Only the Wins
Run a monthly loss review: every deal that did not close, the reason the buyer gave, and the reason you privately believe. The two columns disagree often, and the disagreement is where commercial judgment is made.
Wins teach you what a well-matched buyer sounds like. Losses teach you the shape of the boundary, which is the more useful map when you are still deciding who to serve.
You know it is working when the same stated reason appears three months running and you can name the change that would remove it. The mistake is reviewing losses only after a bad quarter, by which point the pattern has already cost you the quarter.
What the YC Cohort Numbers Do and Do Not Prove
The numbers lean one way and stop short of proving it. In an analysis of 4,360 startups, Jenna Hermann found that across the 2019 to 2023 Y Combinator cohorts, 8.1 percent of companies with a commercial founder became inactive against 16.9 percent of the rest. Among multi-founder teams only, the split was 8.0 percent against 18.3 percent.
Then the caveats, which are hers as much as ours:
- The commercial-founder sample is 62 companies.
- The p-value runs 0.06 to 0.08, narrowly missing the conventional significance threshold.
- Y Combinator is one sample of venture-backed companies, not the market.
- Staying active measures survival, not revenue quality or category leadership.
Her own reading is a claim about scarcity rather than a promise about outcomes: “The technical barrier has collapsed but the commercial moat remains strong.” That is the sentence to argue with, not the percentages.
So the honest use of this data is not as proof. Ask what it costs you to act as though commercial judgment were the scarce thing, and what it costs to be wrong about that, and the second number is almost always larger.
Founder Judgment Compared With a First Revenue Hire
A hire executes judgment far more often than they create it. That distinction is worth holding on to the next time somebody asks why a founder is still running discovery calls personally.
| Decision | Founder With Commercial Judgment | First Revenue Hire |
|---|---|---|
| Which buyer to chase | Can change the target after one call | Works the list they were given |
| What to charge | Can move the price and the packaging | Discounts inside a set band |
| What ships next | Feeds the objection into the roadmap | Files the objection as a request |
| When to stop | Can rule a whole segment out | Keeps the pipeline moving |
A B2B demand generation agency, a fractional revenue leader and a first account executive all work from a thesis about the buyer. If nobody has written that thesis down, the hire writes it for you, and you spend the next year managing a guess you did not make.
When Commercial Judgment Is Not the Bottleneck
Sometimes the market is already understood and the constraint is reach. If you can name the buyer, the trigger, the objection and the price without checking your notes, another ten calls will teach you little. The question has become how many of those buyers ever hear from you.
That is where a B2B demand generation agency earns its fee, because it is scaling a thesis rather than searching for one. An AI SEO service works the same half of the problem, making answers you already hold findable on Google and inside AI answers. The threshold test for a first marketing hire asks the same question in different clothes.
Two other cases sit outside all of this. When the product does not yet do the thing it promises, no amount of buyer contact will rescue it, and in a regulated category where access is gated by approvals, the sequencing problem is procurement rather than vocabulary.
Commercial judgment is also no substitute for capacity. A founder who has the judgment and no hours left has a staffing problem instead. Content marketing for tech companies is one of the first jobs worth delegating, once the thesis is written down and product-market fit is close enough to argue about.
Turn Commercial Judgment Into Your Next Ninety Days
Pick one quarter and run it as an apprenticeship. Ten buyer conversations a month, a two-column vocabulary file, one price test, one published argument a week, and a loss review on the last Friday.
At the end you hold a written thesis about your market that no hire could have handed you, plus a list of the parts you got wrong. Only then is a B2B demand generation agency worth hiring, which is how Better Marketing works: we ask founders for their loss notes before we ask about budget.
