How funded AI and SaaS startups' websites score: 2026 benchmark
We scored 100 funded AI and SaaS homepages against the 12 questions in our scorecard. The average is 14.1 out of 24. Most sites explain the product reasonably well and almost none prove it.
Take the 3-min scorecardA typical funded AI or SaaS homepage in 2026 scores 14.1 out of 24 (59%). The weakest area is credibility: 53% show no customer logos, testimonials or case studies, and only 21% pair a named customer with a number. Only 21% say clearly what they do in the first screen on both desktop and phone. Later-stage companies score higher (15.9 vs 12.8 at seed) because they have more proof and better measurement. They are not clearer.
- Average score
- 14.1/24
- No customer proof
- 53%
- Pass the five-second test on both widths
- 21%
- Reach the top tier (18+)
- 13%
Eight things the data says.
Percentages are of all 100 sites unless a stage is named. Seed n = 60, Series A and later n = 40.
- 14.1 / 24
The average funded AI or SaaS homepage scores 14.1 out of 24.
The median is 15 and the range is 4 to 19. No site reached 20. Thirteen of the 100 cleared 18, the top tier on our scorecard, and 8 scored under 10.
- 53%
Proof is the biggest gap.
Credibility is the weakest pillar at 3.16 out of 6. Half the sample (53%) shows no customer logos, testimonials or case studies at all. Among seed-stage sites it is 77%, and more of them put an investor badge in the first screen (58%) than show any customer proof (23%).
- 21%
Only one in five passes the five-second test on both desktop and phone.
35% of desktop first screens say plainly what the product does. On a 390px phone that drops to 21%, and 15% of sites score lower on mobile than on desktop.
- 26%
Few homepages name a buyer or a difference up front.
26% name a specific role, team or industry in the first screen, and 22% state a specific difference there (a number, a named alternative or a mechanism). Most say who it's for and why it's better somewhere lower on the page, or only in generic words.
- +3.1
Later-stage sites score higher on proof and measurement, not clarity.
Series A and later sites average 15.9 against 12.8 at seed. The gap is credibility (4.0 vs 2.6 out of 6) and conversion (4.0 vs 3.2). Clarity barely moves (3.6 vs 3.4), and fewer later-stage sites pass the five-second test at both widths (15% vs 25%).
- 42%
Fewer than half show the product in the first screen.
42% show product UI, a product video or a live demo in the first desktop screen and 35% on mobile. 19% show no product UI anywhere on the homepage, relying on illustration, abstract art or stock imagery instead.
- 45%
Desktop first screens compete with themselves.
45% put three or more buttons in the first desktop screen (17% on mobile, where the navigation collapses into a menu). Only 27% have one clear primary action at the top and another in the last section on both widths.
- 23%
About a quarter load no analytics at all.
23% load no analytics tag before consent (33% at seed, 8% at Series A and later). 43% run both analytics and a conversion, ads or CRM tag, which is what it takes to know which pages bring in qualified calls.
Where the points go.
Hover a bar for its exact value. Every chart has a table view underneath.
Table view
| Score | Sites |
|---|---|
| 4 | 1 |
| 7 | 2 |
| 8 | 2 |
| 9 | 3 |
| 10 | 6 |
| 11 | 6 |
| 12 | 5 |
| 13 | 16 |
| 14 | 7 |
| 15 | 15 |
| 16 | 14 |
| 17 | 10 |
| 18 | 10 |
| 19 | 3 |
Table view
| Pillar | All | Seed | Series A+ |
|---|---|---|---|
| Clarity | 3.46 | 3.38 | 3.58 |
| Credibility | 3.16 | 2.60 | 4.00 |
| Consistency | 3.98 | 3.70 | 4.40 |
| Conversion | 3.49 | 3.17 | 3.98 |
Table view
| Criterion | Pillar | 0 | 1 | 2 |
|---|---|---|---|---|
| Says what it does in five seconds | Clarity | 6% | 73% | 21% |
| Speaks to one specific buyer | Clarity | 11% | 63% | 26% |
| States why it's different | Clarity | 6% | 72% | 22% |
| Looks established | Credibility | 3% | 16% | 81% |
| Shows specific customer results | Credibility | 36% | 43% | 21% |
| Shows proof that matches buyers | Credibility | 53% | 41% | 6% |
| Brand carried into shared surfaces | Consistency | 3% | 21% | 76% |
| Tight type and colour system | Consistency | 11% | 42% | 47% |
| Ships new content | Consistency | 22% | 67% | 11% |
| One clear next step | Conversion | 14% | 59% | 27% |
| Shows the product | Conversion | 19% | 46% | 35% |
| Measures what converts | Conversion | 23% | 34% | 43% |
Table view
| Check | Desktop | Mobile |
|---|---|---|
| Passes the five-second test | 35% | 21% |
| Product UI or video in the first screen | 42% | 35% |
| Action CTA in the first screen | 86% | 79% |
| Three or more buttons competing in the first screen | 45% | 17% |
No site overflowed horizontally at 390px and 96% kept the H1 in the first mobile screen. Page weight is the bigger performance issue: the median homepage transfers 2.3 MB on load and 24% transfer more than 5 MB.
How we scored 100 homepages.
Captured on 2 October 2026. The rubric was written before any site was scored.
Sample
Sixty companies come from Y Combinator's public company directory: the Winter, Spring and Summer 2026 batches, filtered to B2B companies tagged AI, SaaS, developer tools, enterprise software or workflow automation, with hardware, robotics and biotech removed (147 companies). Forty come from the public a16z portfolio page: companies listed at the Venture or Growth stage, still private, and described as AI or software, with crypto, consumer, bio and defense removed (150 companies). We shuffled each pool with a fixed seed and took companies in order.
A site was skipped and replaced by the next one if it failed to load, blocked automated browsers, redirected to an acquirer, or turned out to be a services firm, a consumer product or hardware. That happened 16 times (2 in the YC pool, 14 in the a16z pool). "Seed" means a 2026 YC batch. "Series A and later" means a16z lists the company at Venture or Growth, which is the stage when a16z invested.
Capture
Each homepage was loaded in headless Microsoft Edge at 1440 by 900 and at 390 by 844 with a mobile user agent. We waited 4.5 seconds, hid cookie banners without accepting them, took the first-screen screenshot, then scrolled the page to load lazy content. A script recorded the H1 and largest headline, the buttons in the first screen, video and canvas elements, logo-like images, numbers and quotes in the text, fonts and button colours, metadata, dated items, page weight, horizontal overflow and the analytics and marketing tags requested before consent.
Scoring
Each of the 12 criteria scores 0, 1 or 2, for a total out of 24. Where a criterion depends on the first screen, we scored both widths and kept the lower score. When unsure, we gave the lower score.
Judgments were assisted by a calibrated classification model (TypeSafe Jev) and spot-checked by hand. The model read each site's captured text and signals and returned an expected score between 0 and 2 for each question. We converted it to a level with fixed thresholds: 1.67 or more is a 2, 0.67 or more is a 1, anything lower is a 0. The model scores clarity, buyer, difference, results, proof and next step. Two criteria are visual, so we scored them by eye from the screenshots: "looks established" and "shows the product". The four proxy criteria use fixed rules on the captured signals. If the headline was not visible five seconds after load, the five-second test and next-step criteria scored 0 at that width (one site).
We scored 20 sites by hand, blind to the model's answers, across the six model-scored criteria: 160 judgments. The model matched the hand score exactly in 68% of them and was within one level in all of them. Where they differed, the model was usually lower (41 times lower, 10 times higher), so treat the clarity, credibility and conversion figures as conservative.
The rubric
- Says what it does in five seconds. The headline, or headline plus one-line subhead, says what the product does at 1440 and at 390.
- Speaks to one specific buyer. The first screen names a role, team or industry. Generic words (teams, businesses) score 1.
- States why it's different. The first screen states a quantified claim, a named alternative or a concrete mechanism. Lower on the page scores 1.
- Looks established. Custom identity and no defects at either width. Template-generic or stock-led scores 1; broken, empty or overflowing scores 0.
- Shows specific customer results. A named customer paired with a number. Unattributed numbers or a named quote without one score 1.
- Shows proof that matches buyers. Customer logos and an attributed testimonial or case study. One of the two scores 1. Investor badges alone score 0.
- Brand carried into shared surfaces. Proxy: og:image, favicon and a linked social profile. All three score 2, two score 1.
- Tight type and colour system. Proxy: two or fewer font families and three or fewer filled-button colours.
- Ships new content. Proxy: a dated post, release or announcement from the last 90 days. A blog or changelog link alone scores 1.
- One clear next step. One primary action CTA in the first screen and another in the last section, at both widths.
- Shows the product. Product UI, a product video or a live demo in the first screen, at both widths. Lower on the page scores 1.
- Measures what converts. Proxy: an analytics tag and a conversion, ads or CRM tag load before any consent. One kind scores 1.
Limitations
- Homepage only. Four scorecard questions are about internal practice (brand guidelines, shipping speed, attribution), which a homepage can't show, so those scores are proxies.
- One capture on one day from one location, on a fast connection. We report page weight, not real-user load times.
- Analytics were counted only if they loaded before consent, so sites that gate tags behind a cookie banner are undercounted.
- The sample is random within two investor portfolios, not all funded startups. YC 2026 companies are months old. With 100 sites, a percentage near 50% carries a margin of about ±10 points.
- Judgments come from one model and one reviewer, with no second human rater. Scroll-driven pages can hide content from captures, so we used scroll-by-scroll frames where a full-page shot failed.
See how your homepage compares.
The same 12 questions are in our free scorecard. Answer them in three minutes and you'll get your score by pillar and the three fixes worth doing first. Compare your total with the 14.1 average here and the 18+ top tier.
Get the data
Is your brand costing you deals?
Answer 12 quick questions and get a score across clarity, credibility, consistency and conversion, plus the three fixes that matter most for your team.