How the ranking works
Every number on this site comes from a Jev answer or a weight in our code. Jev is TypeSafe AI's System One model: you send it a state and typed questions, it returns typed answers with probabilities. It does not generate text. There is no language model anywhere in this product, and nothing you read here was written by one.
Jev calls made
118Input tokens
136,558Scoring version
v1Ranking version
v11. The score, 0 to 100
When a product is submitted we fetch its landing page, reduce it to text, and send one request to Jev containing the product fields and every question below. They are all answered in parallel, in that one call. The weights are ours and live in code, not in the model.
- problem_clarityweight 22.2 . score
How clearly does the product text state the specific problem it solves?
- 0No problem stated
- 1Problem implied but vague
- 2Specific problem named
- 3Specific problem named with who has it and when
- audience_specificweight 16.7 . noul
Does the product text name a specific type of user or business as its target?
yesA concrete role, industry, or situation is named (e.g. freelance designers, Shopify stores)noNo target named, or generic words like everyone, businesses, teams, users - differentiationweight 22.2 . score
How clearly does the product text explain why to choose it over an existing alternative?
- 0No alternative or difference mentioned
- 1Claims to be better with no specifics
- 2Names a concrete difference from alternatives
- 3Names the alternative and a concrete difference
- pricing_clearweight 11.1 . noul
Does the landing page excerpt state a price, a free tier, or a clear way to start?
yesA price, free plan, trial, or open-source license is statednoNo pricing or start path is visible - copy_qualityweight 16.7 . score
How professional and readable is the landing page excerpt as marketing copy?
- 0Broken, placeholder, or unreadable
- 1Readable but generic or sloppy
- 2Clear and competent
- 3Sharp, specific, and confident
- looks_realweight 11.1 . noul
Does the text describe a product that exists and can be used today, rather than an idea or waitlist?
yesDescribes working features, signup, or usagenoComing soon, waitlist, idea stage, or unclear - hype_onlypenalty, up to -17 . noul
Is the product text mostly buzzwords and claims with no concrete features or details?
yesMostly adjectives, superlatives, or AI buzzwords without specificsnoConcrete features or workflows are described
Turning answers into a score
- -A yes/no question contributes its probability times its weight. 0.8 on a weight of 15 is 12 points.
- -A score question contributes its position on the ladder times its weight. Level 2 of 4 is two thirds, so two thirds of the weight.
- -hype_only is subtracted, up to 17 points.
- -The weights in the table were specified against a 90 point total. They are normalised onto 100 so a perfect product actually reaches 100, with their ratios unchanged. The weight shown beside each question is the one applied.
- -All of this arithmetic happens in our code. Jev is never asked to add, count or compare numbers, because it is not reliable at that and the docs say so.
If the landing page cannot be read, pricing_clear and copy_quality are dropped and the remaining weights are scaled back up to 100, so a product without a readable page is judged on the same scale rather than being capped. The scorecard says so.
2. Extra questions per category
Some categories ask up to two more questions, worth 10 each. That weight comes out of pricing_clear and copy_quality in proportion, so the total stays 100 and both of those still count.
Dev Tools
has_install_pathDoes the text show how a developer would install or integrate it (CLI command, SDK, API, or code snippet)?
AI
ai_is_coreIs the AI functionality the main thing the product does, rather than a feature added to something else?
Finance
states_compliance_or_trustDoes the text mention security, compliance, or how money is protected?
Every other category uses the base rubric only.
3. The rank
The score is a first pass. The rank combines three things:
+ 0.4 x eloNormalised
+ 0.1 x shippingRecent x 100
- -Score is the 100 point total above.
- -Elo is a head to head rating inside the category. Everyone starts at 1000. Each day Jev compares up to 20 pairs per category: new entrants against their nearest neighbours, plus a few random pairs among the top ten to keep the head of the table honest. Each pair is one call with three questions, and the majority wins. K is 32, scaled by how confident Jev was, so an unsure verdict barely moves the rating.
- -Elo is normalised inside each category, over a span of at least 200 points. Without that floor, a two point rating gap in a young category would become a hundred point swing.
- -Freshness comes from shipping_recent, asked at the weekly re-judge.
Ranks are recomputed at 09:00 UTC every day and one row per product per day is kept, which is where the movement arrows and the 30 day sparkline come from.
4. The weekly re-judge
Every product is re-scraped and re-judged once a week, spread across the week so roughly a seventh of the catalogue goes each day. Two extra questions are added to that call:
- shipping_recent
Has the landing page changed in a way that shows new features, pricing, or content since the previous version?
- still_live
Does the current excerpt describe a working product that can be used today?
If still_live comes back below 0.3 twice in a row, or the site is genuinely gone twice in a row, the product is marked dormant and drops off the leaderboard. Its scorecard stays up. A page that merely blocks our scraper does not count: that says something about us, not about the product.
5. Product of the Day
- -Candidates are products submitted in the last 7 days that have not won before.
- -If the top two are 5 points or more apart, the composite score decides it and Jev is not asked.
- -Otherwise the top three go into one Jev choice question. This is the only place where more than two products go into a single call, because a bigger state costs accuracy.
- -Every pick says which of the two decided it.
6. The plain English finder
Your words go into the state, never into the instruction. We take up to 40 ranked products and ask Jev one choice question over them, plus one yes/no question: does any of these actually solve this. A choice always picks something, even from a bad field, so if that second answer is under 0.4 we say nothing fits instead of showing you the least bad option.
7. What this does not tell you
- -Jev reads what a founder wrote and what their landing page says. It does not use the product, so this measures how clearly a product explains itself, not whether it is any good.
- -Score levels are good for ordering and thresholds, not for reading an exact magnitude. A 62 and a 65 are not meaningfully different.
- -Yes/no answers do not carry a confidence figure from the API, so the unsure tag can only appear on the laddered questions. We say which is which on every scorecard rather than inventing a number.
- -Founder text is untrusted. Every question carries a note saying marketing claims are not evidence, but a determined copywriter can still move an answer.
Jev is made by TypeSafe AI. Back to the leaderboard.