Skip to content
UncommonBits
Technology, tested differently

How we test

Every score on UncommonBits comes from hands-on use. This page explains exactly how we get to a number, so you can decide how much weight to give it.

The short version

We buy or sign up for the product ourselves. We use it for real work over a defined period, run a fixed set of tasks against it, and record what happened. We publish the number of days, the number of tasks, and the date we last tested. If we have not used a product, we do not score it.

What every review must document

No review is published without all of the following recorded, and most of it is printed on the page:

  • The product and the exact version or model tested
  • The date testing began and the date it ended
  • How many days the product was in real use
  • How many discrete tasks were run against it
  • The test environment, including operating system, browser and plan tier
  • What it did well, with examples
  • What failed, with examples
  • Screenshots or transcripts of the runs that informed the score
  • The reasoning behind the score, including why it differs from the average of the sub-scores

The five things we score

Every review is scored across the same five categories, on a scale of 0 to 10, in whole and half points only. We use half points because our testing cannot honestly distinguish 8.3 from 8.4, and pretending otherwise would be theatre.

Quality

Is the output correct, and is it good? For an AI product this means factual accuracy, instruction following, and whether the result needs so much editing that starting from scratch would have been faster.

Reliability

Does it behave the same way twice? We run a subset of tasks repeatedly and record variance, refusals, timeouts and outright failures. A product that is brilliant four times in five scores badly here.

Ease of use

How long from signing up to a genuinely useful result, without reading documentation? We time this.

Features

Depth measured against what the category actually offers today, not against a wishlist. A missing feature counts against a product only if a direct competitor ships it.

Value

What you get for the money at a realistic level of use, not on the free tier and not on the enterprise plan. Where pricing is usage-based, we publish what our own testing cost.

What the overall score is

The overall UB Score is a judgement, not an average. We weight the categories according to what the product is for: reliability matters more in something people run unattended, ease of use matters more in something aimed at non-specialists. Where the overall score differs noticeably from the mean of the sub-scores, the review says why.

What the bands mean

  • 9.0 and above, Exceptional. Best in its category. We would use it ourselves and recommend it without hedging.
  • 8.0 to 8.9, Excellent. A strong choice with known, stated limitations.
  • 7.0 to 7.9, Good. Does the job. Something else probably does it better for your specific case.
  • 6.0 to 6.9, Mixed. Real strengths undercut by real problems. Read the detail before buying.
  • 4.0 to 5.9, Weak. We would not spend our own money on it in its current state.
  • Below 4.0, Avoid. Broken, misleading, or dishonest about what it does.

How we buy what we test

We pay for products with our own money at standard retail prices. We do not accept free permanent licences, paid placements, or review units offered on condition of coverage. Where a vendor gives us temporary access we could not otherwise obtain, we say so in the review itself, at the top, not in a footnote.

Retesting

Software changes. Every review carries a last-tested date, and we retest anything in a fast-moving category at least every six months. When a retest changes a score, we update the review, change the date, and add a note explaining what moved and why. We do not quietly edit scores.

What we will not do

  • Score a product we have not used.
  • Describe a feature we have not verified ourselves.
  • Publish a review written or substantially drafted by an AI system.
  • Let a commercial relationship influence a score, a ranking, or whether we publish at all.
  • Present sponsored content as editorial.

Tell us we got it wrong

If you think a score is unfair or a fact is wrong, we want to hear it. Write to contact@uncommonbits.com and we will look again. Anything we get wrong is logged publicly on our corrections page.

Last updated 6 September 2026.