The short answer
Every nutrition score does the same operation: it takes a dozen label values, multiplies each by a weight somebody chose, adds the result, and prints one symbol. Compression is the point—and compression destroys information. Two products with opposite strengths can land on the identical letter, star count, or number, because the trade-off between them was averaged away before you ever saw it. The score is not wrong arithmetic; it is someone else's priorities presented as your answer.
How the major scoring systems actually work
These are not hypothetical. Nutri-Score, used on packaging across much of Europe, assigns a grade from A (green) to E (red). Its algorithm subtracts points for energy, sugars, sodium, and saturated fat per 100g, then adds points back for fiber, protein, and fruit, vegetable, and legume content; the net point total maps to the letter. Australia's and New Zealand's Health Star Rating runs a similar points model and displays 0.5 to 5 stars. In the United States, grocery chains and apps show proprietary 1–100 scores on shelf tags and scan screens. Each system bakes in a fixed weighting—and each weighting is a policy decision about what “healthy” means for everyone at once.
Same score, opposite products
Consider two yogurts. Yogurt A: 170 kcal, 15g protein, 9g added sugar per serving. Yogurt B: 100 kcal, 6g protein, 2g added sugar. Under a points model, A's protein earns points back while its sugar and energy cost points; B's low sugar earns a favorable base while its modest protein adds little. Both can plausibly net out to the same band—the same letter, the same 3.5 stars, the same low-70s score. Yet these products answer different questions. A trades sugar for protein; B trades protein for sugar. A shopper rebuilding after training and a shopper managing blood glucose should pick opposite products off the same row of identical grades.
Fixed weights cannot know your goal
A weight is an opinion frozen into arithmetic. When Nutri-Score's designers decided how many points sodium costs relative to protein, they made a reasonable population-level judgment. But a shopper whose clinician told them to watch sodium does not experience sodium as one weighted input among many—it is the headline number. Averaging cannot represent a priority; it can only dilute it. This is why the same score that usefully sorts a shelf for a general audience can actively mislead a person with one binding constraint.
What scores are genuinely good at
None of this makes scores useless. A letter or star count is a fast triage tool: it can shrink a wall of forty cereals to a shortlist of six in seconds, which is better than giving up and grabbing the familiar box. The mistake is treating triage as the decision. The defensible workflow is: let the score narrow the field, then open the two or three surviving labels and compare the actual numbers on a common basis—per serving or per 100g—with your own goal doing the weighting. That is the step a score cannot perform for you, and it is the step AptBite is built around: it shows both products' numbers and the trade-off between them, and never collapses that into a verdict.
“Healthiest” changes when the question changes
Take the same pair of products and ask three questions: which provides more protein per 100 calories, which has less added sugar in the portion you eat, and which costs less per 25g of protein? Each question can produce a different leader without any contradiction. The labels did not change; the goal and denominator did. A single score cannot preserve all three answers because its purpose is to collapse them. Before accepting a winner, translate the claim into a unit the label can answer—grams, calories, milligrams, or price—and state the basis.
Transparent comparison is more than showing raw numbers
A useful comparison needs four visible parts: the original label inputs, the common basis used for normalization, the size of the difference, and the trade-off that could reverse the decision. Raw numbers without a shared basis can still mislead, and a percentage difference without the original values can exaggerate a tiny gap. Keeping all four together makes the result auditable: another shopper can change the goal, update the price, or correct a serving size without rebuilding an opaque score.
A practical final check
- Name the score system and find its weights—if you cannot, you are trusting an unknown opinion.
- Ask whether the score's priority matches your constraint for this purchase.
- Pull the two labels and compare the one or two nutrients that actually matter to you.
- Check the basis: a score computed per 100g can rank differently than your real portion would.
For what the underlying label lines mean, see the FDA Nutrition Facts label guidance. Scores are a sorting aid, not a diagnosis; for medically required diets, pregnancy, or medication concerns, work from the label with a qualified professional.