We searched twenty common foods across four apps, compared each returned entry against reference composition data or the manufacturer’s own published figures, and noted where the correct one ranked.
On crowd-sourced databases, ranking told us nothing
The correct entry appeared first about as often as it appeared fourth. Position appeared to reflect popularity — how many people had previously selected an entry — which is a measure of what other users guessed, not of what is true.
That is a self-reinforcing loop. An entry selected often ranks higher, ranks higher so is selected more often, and its correctness never enters the calculation at any point.
On one chain sandwich we found eleven entries spanning about two hundred calories. Nothing in the interface indicated which came from the restaurant’s own nutrition data and which came from a stranger in 2016.
On curated databases the question does not arise
Search a chicken breast in a curated catalogue and you get one record. There is nothing to rank because there is nothing to choose between.
This is the practical argument for verified data that is easier to feel than the accuracy statistics are. It is not mainly that curated entries are individually more accurate — though they are. It is that you are not being asked to adjudicate, and you were never equipped to adjudicate.
The habit to break
Tapping the first result.
Everyone does it, including us, including while running this test. It is the whole design of a search interface and resisting it costs attention you would rather spend on anything else.
Which is why, if you are on a crowd-sourced database, the honest advice is uncomfortable: you have to check, every time, against the packet in your hand or the restaurant’s published figures. If you are not going to do that, the database’s size is not an advantage to you — it is a wider distribution to draw randomly from.
The three-second version
Before selecting an entry, look at whether the calorie figure is plausible for the portion in front of you. Not precise — plausible. A 200-calorie difference on a sandwich is visible if you glance at it, and glancing catches most of the bad picks.
That is a lower standard than verification and it catches the errors that matter.
Inés Okonkwo
Editor
Founded this because product writing had stopped distinguishing between a measurement and a press release. Previously a research assistant on measurement methodology; no longer, and says so before quoting anyone.
Elsewhere in the magazine
- We scanned sixty packets into five apps. The failure rate was not where we expected.
- Fibre is the test: how to find out in sixty seconds whether your food app has a real database
- The deep bowl problem: why every food camera fails the same way and none of them mention it
- We weighed everything for a week before logging it. The apps were not the biggest source of error.