AI Academy
Dionysus reclining with a wine cup raised, a leopard resting beside him.
Expert

Why language models fumble numbers

A language model produces text that looks like math, which is different from doing math. Common calculations come out right because they appear in training data a million times. Multi-step ones, unit conversions, anything unusual: those can silently go wrong.

Hi-fi runs on exactly the math models fumble. Decibels are logarithmic. A model may tell you a 100 W amp plays twice as loud as a 50 W amp. It plays about 3 dB louder, a just-noticeable step. Doubling perceived loudness takes roughly ten times the power.

The habit to build: when an answer hinges on a number, check the number. Recalculate it, or ask the model to show its steps, which exposes errors fast. We anchor the Assistant to real specs for the same reason. Numbers should come from data, never from memory.

Go deeper

Text that looks like arithmetic

A model generates numbers the way it generates words: by predicting what usually comes next. "2 plus 2 equals" is followed by "4" in the training data every time, so it gets that right. "What is 20 times log base 10 of 3" appears rarely, so the model produces something that looks like a result. Sometimes it is 9.54. Sometimes it is a plausible neighbor. The sentence around it reads identically in both cases.

Multi-step problems compound the risk. Each step is a fresh prediction, and an error early in the chain travels quietly to the end. Unit conversions are a classic failure, because the model has seen the same quantity expressed in different units and blends them.

Practical takeaway: a model producing a number is writing. Treat the output as a draft of the math, to be checked.

The hi-fi math models get wrong

Decibels are the first trap. Power and loudness are logarithmic, so doubling amplifier power adds 3 dB, and doubling perceived loudness needs about ten times the power. A model that says a 100 W amp plays twice as loud as a 50 W amp has confused the two scales, which happens often because the sentence "twice the power, twice as loud" appears everywhere online.

Sensitivity ratings are the second. A speaker rated 88 dB at 2.83 V and one rated 88 dB at 1 W are the same speaker only if the impedance is 8 ohms. At 4 ohms, 2.83 V is 2 W, so the 2.83 V rating flatters the speaker by 3 dB. Models mix the two conventions constantly.

Distance is the third. Sound level drops roughly 6 dB each time you double the distance from the speaker in free space, so a listening seat at 3 m sits about 9.5 dB below the 1 m rating. Models often quote the 1 m figure as if that is what you hear.

Practical takeaway: for any answer involving dB, watts, or sensitivity, check which convention the model used. The wrong one is off by 3 dB or more, and 3 dB is a real difference.

The pattern that catches errors

Ask the model to show its working. "Calculate the peak level at 3 m for an 86 dB at 2.83 V speaker with a 4 ohm load on a 60 W amp, step by step." Now each step is visible: 60 W into 4 ohms, the 2.83 V convention correction, the distance loss. When one step is wrong, you can see it and correct it. When the steps are right, you have a checkable calculation instead of a bare number.

Reasoning models do this more reliably than fast ones, because writing out the steps is what they were trained to do. It is one of the clearest cases where the expensive model earns its price. A calculation with a real consequence, like whether an amplifier will clip at your listening level, is worth the extra tokens.

Practical takeaway: for a number that decides a purchase, say "step by step" and read every step. If the model skips one, ask for it.

Why the Assistant reads specs instead of remembering them

The Assistant reads specs from the database record rather than from the model’s memory. The speaker’s sensitivity, its nominal impedance, the amplifier’s rated power into 4 and 8 ohms, all read from recorded data at the moment of the question. The model still does the reasoning, and the checks above still apply to its arithmetic. What it will not do is invent the inputs. A calculation can only be as good as the numbers going in, and those, at least, are not predictions.

Practical takeaway: get the inputs from data and the working shown in steps. Between the two, most numerical errors have nowhere to hide.