Measures of Dispersion for Ungrouped Data · 8.2.2
Pros and cons of dispersion measures
Students discuss when each measure of dispersion, range, interquartile range, variance, standard deviation, is useful and where it falls short. For example, range is quick to find but easily distorted by extreme values, interquartile range resists outliers but ignores much of the data, and standard deviation uses every value yet is still affected by extreme data.
The official learning standard (8.2.2)
“Explain the advantages and disadvantages of various measures of dispersion in describing ungrouped data.”
What it means
Students discuss when each measure of dispersion, range, interquartile range, variance, standard deviation, is useful and where it falls short. For example, range is quick to find but easily distorted by extreme values, interquartile range resists outliers but ignores much of the data, and standard deviation uses every value yet is still affected by extreme data.
How it is examined
This is examined in SPM Paper 2, usually as a short reasoning part of a statistics question asking students to justify which measure of dispersion is more suitable for a given data set, especially one containing an outlier or extreme value, rather than to perform a fresh calculation.
Worked example
A set of 20 data values contains one extremely large outlier. Explain why the interquartile range is more suitable than the range to describe the dispersion of this data.
- The range only uses the maximum and minimum values, so a single extremely large outlier directly increases the range and gives a misleadingly large value.
- The interquartile range uses only the middle 50% of the ordered data (between Q1 and Q3), so it is not affected by extreme values at either end.
- Since the data set contains an outlier, IQR better represents the spread of the majority of the data.
Source:DSKP KSSM Mathematics Form 4 and 5 (Versi English)
Frequently asked questions
Why is standard deviation usually preferred over range?
Standard deviation uses every value in the data set, so it reflects the overall spread more accurately than range, which only depends on the two extreme values. This makes standard deviation a more reliable measure for most SPM comparison questions, unless the data has a serious outlier.
Is IQR always better than standard deviation?
Not always, IQR is better when a data set has outliers, since it ignores the extreme 25% at each end. But standard deviation uses all the data, so it is often preferred for data without outliers, as it captures the spread more completely.
Why is range considered the weakest measure of dispersion?
Range only depends on the two extreme values (the maximum and minimum) and ignores every other value in between, so a single unusually large or small value can make the range misleading. Despite this weakness, it remains useful because it is quick and simple to calculate.