Measures of Dispersion for Grouped Data · Form 5
Measures of Dispersion for Grouped Data: Worked Examples (KBAT)
Real situations where the method is not named: judging a service claim with the median and interquartile range, finding a cut-off mark from a percentile, and comparing the mean with the median to describe the shape of a distribution. For students ready to decide which measure a problem actually needs.
Worked example 1
A clinic recorded the waiting times (minutes) of 60 patients: 0–9 (6), 10–19 (14), 20–29 (20), 30–39 (10), 40–49 (6), 50–59 (4). The clinic advertises that 'half of our patients are seen within 25 minutes.'
Estimate the median waiting time to test this claim, and find the interquartile range as a measure of consistency.
- 'Half of our patients' points to the median; consistency points to the interquartile range.
- Cumulative frequency: 6, 20, 40, 50, 56, 60; n = 60.
- n ÷ 2 = 30 → median class 20–29 (L = 19.5, F = 20, f = 20, c = 10).
- Median = 19.5 + ((30 − 20) ÷ 20) × 10 = 19.5 + 5 = 24.5 minutes.
- 24.5 min < 25 min, so the claim is supported by the data.
- Q1 at n ÷ 4 = 15 → class 10–19: Q1 = 9.5 + ((15 − 6) ÷ 14) × 10 = 15.93 min.
- Q3 at 3n ÷ 4 = 45 → class 30–39: Q3 = 29.5 + ((45 − 40) ÷ 10) × 10 = 34.5 min.
- Interquartile range = 34.5 − 15.93 = 18.57 minutes.
Worked example 2
A school scored 50 projects: 40–49 (4), 50–59 (10), 60–69 (16), 70–79 (12), 80–89 (6), 90–99 (2). A 'distinction' is awarded to the top 20% of projects.
Using interpolation, estimate the minimum mark needed for a distinction.
- The top 20% lie above the 80th percentile, so find P80 (the mark with 80% of projects below it).
- Position = 80% × 50 = 0.8 × 50 = 40th value.
- Cumulative frequency: 4, 14, 30, 42, 48, 50; the 40th value is in class 70–79.
- L = 69.5, F = 30, f = 12, c = 10.
- P80 = 69.5 + ((40 − 30) ÷ 12) × 10 = 69.5 + 8.33 = 77.83.
- So a project needs about 78 marks (77.83 rounded up) to fall in the top 20%.
Worked example 3
A ranger recorded the masses (kg) of 32 monitor lizards: 1.0–1.9 (8), 2.0–2.9 (12), 3.0–3.9 (7), 4.0–4.9 (3), 5.0–5.9 (2). Find the estimated mean and the median, then use them to describe the shape of the distribution.
- Midpoints x: 1.45, 2.45, 3.45, 4.45, 5.45.
- Σf = 32; Σfx = 11.6 + 29.4 + 24.15 + 13.35 + 10.9 = 89.4.
- Estimated mean = 89.4 ÷ 32 = 2.79 kg.
- Cumulative frequency: 8, 20, 27, 30, 32; n ÷ 2 = 16 → median class 2.0–2.9.
- Median = 1.95 + ((16 − 8) ÷ 12) × 1.0 = 1.95 + 0.67 = 2.62 kg.
- Mean 2.79 > median 2.62, so the distribution is skewed to the right (a few heavier lizards pull the mean up).
Worked example 4
The table shows the lifespans, in hours, of 40 light bulbs from Batch A: 500–599 (4), 600–699 (10), 700–799 (14), 800–899 (8), 900–999 (4). (a) Estimate the mean and standard deviation of Batch A.
(b) Batch B has the same mean but a standard deviation of 95 hours. State, with a reason, which batch is more consistent.
- Midpoints x: 549.5, 649.5, 749.5, 849.5, 949.5; Σf = 40.
- Σfx = 4(549.5) + 10(649.5) + 14(749.5) + 8(849.5) + 4(949.5) = 29 780, so mean = 29 780/40 = 744.5 hours.
- Σfx² = 22 670 210; variance = 22 670 210/40 − 744.5² = 566 755.25 − 554 280.25 = 12 475.
- Standard deviation of A = √12 475 = 111.69 ≈ 111.7 hours.
- (b) Batch B's standard deviation (95 h) is smaller than Batch A's (111.7 h), so Batch B is more consistent because a smaller standard deviation means the lifespans are less spread out.
Worked example 5
In an examination taken by 80 candidates the marks were grouped: 30–39 (6), 40–49 (14), 50–59 (22), 60–69 (20), 70–79 (12), 80–89 (6). Only the top 15% of candidates receive a merit award.
Using interpolation, find the minimum mark needed for a merit award, then round up to the nearest whole mark.
- The top 15% start at the 85th percentile, at position 0.85 × 80 = the 68th value from the bottom.
- Cumulative frequencies: 6, 20, 42, 62, 74, 80; the 68th value lies in the class 70–79 (cf just below = 62).
- For this class: L = 69.5, cf before F = 62, frequency f = 12, class width c = 10.
- P₈₅ = 69.5 + ((68 − 62)/12) × 10 = 69.5 + 5 = 74.5, so the minimum whole mark is 75.
Worked example 6
A survey of 60 households recorded the number of electronic devices owned, grouped as: 0–4 (8), 5–9 (p), 10–14 (18), 15–19 (q), 20–24 (6). The estimated mean number of devices is 11.
Find the values of p and q.
- Total frequency: 8 + p + 18 + q + 6 = 60, so p + q = 28. ... (1)
- Midpoints: 2, 7, 12, 17, 22; Σfx = 8(2) + 7p + 18(12) + 17q + 6(22) = 364 + 7p + 17q.
- Mean = 11 means Σfx = 11 × 60 = 660, so 364 + 7p + 17q = 660, giving 7p + 17q = 296. ... (2)
- From (1), p = 28 − q; substitute into (2): 7(28 − q) + 17q = 296 → 196 + 10q = 296 → q = 10, then p = 18.
Worked example 7
A factory recorded the number of defective items found per batch over 100 batches: 0–4 (10), 5–9 (p), 10–14 (34), 15–19 (q), 20–24 (8). The median number of defects per batch is 12.
Find the values of p and q.
- Total frequency: 10+p+34+q+8 = 100 → p+q = 48.
- The median (50th value) lies in the class 10–14 (L = 9.5, F = 10+p, f = 34, c = 5).
- Apply the median formula: 12 = 9.5 + [(50 − (10+p))/34] × 5 → 2.5 = [(40−p)/34] × 5.
- (40−p)/34 = 0.5 → 40−p = 17 → p = 23.
- From p+q = 48: q = 48 − 23 = 25.
Worked example 8
A driving school recorded the number of hours 60 learners needed before passing their test: 5–9 (6), 10–14 (14), 15–19 (22), 20–24 (12), 25–29 (6). Estimate the number of learners who needed more than 18 hours, and express this as a percentage of all learners.
- Cumulative frequency up to 14.5 (end of 10–14 class) = 6+14 = 20.
- Within the 15–19 class (boundaries 14.5–19.5, frequency 22), the fraction below 18 is (18−14.5)/(19.5−14.5) = 3.5/5 = 0.7, giving about 0.7 × 22 = 15.4 learners below 18 within that class.
- So the number below 18 hours ≈ 20 + 15.4 = 35.4, meaning the number above 18 hours ≈ 60 − 35.4 = 24.6 ≈ 25 learners.
- Percentage = (24.6/60) × 100% ≈ 41%.
Source:DSKP KSSM Mathematics Form 4 and 5 (Versi English)
Frequently asked questions
What extra demand do KBAT questions add to grouped dispersion?
KBAT items usually give two grouped datasets, say, two classes' exam scores or two factories' output, and ask you to compare their spread, not just calculate it. You need to compute a measure like standard deviation or interquartile range for each, then write a conclusion linking the smaller value to greater consistency or less variation.
What does the examiner reward in a KBAT comparison answer?
Full, correctly labelled calculations for both datasets, followed by an explicit comparative statement using the actual figures, for example, naming which group has the smaller standard deviation and stating what that means in context. A calculation with no concluding sentence, or a conclusion without the numbers to back it, loses marks.
What trips students up in these KBAT questions?
Calculating both datasets' measures correctly but forgetting to compare them in words, mixing up which measure indicates more consistency (a smaller spread means more consistent, not the larger one), and rounding intermediate values too early, which throws off the comparison. Always link the final numbers back to the real-world context asked about.