Form 5 · Statistics and Probability

Measures of Dispersion for Grouped Data

This chapter measures spread for grouped data, using class intervals, ogives and the graphs that come with them.

What is Measures of Dispersion for Grouped Data?

When data is grouped into class intervals, you cannot use each value directly. This chapter shows how to find the mean, variance and standard deviation from a frequency table, and how to draw and read a histogram, frequency polygon and ogive (cumulative frequency curve).

Content standards (DSKP)

The DSKP KSSM sets these content standards for this chapter:

The key ideas

Midpoints stand in for values

For grouped data you use the class midpoint as the value for the whole interval in every calculation.

The ogive

A cumulative-frequency curve lets you read off the median and quartiles, and therefore the interquartile range.

Standard deviation from a table

The given grouped formula uses Σfx and Σfx². A clear frequency table with these columns makes it almost mechanical.

How this chapter is examined

Paper 2 gives a frequency table and asks for the mean and standard deviation, or provides data to plot an ogive and then read the median and interquartile range from it. A carefully built table and a smooth, accurate curve are where the marks are.

Formulae given in the exam for this chapter

Mean (grouped)
x̄ = Σfx / Σf
Given in the exam
Variance (grouped)
σ2 = Σf(x−x̄)2 / Σf = Σfx2/Σf − x̄2
Given in the exam
Standard deviation (grouped)
σ = √Σf(x−x̄)2 / Σf = √Σfx2/Σf − x̄2
Given in the exam

Common mistakes to avoid

  • Using class boundaries when the midpoint is needed (or vice versa)
  • Plotting an ogive against the wrong x-value
  • Reading the median off a rushed, uneven curve

Three x-values, three jobs: midpoint, class boundary, upper boundary

Grouped-data questions punish students who mix up the three x-values a class can supply. The midpoint, the average of a class's lower and upper limits, stands in for every value in the class and is the x you put into the mean and standard deviation calculations.

The class boundaries, found by closing the half-unit gaps between classes, set the width of each bar in a histogram and are what you use to identify and calculate the modal class. The upper boundary of each class is the x-value you plot the cumulative frequency against on an ogive, because everything up to that boundary has already been counted.

Match the right x to the right graph and the marks follow.

Reading the median and quartiles off an ogive

An ogive turns a frequency table into a smooth cumulative curve, and its main job in Paper 2 is to hand you the median and the quartiles by reading, not by formula. With N values in total, the median sits at the N/2 mark: find N/2 on the vertical cumulative-frequency axis, run a horizontal line across to the curve, then drop straight down to read the value on the horizontal axis.

The first quartile Q1 uses N/4 and the third quartile Q3 uses 3N/4, read the same way. Note that for grouped data from an ogive you use N/4, N/2 and 3N/4, not the (N + 1) positions used for a short ungrouped list.

The interquartile range is then simply Q3 − Q1.

Comparing two groups with the mean and standard deviation together

A common final part asks you to compare two sets, two classes, two machines, two months, using the figures you have just calculated, and full marks need both numbers read together. The mean tells you the centre: which group is higher on average.

The standard deviation tells you the consistency: a smaller standard deviation means the values cluster tightly around the mean, so the group is more uniform and predictable, while a larger one means the values are more spread out. So a class with the same mean but a smaller standard deviation is the more consistent performer, and a machine with a larger standard deviation is the less reliable one.

Always state what each number says about the situation, not just which is bigger.

A worked exam-style example

A frequency table with clean class midpoints, taken through to the mean and standard deviation.

  1. Write down the midpoint x of each class (the average of its limits): 35, 45, 55, 65, 75.
  2. Form the fx column: 4×35 = 140, 10×45 = 450, 14×55 = 770, 8×65 = 520, 4×75 = 300. Totals: Σf = 40, Σfx = 140 + 450 + 770 + 520 + 300 = 2180.
  3. Mean x̄ = Σfx ÷ Σf = 2180 ÷ 40 = 54.5 kg.
  4. Form the fx² column (f × x²): 4×35² = 4900, 10×45² = 20250, 14×55² = 42350, 8×65² = 33800, 4×75² = 22500. Sum: Σfx² = 123800.
  5. Variance σ² = Σfx²/Σf − x̄² = 123800 ÷ 40 − 54.5² = 3095 − 2970.25 = 124.75.
  6. Standard deviation σ = √124.75 = 11.2 kg (3 s.f.).

How to study this chapter

Frequently asked questions

How this chapter is examined

SPM Mathematics assesses this chapter across Mathematics Paper 1 (Objective) and Mathematics Paper 2 (Subjective), drawing on the DSKP content standards above. Paper 2 gives marks for working, so showing every step matters.

Common mistakes to avoid

Using class boundaries when the midpoint is needed (or vice versa); Plotting an ogive against the wrong x-value; Reading the median off a rushed, uneven curve.

Formulae given in the exam for this chapter

Yes, Mean (grouped), Variance (grouped), Standard deviation (grouped) appear on the formula sheet the exam provides. Anything else in this chapter you are expected to know.

Do I use the midpoint or the class boundary in the mean formula?

Always the midpoint, the average of the class's lower and upper limits. The midpoint represents every value in that class, so it is the x in both the mean and the standard deviation.

Class boundaries are only for drawing the histogram and the ogive and for the modal class; they never go into the mean.

For an ogive, what do I plot the cumulative frequency against?

Against the upper boundary of each class, not the midpoint. The cumulative frequency counts everything up to the top of that class, so it must sit above the upper boundary.

Also add a starting point with zero cumulative frequency at the lower boundary of the first class, so the curve begins on the axis.

How do I read the median and interquartile range from the ogive?

With N values, go to N/2 on the cumulative-frequency axis, read across to the curve, then down to the horizontal axis for the median. Read Q1 at N/4 and Q3 at 3N/4 the same way.

The interquartile range is Q3 − Q1. Use a sharp, smoothly drawn curve so the readings are accurate.

Source:DSKP KSSM Mathematics Form 4 and 5 (Versi English)· SPM: Format Pentaksiran mulai 2021, Matematik (1449)

Get help with Measures of Dispersion for Grouped Data

One-to-one, in English, with your working checked line by line.

Get help with Measures of Dispersion for Grouped Data
One-hour paid trial · Same-day replyfrom RM50/hr
Book a Trial Class