Model G20 2027 at FLAME University, registrations now open

Statistics | CBSE Class 10 Maths Notes

25 min read

On this page

Watch & explore

Start with a few high-quality watches, then dive into the notes below.

What Is Statistics: Crash Course Statistics #1 · CrashCourse

These Mathematics notes cover measures of central tendency, class marks, the direct, assumed mean and step-deviation methods, mode, cumulative frequency, median, missing frequencies, continuous class intervals, and the choice of a suitable average.

What do grouped data and measures of central tendency represent?

Mean, median and mode are numerical representatives of a collection of observations. They describe central tendency in different ways. The mean uses the values of all observations, the median identifies the middle of an ordered distribution, and the mode concerns the most frequent value.

Data in many real situations are large enough to need condensation. A grouped frequency distribution organises observations into class intervals and records how many fall in each interval. It is easier to read, but it no longer displays every individual value.

How are class intervals read?

The lower limit and upper limit identify the ends of an interval. With the grouping convention used here, an observation at an upper limit belongs to the next class. For example, a score of 40 belongs to 40 to 55, rather than 25 to 40.

A frequency counts observations; it is not the value of an observation. A frequency of seven in the class 40 to 55 means seven students have scores in that interval. It does not say that each of them scored seven marks.

Result: A class mark represents a class

Let ii identify a class, and let xix_i denote its class mark, or midpoint. The class mark is found by averaging the upper and lower limits. For calculating a grouped mean, the frequency of the class is assumed to be centred at this midpoint.

  1. Write the midpoint rule: xi=upper class limit+lower class limit2.x_i=\frac{\text{upper class limit}+\text{lower class limit}}{2}.
  2. For the class 10 to 25, substitute the limits: xi=25+102.x_i=\frac{25+10}{2}.
  3. Add and divide: xi=352=17.5.x_i=\frac{35}{2}=17.5.

Note: Grouping can change the calculated mean because class marks replace individual observations. An approximate grouped mean and an exact mean from the original observations need not agree.

This distinction concerns the information used, rather than the calculation method. Once the same grouped table and class marks are fixed, the three methods of calculating its mean give the same result. Their arithmetic is arranged differently to make the work more convenient.

How is the mean calculated by the direct method?

Let fif_i be the frequency corresponding to xix_i, and let xˉ\bar{x} denote the mean. The symbol ∑\sum, called sigma, means summation over all the observations or classes listed. The total frequency is the number of observations, not the number of rows.

Result: The direct mean is a weighted average

Each value contributes once for every observation having that value. Therefore, multiply before adding: the product of a value and its frequency gives its contribution to the total. For grouped data, use the class mark as the value in this calculation.

  1. Find the total frequency: number of observations=∑fi.\text{number of observations}=\sum f_i.
  2. Find the total of frequency-weighted values: weighted total=∑fixi.\text{weighted total}=\sum f_i x_i.
  3. Divide the weighted total by the total frequency: xˉ=∑fixi∑fi.\bar{x}=\frac{\sum f_i x_i}{\sum f_i}.

The following distribution gives the individual marks obtained by 30 students in a paper out of 100. Here the listed values are actual marks, so the midpoint assumption is unnecessary. The product column preserves the contribution of every student to the overall total.

Marks xix_iStudents fif_iProduct fixif_i x_i
1011×10=101\times 10=10
2011×20=201\times 20=20
3633×36=1083\times 36=108
4044×40=1604\times 40=160
5033×50=1503\times 50=150
5622×56=1122\times 56=112
6044×60=2404\times 60=240
7044×70=2804\times 70=280
7211×72=721\times 72=72
8011×80=801\times 80=80
8822×88=1762\times 88=176
9233×92=2763\times 92=276
9511×95=951\times 95=95

Worked example 1. Find the exact mean of the individual marks in the table.

  1. Calculate each product as shown in the final column, then add frequencies: ∑fi=1+1+3+4+3+2+4+4+1+1+2+3+1=30.\sum f_i=1+1+3+4+3+2+4+4+1+1+2+3+1=30.
  2. Add the weighted marks: ∑fixi=10+20+108+160+150+112+240+280+72+80+176+276+95=1779.\sum f_i x_i=10+20+108+160+150+112+240+280+72+80+176+276+95=1779.
  3. Apply the direct formula: xˉ=177930=59.3.\bar{x}=\frac{1779}{30}=59.3.

Answer: The exact mean is 59.359.3 marks.

What changes when these marks are grouped?

The same scores can be placed in six intervals of width 15. The class marks now stand in for the scores in each interval. Calculate each midpoint and multiply it by the corresponding frequency, as shown below.

MarksFrequency fif_iClass mark xix_iProduct fixif_i x_i
10 to 25217.517.52×17.5=352\times 17.5=35
25 to 40332.532.53×32.5=97.53\times 32.5=97.5
40 to 55747.547.57×47.5=332.57\times 47.5=332.5
55 to 70662.562.56×62.5=3756\times 62.5=375
70 to 85677.577.56×77.5=4656\times 77.5=465
85 to 100692.592.56×92.5=5556\times 92.5=555

Worked example 2. Find the grouped mean for the marks table.

  1. Find midpoints: (10+25)/2=17.5,(25+40)/2=32.5,(40+55)/2=47.5.(10+25)/2=17.5,\quad(25+40)/2=32.5,\quad(40+55)/2=47.5. (55+70)/2=62.5,(70+85)/2=77.5,(85+100)/2=92.5.(55+70)/2=62.5,\quad(70+85)/2=77.5,\quad(85+100)/2=92.5.
  2. Add frequencies: ∑fi=2+3+7+6+6+6=30.\sum f_i=2+3+7+6+6+6=30.
  3. Add the products in the table: ∑fixi=35+97.5+332.5+375+465+555=1860.\sum f_i x_i=35+97.5+332.5+375+465+555=1860.
  4. Divide by the total frequency: xˉ=186030=62.\bar{x}=\frac{1860}{30}=62.

Answer: The grouped mean is 6262 marks, an approximation to the exact mean of 59.359.3 marks.

Why does the assumed mean method work?

The assumed mean method reduces the size of the numbers used in multiplication. Choose a convenient value, usually a class mark near the centre, and subtract it from each class mark. This changes the arithmetic without changing the final grouped mean.

Let aa denote the assumed mean, did_i the deviation of a class mark from it, and dˉ\bar{d} the mean of these deviations. The deviation is signed: a class mark below the assumed mean gives a negative value.

Derivation: How is the assumed mean formula obtained?

  1. Define each deviation: di=xi−a.d_i=x_i-a.
  2. Take the frequency-weighted mean of the deviations: dˉ=∑fidi∑fi=∑fi(xi−a)∑fi.\bar{d}=\frac{\sum f_i d_i}{\sum f_i}=\frac{\sum f_i(x_i-a)}{\sum f_i}.
  3. Separate the terms and take the fixed assumed mean outside the sum: dˉ=∑fixi∑fi−a∑fi∑fi=xˉ−a.\bar{d}=\frac{\sum f_i x_i}{\sum f_i}-\frac{a\sum f_i}{\sum f_i}=\bar{x}-a.
  4. Restore the subtracted value: xˉ=a+dˉ=a+∑fidi∑fi.\bar{x}=a+\bar{d}=a+\frac{\sum f_i d_i}{\sum f_i}.

The correction is added to the assumed mean. It may itself be positive or negative. Choosing an assumed mean does not declare that it is the actual mean; the weighted deviations supply the required correction.

How are the deviations used in a table?

Use the grouped marks from the preceding section and choose a=47.5a=47.5. Retain the original frequencies. The new columns contain the deviations and their products with frequencies. The zero deviation contributes zero even though its class has seven students.

MarksFrequency fif_iDeviation did_iProduct fidif_i d_i
10 to 25217.5−47.5=−3017.5-47.5=-302×(−30)=−602\times(-30)=-60
25 to 40332.5−47.5=−1532.5-47.5=-153×(−15)=−453\times(-15)=-45
40 to 55747.5−47.5=047.5-47.5=07×(0)=07\times(0)=0
55 to 70662.5−47.5=1562.5-47.5=156×(15)=906\times(15)=90
70 to 85677.5−47.5=3077.5-47.5=306×(30)=1806\times(30)=180
85 to 100692.5−47.5=4592.5-47.5=456×(45)=2706\times(45)=270

Worked example 3. Calculate the mean of the grouped marks using the assumed mean 47.5.

  1. Compute the six deviations and products shown in the table, and total the frequencies: ∑fi=2+3+7+6+6+6=30.\sum f_i=2+3+7+6+6+6=30.
  2. Add signed products: ∑fidi=−60−45+0+90+180+270=435.\sum f_i d_i=-60-45+0+90+180+270=435.
  3. Calculate the mean deviation: dˉ=43530=14.5.\bar{d}=\frac{435}{30}=14.5.
  4. Restore the assumed mean: xˉ=47.5+14.5=62.\bar{x}=47.5+14.5=62.

Answer: The mean is 6262 marks, matching the direct calculation.

The choice of assumed mean does not affect the answer. A central class mark is convenient because it often keeps the deviations small. Changing that choice changes the correction, but adding the correction back restores the same mean.

How does step-deviation simplify the mean calculation?

The step-deviation method makes another simplification after subtracting the assumed mean. When deviations have a common factor, dividing by that factor produces smaller numbers to multiply. For equal class sizes, the common class width is often a convenient divisor.

Let hh denote this non-zero divisor, taken here as the class width. Let uiu_i denote the scaled deviation and uˉ\bar{u} its frequency-weighted mean. The final calculation must restore both changes: the division by the divisor and the subtraction of the assumed mean.

Derivation: How is the step-deviation formula obtained?

  1. Define the scaled deviation: ui=xi−ah.u_i=\frac{x_i-a}{h}.
  2. Calculate its weighted mean: uˉ=∑fiui∑fi=1h(∑fixi−a∑fi∑fi).\bar{u}=\frac{\sum f_i u_i}{\sum f_i}=\frac{1}{h}\left(\frac{\sum f_i x_i-a\sum f_i}{\sum f_i}\right).
  3. Simplify using the direct mean formula: uˉ=xˉ−ah.\bar{u}=\frac{\bar{x}-a}{h}.
  4. Multiply by the divisor and restore the assumed mean: huˉ=xˉ−a,xˉ=a+huˉ.h\bar{u}=\bar{x}-a,\qquad \bar{x}=a+h\bar{u}.
  5. Substitute the weighted mean of scaled deviations: xˉ=a+h∑fiui∑fi.\bar{x}=a+h\frac{\sum f_i u_i}{\sum f_i}.

The divisor must be restored. The mean of the scaled deviations is not the mean of the original observations. It describes how far the mean lies from the assumed mean, measured in the chosen steps.

How does this work for the marks distribution?

For the same six marks intervals, keep a=47.5a=47.5 and take h=15h=15. The ordinary deviations are all multiples of the class width. Their scaled values are small integers, making the products particularly easy to calculate.

MarksFrequency fif_iScaled deviation uiu_iProduct fiuif_i u_i
10 to 252(17.5−47.5)/15=−2(17.5-47.5)/15=-22×(−2)=−42\times(-2)=-4
25 to 403(32.5−47.5)/15=−1(32.5-47.5)/15=-13×(−1)=−33\times(-1)=-3
40 to 557(47.5−47.5)/15=0(47.5-47.5)/15=07×(0)=07\times(0)=0
55 to 706(62.5−47.5)/15=1(62.5-47.5)/15=16×(1)=66\times(1)=6
70 to 856(77.5−47.5)/15=2(77.5-47.5)/15=26×(2)=126\times(2)=12
85 to 1006(92.5−47.5)/15=3(92.5-47.5)/15=36×(3)=186\times(3)=18

Worked example 4. Find the mean marks using step-deviation.

  1. Calculate the scaled deviations and products in the table, then sum the frequencies: ∑fi=2+3+7+6+6+6=30.\sum f_i=2+3+7+6+6+6=30.
  2. Add the signed products: ∑fiui=−4−3+0+6+12+18=29.\sum f_i u_i=-4-3+0+6+12+18=29.
  3. Find and rescale the mean deviation: uˉ=2930,huˉ=15×2930=14.5.\bar{u}=\frac{29}{30},\qquad h\bar{u}=15\times\frac{29}{30}=14.5.
  4. Add the assumed mean: xˉ=47.5+14.5=62.\bar{x}=47.5+14.5=62.

Answer: The mean is 6262 marks by the third method as well.

Which mean method is suitable when class sizes differ?

The direct method is convenient when the class marks and frequencies make multiplication straightforward. The assumed mean and step-deviation methods simplify the same calculation when the numbers are less convenient. A method should be chosen for manageable arithmetic, not to obtain a different mean.

Result: All three mean methods agree

For a fixed grouped table, the methods use the same class marks and frequencies. The assumed mean method subtracts and restores a fixed value. Step-deviation also divides and multiplies by a fixed non-zero number. These transformations preserve the grouped mean.

Unequal class widths do not prevent calculation of the mean. Find the midpoint of each interval separately. Step-deviation can still be used with a suitable divisor; that divisor need not be the width of every interval, and scaled deviations need not all be integers.

The following distribution records wickets taken by 45 bowlers in one-day cricket matches. Its class widths vary. Choose a=200a=200 and h=20h=20, then calculate the class marks and scaled deviations. In particular, the midpoint of 100 to 150 is 125.

WicketsBowlers fif_iMidpoint xix_iScaled deviation uiu_iProduct fiuif_i u_i
20 to 607(20+60)/2=40(20+60)/2=40(40−200)/20=−8(40-200)/20=-87×(−8)=−567\times(-8)=-56
60 to 1005(60+100)/2=80(60+100)/2=80(80−200)/20=−6(80-200)/20=-65×(−6)=−305\times(-6)=-30
100 to 15016(100+150)/2=125(100+150)/2=125(125−200)/20=−3.75(125-200)/20=-3.7516×(−3.75)=−6016\times(-3.75)=-60
150 to 25012(150+250)/2=200(150+250)/2=200(200−200)/20=0(200-200)/20=012×(0)=012\times(0)=0
250 to 3502(250+350)/2=300(250+350)/2=300(300−200)/20=5(300-200)/20=52×(5)=102\times(5)=10
350 to 4503(350+450)/2=400(350+450)/2=400(400−200)/20=10(400-200)/20=103×(10)=303\times(10)=30

Worked example 5. Find the mean number of wickets from this unequal-width distribution.

  1. Use the midpoint and scaled-deviation calculations in the table. Sum the frequencies: ∑fi=7+5+16+12+2+3=45.\sum f_i=7+5+16+12+2+3=45.
  2. Add the weighted scaled deviations: ∑fiui=−56−30−60+0+10+30=−106.\sum f_i u_i=-56-30-60+0+10+30=-106.
  3. Substitute without dropping the negative sign: xˉ=200+20(−10645).\bar{x}=200+20\left(\frac{-106}{45}\right).
  4. Simplify, retaining the fraction until the final rounding: xˉ=200−212045≈152.89.\bar{x}=200-\frac{2120}{45}\approx152.89.

Answer: The average is approximately 152.89152.89 wickets per bowler.

The negative correction means that the mean is below the assumed mean. It is not a negative number of wickets. Interpret the final value in its context: it summarises the average for these bowlers, rather than the number taken by every individual bowler.

How are the modal class and mode found?

Definition: The mode is the observation value occurring most frequently. In grouped data, the class with the highest frequency is the modal class, within which the mode is calculated.

Finding the largest frequency identifies a class, not the numerical mode itself. The grouped formula also uses the frequencies of the classes immediately before and after that class. Keep these neighbouring frequencies in their correct positions.

Which quantities enter the mode formula?

For mode, let ll be the lower limit of the modal class and hh its class size, assuming equal class sizes. Let f1f_1 be the modal-class frequency, f0f_0 the preceding-class frequency and f2f_2 the succeeding-class frequency.

The grouped mode formula is Mode=l+f1−f02f1−f0−f2 h.\text{Mode}=l+\frac{f_1-f_0}{2f_1-f_0-f_2}\,h. The intervals must be continuous before this formula is applied. The calculation here concerns a distribution with a single mode.

A survey of 20 households gives the following family-size distribution. The highest frequency is eight, belonging to the interval 3 to 5. Seven households belong to the preceding interval, while two belong to the succeeding interval.

Family sizeHouseholds
1 to 37
3 to 58
5 to 72
7 to 92
9 to 111

Worked example 6. Find the mode of the family-size distribution.

  1. Identify the modal class as 3 to 5 and record its quantities: l=3,h=5−3=2,f1=8,f0=7,f2=2.l=3,\quad h=5-3=2,\quad f_1=8,\quad f_0=7,\quad f_2=2.
  2. Substitute into the formula: Mode=3+8−72×8−7−2×2.\text{Mode}=3+\frac{8-7}{2\times8-7-2}\times2.
  3. Simplify the numerator and denominator: Mode=3+17×2.\text{Mode}=3+\frac{1}{7}\times2.
  4. Finish the calculation: Mode=3+27≈3.286.\text{Mode}=3+\frac{2}{7}\approx3.286.

Answer: The grouped mode is approximately 3.2863.286 family members.

How can mode and mean differ?

For the grouped marks distribution used earlier, the largest frequency is seven in the interval 40 to 55. Its neighbouring frequencies are three and six. The mode describes the concentration of marks, while the mean combines contributions from the whole distribution.

Worked example 7. For marks intervals 10 to 25, 25 to 40, 40 to 55, 55 to 70, 70 to 85 and 85 to 100, the frequencies are respectively 2, 3, 7, 6, 6 and 6. Find the mode.

  1. Identify the modal class and parameters: l=40,h=55−40=15,f1=7,f0=3,f2=6.l=40,\quad h=55-40=15,\quad f_1=7,\quad f_0=3,\quad f_2=6.
  2. Substitute the values: Mode=40+7−32×7−3−6×15.\text{Mode}=40+\frac{7-3}{2\times7-3-6}\times15.
  3. Simplify the fraction: Mode=40+45×15.\text{Mode}=40+\frac{4}{5}\times15.
  4. Calculate the result: Mode=40+12=52.\text{Mode}=40+12=52.

Answer: The grouped mode is 5252 marks, below the grouped mean of 6262 marks.

Mode need not be below mean in every distribution; it may be equal to or greater than mean. Data may also be multimodal when more than one value shares the maximum frequency. Do not assume that every collection has a unique mode.

What does cumulative frequency tell us about the median?

The median gives the middle-most observation in ordered data. For ungrouped observations, first arrange the values in ascending order. Frequencies can then locate the middle position without writing each repeated observation separately.

Let nn denote the total number of observations. For odd totals, use the observation in position (n+1)/2(n+1)/2. For even totals, average the observations in positions n/2n/2 and n/2+1n/2+1. These expressions give positions, not the values at those positions.

How is a running total formed?

A cumulative frequency combines frequencies up to a specified point. In an ascending table of individual marks, it records how many students have that mark or a lower mark. Add the current frequency to the total already accumulated.

The following marks, out of 50, belong to 100 students. The cumulative total reaches 50 at 28 marks and then reaches 78 at 29 marks. These totals identify the two central observations without any interpolation within a class interval.

MarksStudentsCumulative frequency
2060+6=60+6=6
25206+20=266+20=26
282426+24=5026+24=50
292850+28=7850+28=78
331578+15=9378+15=93
38493+4=9793+4=97
42297+2=9997+2=99
43199+1=10099+1=100

Worked example 8. Find the median of the 100 individual marks in the table.

  1. Read the final cumulative frequency: n=100.n=100.
  2. Locate the two middle positions because the total is even: n/2=50,n/2+1=51.n/2=50,\qquad n/2+1=51.
  3. The running totals show that the 50th observation is 28 and the 51st is 29.
  4. Average these values: Median=28+292=572=28.5.\text{Median}=\frac{28+29}{2}=\frac{57}{2}=28.5.

Answer: The median is 28.528.5 marks.

How do less-than and more-than distributions differ?

For grouped data, a less-than distribution attaches cumulative frequencies to upper class limits. Start with the first frequency and add successive frequencies. It counts observations below each upper limit, with observations at that limit assigned to the next class.

A more-than distribution attaches totals to lower class limits. Begin with the total frequency and subtract the frequencies of classes already passed. In the marks example with 53 students, subtracting the five below 10 leaves 53−5=4853-5=48 at or above 10.

These tables describe the same observations from opposite ends. Distinguish a class frequency, which counts observations in one interval, from a cumulative frequency, which includes observations across several intervals. This distinction is essential when selecting quantities for the median formula.

How is the median of grouped data calculated?

Grouped data place the middle observation inside an interval, so the exact individual observation may not be visible. Use cumulative frequencies to identify the median class, then calculate a value inside that class using the grouped median formula.

Which class and quantities should be selected?

Find half the total frequency. Select the class whose cumulative frequency is greater than, and nearest to, that half-total. For median, ll denotes the lower limit of that class, ff its frequency, and hh its width.

Let cf\mathrm{cf} denote the cumulative frequency of the class immediately preceding the median class. The total number of observations is nn, as before. With continuous intervals and the equal-width setting used here, the formula is Median=l+n/2−cff h.\text{Median}=l+\frac{n/2-\mathrm{cf}}{f}\,h.

The subtraction removes the observations already counted before the median class. The denominator is the frequency within the median class, not its cumulative total. Substituting the wrong one changes the portion of the class width added to the lower limit.

How are ordinary frequencies recovered from cumulative data?

A height survey of 51 girls gives the less-than totals below. Subtract consecutive totals to recover each class frequency. Keep the first class as below 140 cm; there is no need to invent a lower limit for this open first interval.

Height (cm)Given cumulative frequencyRecovered class frequency
Below 140444
140 to 1451111−4=711-4=7
145 to 1502929−11=1829-11=18
150 to 1554040−29=1140-29=11
155 to 1604646−40=646-40=6
160 to 1655151−46=551-46=5

Worked example 9. Find the median height from the less-than height distribution.

  1. Recover the class frequencies by the successive subtractions shown in the table. The final cumulative total gives n=51,n/2=51/2=25.5.n=51,\qquad n/2=51/2=25.5.
  2. The first cumulative frequency above 25.5 is 29, so the median class is 145 to 150 cm. Record l=145,cf=11,f=18,h=150−145=5.l=145,\quad\mathrm{cf}=11,\quad f=18,\quad h=150-145=5.
  3. Substitute into the formula: Median=145+25.5−1118×5.\text{Median}=145+\frac{25.5-11}{18}\times5.
  4. Simplify the numerator and multiply: Median=145+14.5×518=145+72.518.\text{Median}=145+\frac{14.5\times5}{18}=145+\frac{72.5}{18}.
  5. Complete the division and round: Median≈149.03 cm.\text{Median}\approx149.03\text{ cm}.

Answer: The median height is approximately 149.03149.03 cm. About half the girls are shorter and about half are taller than this height.

The median class and the modal class answer different questions. The median class is selected using cumulative position. The modal class is selected using the largest ordinary frequency. A calculation should state the selection rule before substituting values.

How can a known median determine missing frequencies?

A missing-frequency question reverses the usual calculation. Instead of using a completed table to find the median, use the given median to form an equation. The stated total frequency supplies another equation when two frequencies are unknown.

Let xx and yy denote the unknown frequencies in the intervals 200 to 300 and 600 to 700, respectively. They are counts of observations, not class marks. The median is 525 and the total frequency is 100 in this distribution.

Class intervalFrequency
0 to 1002
100 to 2005
200 to 300xx
300 to 40012
400 to 50017
500 to 60020
600 to 700yy
700 to 8009
800 to 9007
900 to 10004

How are the two equations formed and solved?

First add every frequency, keeping unknowns as symbols. Next identify the class containing the stated median. Build the cumulative frequency before this class with particular care: it includes the first unknown but excludes the frequency of the median class itself.

Worked example 10. Find both missing frequencies in the table, given median 525 and total frequency 100.

  1. Add all frequencies: 2+5+x+12+17+20+y+9+7+4=100.2+5+x+12+17+20+y+9+7+4=100. Therefore 76+x+y=100,x+y=24.76+x+y=100,\qquad x+y=24.
  2. The median lies in 500 to 600. Its lower limit, frequency, width and preceding cumulative frequency are l=500,f=20,h=100,cf=2+5+x+12+17=36+x.l=500,\quad f=20,\quad h=100,\quad\mathrm{cf}=2+5+x+12+17=36+x.
  3. Use half the total frequency and substitute: n/2=100/2=50,525=500+50−(36+x)20×100.n/2=100/2=50,\qquad 525=500+\frac{50-(36+x)}{20}\times100.
  4. Subtract the lower limit and simplify: 525−500=5(14−x),25=70−5x.525-500=5(14-x),\qquad25=70-5x.
  5. Isolate the first unknown: 5x=70−25=45,x=45/5=9.5x=70-25=45,\qquad x=45/5=9.
  6. Use the total-frequency equation: 9+y=24,y=24−9=15.9+y=24,\qquad y=24-9=15.

Answer: The missing frequencies are x=9x=9 and y=15y=15.

Check the result in both conditions. Adding the recovered frequencies must give the stated total. Substituting the resulting preceding cumulative frequency into the median formula must also reproduce the given median. Satisfying just one condition is insufficient when the problem supplies two.

  1. Check the total: 76+9+15=100.76+9+15=100.
  2. Check the preceding cumulative frequency: cf=36+9=45.\mathrm{cf}=36+9=45.
  3. Check the median: 500+50−4520×100=500+25=525.500+\frac{50-45}{20}\times100=500+25=525.

How should class boundaries and the choice of average be handled?

Why must median and mode classes be continuous?

Before applying the grouped median or mode formula, ensure that the class intervals are continuous. Some tables use inclusive-looking intervals for measurements rounded to a stated unit. Their written limits need adjustment before the formula is used.

For leaf lengths measured to the nearest millimetre, the intervals 118 to 126, 127 to 135 and 136 to 144 become 117.5 to 126.5, 126.5 to 135.5 and 135.5 to 144.5. Use the converted lower boundary and width in the calculation.

The adjustment makes adjacent boundaries meet while retaining the recorded frequencies. Do not remove observations or alter their counts during conversion. In this example, the remaining intervals are treated in the same way, ending with 171.5 to 180.5.

Which measure answers the question being asked?

The mean takes all observations into account and lies between the smallest and largest observations. It is frequently used to compare distributions, such as average results from different schools. However, extreme values affect it, so it may not represent a typical observation well.

The median is more appropriate when the interest is in a typical observation and extreme values may be present. Examples include typical worker productivity or average wages. The position of the middle observation is central to this interpretation.

The mode is useful when the interest is the most frequent value or most popular item. Examples include a television programme watched most often, a consumer item in greatest demand, or a vehicle colour used by most people.

What is the empirical relationship between the averages?

An empirical relationship connects the three measures: 3 Median=Mode+2 Mean.3\,\text{Median}=\text{Mode}+2\,\text{Mean}. Treat this as an empirical relationship, rather than an exact identity guaranteed for every collection of observations. The direct formulas describe how the measures are calculated from a distribution.

A useful final interpretation states which average has been found and what it represents. Keep the unit or context with the answer: marks, height in centimetres, or wickets per bowler. For grouped calculations, remember the information lost when individual observations were placed into intervals.

Glossary

  • Mean — Sum of observation values divided by the total number of observations in the distribution.
  • Median — Central value found from the positions of observations after arranging the data in order.
  • Mode — Observation value occurring most frequently, or a value calculated within the modal class for grouped data.
  • Frequency — Number of observations corresponding to a particular value or falling within a particular class interval.
  • Class interval — Range used to group observations between specified lower and upper class limits.
  • Class mark — Midpoint of a class interval, obtained by averaging its lower and upper limits.
  • Assumed mean — Convenient fixed value subtracted from class marks to simplify the calculation of the mean.
  • Deviation — Signed difference obtained by subtracting the assumed mean from a class mark.
  • Step-deviation — Deviation divided by a chosen non-zero divisor to simplify frequency-weighted arithmetic.
  • Cumulative frequency — Running frequency total counting observations up to, or from, a specified value or class boundary.
  • Modal class — Class interval having the greatest frequency in a grouped distribution with a single modal class.
  • Median class — Class located using the first cumulative frequency greater than half the total number of observations.

Common errors and misconceptions

  • Misconception: Divide the weighted total by the number of classes. Correct: Divide it by the sum of frequencies, which counts all observations.
  • Misconception: The grouped mean must equal the exact mean. Correct: The midpoint assumption can change the result when individual observations are condensed into classes.
  • Misconception: Negative deviations should be made positive. Correct: Keep their signs when multiplying and adding; the correction may lower the assumed mean.
  • Misconception: Add the mean scaled deviation straight to the assumed mean. Correct: Multiply it by the chosen divisor first, restoring the scale of the original values.
  • Misconception: The highest frequency itself is the mode. Correct: It identifies the modal class; calculate the grouped mode using that class and its neighbours.
  • Misconception: Use the cumulative frequency of the median class in the numerator. Correct: Subtract the cumulative frequency of the preceding class from half the total frequency.
  • Misconception: Discontinuous class limits can be used unchanged for grouped median or mode. Correct: Convert to continuous intervals before applying these formulas.

Exam-style questions with model answers

Q1. What is a class mark, and what assumption is used when calculating a grouped mean? [2 marks]
  1. A class mark is the midpoint of an interval: xi=upper limit+lower limit2,x_i=\frac{\text{upper limit}+\text{lower limit}}{2}, where xix_i denotes that midpoint.
  2. The frequency of each class is assumed to be centred at its midpoint, which represents observations in that class.
Q2. Marks intervals 10 to 25, 25 to 40, 40 to 55, 55 to 70, 70 to 85 and 85 to 100 have frequencies 2, 3, 7, 6, 6 and 6 respectively. Find the mean by the direct method. [4 marks]
  1. Let xix_i be a class mark and fif_i its frequency. Average the limits of each interval: (10+25)/2=17.5, (25+40)/2=32.5, (40+55)/2=47.5.(10+25)/2=17.5,\ (25+40)/2=32.5,\ (40+55)/2=47.5. (55+70)/2=62.5, (70+85)/2=77.5, (85+100)/2=92.5.(55+70)/2=62.5,\ (70+85)/2=77.5,\ (85+100)/2=92.5.
  2. Multiply each midpoint by its frequency: 2(17.5)=35, 3(32.5)=97.5, 7(47.5)=332.5.2(17.5)=35,\ 3(32.5)=97.5,\ 7(47.5)=332.5. 6(62.5)=375, 6(77.5)=465, 6(92.5)=555.6(62.5)=375,\ 6(77.5)=465,\ 6(92.5)=555.
  3. Add frequencies and products separately: ∑fi=2+3+7+6+6+6=30.\sum f_i=2+3+7+6+6+6=30. ∑fixi=35+97.5+332.5+375+465+555=1860.\sum f_i x_i=35+97.5+332.5+375+465+555=1860.
  4. Divide the weighted total by the student total: xˉ=∑fixi∑fi=186030=62.\bar{x}=\frac{\sum f_i x_i}{\sum f_i}=\frac{1860}{30}=62. Thus the grouped mean is 62 marks; it uses class midpoints to represent the scores.
Q3. Family-size classes 1 to 3, 3 to 5, 5 to 7, 7 to 9 and 9 to 11 have frequencies 7, 8, 2, 2 and 1 respectively. Calculate the mode. [3 marks]
  1. The greatest frequency is eight, so the modal class is 3 to 5. Its lower limit is l=3l=3, and its width is h=5−3=2h=5-3=2.
  2. The modal frequency is f1=8f_1=8, the preceding frequency is f0=7f_0=7, and the succeeding frequency is f2=2f_2=2. Substitute: Mode=l+f1−f02f1−f0−f2h=3+8−716−7−2×2.\text{Mode}=l+\frac{f_1-f_0}{2f_1-f_0-f_2}h=3+\frac{8-7}{16-7-2}\times2.
  3. Simplify before dividing: Mode=3+17×2=3+27≈3.286.\text{Mode}=3+\frac{1}{7}\times2=3+\frac{2}{7}\approx3.286. The grouped mode is approximately 3.286 family members, lying within the modal interval.
Q4. The numbers of girls shorter than 140, 145, 150, 155, 160 and 165 cm are respectively 4, 11, 29, 40, 46 and 51. Find the median height. [5 marks]
  1. These are cumulative frequencies. Subtract consecutive totals to obtain class frequencies: 4,11−4=7,29−11=18,40−29=11,46−40=6,51−46=5.4,\quad11-4=7,\quad29-11=18,\quad40-29=11,\quad46-40=6,\quad51-46=5. The first class remains below 140 cm.
  2. The final cumulative frequency gives the total number of girls, n=51n=51. Half the total is n/2=51/2=25.5.n/2=51/2=25.5. Use this value to locate the median class.
  3. The first cumulative total greater than 25.5 is 29, attached to the class 145 to 150 cm. Thus the lower limit is l=145l=145 and the class width is h=150−145=5h=150-145=5.
  4. The median-class frequency is f=18f=18, while the preceding cumulative frequency is cf=11\mathrm{cf}=11. Substitute in the grouped formula: Median=l+n/2−cffh=145+25.5−1118×5.\text{Median}=l+\frac{n/2-\mathrm{cf}}{f}h=145+\frac{25.5-11}{18}\times5.
  5. Simplify and round only at the end: Median=145+14.5×518=145+72.518≈149.03 cm.\text{Median}=145+\frac{14.5\times5}{18}=145+\frac{72.5}{18}\approx149.03\text{ cm}. About half the girls are shorter and about half taller than this height.
Q5. The successive classes 0 to 100, 100 to 200, 200 to 300, 300 to 400, 400 to 500, 500 to 600, 600 to 700, 700 to 800, 800 to 900 and 900 to 1000 have frequencies 2,5,x,12,17,20,y,9,7,42,5,x,12,17,20,y,9,7,4. Here xx and yy are missing frequencies. The total is 100 and the median is 525. Find both missing frequencies. [5 marks]
  1. Adding all frequencies gives 2+5+x+12+17+20+y+9+7+4=100.2+5+x+12+17+20+y+9+7+4=100. The known frequencies total 76, so 76+x+y=100,x+y=24.76+x+y=100,\qquad x+y=24. This is the first condition on the unknown counts.
  2. The given median lies in 500 to 600. Its lower limit is l=500l=500, its frequency is f=20f=20, and its width is h=100h=100. The preceding cumulative frequency is cf=2+5+x+12+17=36+x.\mathrm{cf}=2+5+x+12+17=36+x.
  3. The total count is n=100n=100, giving n/2=50n/2=50. Substitute in the median formula: 525=500+50−(36+x)20×100.525=500+\frac{50-(36+x)}{20}\times100. Subtract 500 and simplify to obtain 25=5(14−x)=70−5x.25=5(14-x)=70-5x.
  4. Isolate the first missing frequency: 5x=70−25=45,x=45/5=9.5x=70-25=45,\qquad x=45/5=9. Then substitute into the total-frequency condition: 9+y=24,y=15.9+y=24,\qquad y=15.
  5. Check both original conditions using the recovered counts: 76+9+15=100,cf=36+9=45.76+9+15=100,\qquad\mathrm{cf}=36+9=45. The median is 500+50−4520×100=525.500+\frac{50-45}{20}\times100=525. Both checks hold, so the missing frequencies are nine and fifteen respectively.
Q6. Explain when mean, median and mode are suitable measures of central tendency. [3 marks]
  1. The mean uses all observations and can compare distributions, such as the average results of schools. Extreme observations can affect it, so its usefulness depends on the distribution.
  2. The median is suitable when a typical observation is wanted and extreme values may be present, such as in wages or worker productivity. It represents the middle of ordered data.
  3. The mode is suitable when the most frequent value or most popular item is required, such as the television programme watched most often or a consumer item in greatest demand.

Key takeaways

  • Use class midpoints to represent grouped observations, remembering that grouping can make the resulting mean approximate.
  • The direct, assumed mean and step-deviation methods give the same mean for the same grouped frequency distribution.
  • Keep deviation signs, multiply by frequencies, and restore every transformation before interpreting the final mean.
  • The modal class has the largest ordinary frequency; its neighbours help determine the grouped mode.
  • Cumulative frequencies locate the median class by showing how many observations have already been counted.
  • Use the preceding cumulative frequency and the median-class frequency as separate quantities in the median formula.
  • Check continuity of intervals before calculating grouped median or mode, adjusting boundaries when the data require it.
  • Choose mean for overall averaging, median for a typical middle observation, and mode for the most frequent value.

Test yourself

Which interval contains a score of 40: 25 to 40 or 40 to 55?

The score belongs to 40 to 55 under the convention that assigns an upper-limit observation to the next class.

Why can an exact mean differ from the grouped mean of the same observations?

The grouped calculation replaces individual observations by class midpoints, introducing the midpoint assumption and potentially changing the mean.

Does changing the assumed mean change the final mean?

No. With correct signed deviations and restoration of the assumed mean, the final mean remains the same.

Can step-deviation be used when class widths are unequal?

Yes. Calculate each midpoint separately and choose a suitable non-zero divisor for the deviations; the divisor need not equal every class width.

What is the difference between modal class and mode?

The modal class is the highest-frequency interval. The grouped mode is a calculated numerical value within that interval.

Which cumulative frequency is substituted in the median formula?

Use the cumulative frequency of the class immediately before the median class, rather than that of the median class itself.

How are class frequencies recovered from a less-than cumulative distribution?

Keep the first cumulative total as the first frequency, then subtract each preceding cumulative total from the next one.

Why should median and mode calculations begin with a check of class intervals?

The grouped formulas require continuous intervals, so discontinuous written limits must be converted before substitution.