Model G20 2027 at FLAME University, registrations now open

Data Handling | ICSE Class 7 Maths Notes

24 min read

On this page

This note covers collecting and organising data, frequency tables, representative values, arithmetic mean, median, mode, constructing and interpreting bar graphs, and exploring chance through coin and dice experiments.

What is data, and how do we choose what to collect?

Data is a collection of facts, numbers, measurements or observations that gives information. An observation is an individual recorded item. Heights, runs scored in matches and people's choices of favourite games are examples of data.

Statistics is the study of collecting, organising, analysing, interpreting and presenting data. Analysing means examining the information for patterns; interpreting means explaining what those patterns tell us. Collecting suitable information comes before performing calculations or drawing a graph.

What question will the data answer?

A statistical question can be answered by collecting data. Asking how tall Grade 7 students in a school are requires measurements. The heights need not all be the same, so a single student's height does not describe the whole group.

A hypothesis is a statement proposed for checking. For example, Navya thinks cricket is the most popular game in her class. To check this, she and Naresh ask each classmate which game they prefer, then organise the responses.

The data collected must match the question. To investigate favourite games, record choices of games. To investigate heights, record heights. Choosing the information to collect is part of the investigation, not merely something done before the mathematical work begins.

How does an investigation proceed?

  1. State the question or proposed claim clearly, including the group being investigated.
  2. Decide which observations would help answer that question.
  3. Collect and record the observations, keeping each response attached to its intended category.
  4. Organise the information so that counts, comparisons and patterns can be examined.
  5. Explain the conclusion supported by the collected data.

A category is a named group, such as a particular game or sweet. A conclusion about one class describes that class's recorded responses. Data from a small group does not, by itself, justify a claim about all children.

How do tables and tally marks organise observations?

Ungrouped data records individual values without combining them into numerical intervals. Such a list can be rearranged or counted to make it easier to read. Organisation preserves the observations while making particular features of the information easier to see.

The frequency of a value or category is the number of times it occurs. A frequency table places values or categories beside their counts. This allows comparisons without repeatedly counting the original responses.

How are tally marks used?

A tally mark is a counting stroke used when recording an observation. Write one stroke for each occurrence. The fifth stroke crosses the previous four, making a group of five. Completed groups and remaining strokes can then be counted together.

For sweet preferences, list each sweet in its own row. Add the appropriate tally when a student names that sweet. Finally, write the frequency as a number. The finished table shows how many students chose each sweet.

Table: Sweet preferences

SweetNumber of students
Jalebi6
Gulab jamun9
Gujiya13
Barfi3
Rasgulla7

The table helps someone buy the required numbers of sweets. However, counts alone do not identify which student chose which sweet. Keep the individual responses if the sweets must later be distributed to the correct people.

What does sorting reveal?

Ascending order means arranging numbers from smallest to largest. Repeated values remain repeated: sorting changes their positions, not how many observations there are. This is particularly useful when locating the middle of a numerical list.

The minimum is the smallest value and the maximum is the largest. Sorting places these at opposite ends. A frequency table highlights repetition instead. Choose the arrangement according to whether the question concerns order, frequency or individual identity.

Note: A frequency and a data value have different meanings. In the sweet table, Gujiya is the category and 13 is its frequency. The count describes how many responses belong to that category.

What does the arithmetic mean represent?

A representative value is a single value used to summarise a collection of observations. Finding one is useful in most daily-life situations involving data. However, the meaning of the chosen value matters: different summaries need not give the same result.

Definition: The arithmetic mean, or simply mean, is the sum of the numerical observations divided by the number of observations. Here, average means arithmetic mean unless another measure is named.

Result: The mean is an equal-share value

Mean = sum of observations ÷ number of observations. The sign = means “is equal to”, and ÷ means division. The sum is the total obtained by adding every observation. Count repetitions separately when finding the number of observations.

Imagine sharing a total equally across all members of a group. The amount received by each member is the mean. This interpretation explains both parts of the calculation: first find the total, then divide by the number of shares.

Worked example 1. Shreyas and four friends collect 3, 8, 10, 5 and 4 guavas. Find each person's equal share.

Answer: The sign + means addition. Total guavas = 3 + 8 + 10 + 5 + 4 = 30. There are five people, so mean = 30 ÷ 5 = 6 guavas per person.

Worked example 2. Parag and five friends collect 5, 4, 6, 3, 4 and 8 guavas. Find each person's equal share.

Answer: Total guavas = 5 + 4 + 6 + 3 + 4 + 8 = 30. There are six people, so mean = 30 ÷ 6 = 5 guavas per person.

Both groups collect the same total, but their equal shares differ because the group sizes differ. Shreyas's group receives one more guava per person. Comparing totals alone would hide this difference; the mean takes account of the number of people.

The original collections do not have to equal the mean. Equal sharing describes how the total could be redistributed. It does not claim that each person originally collected the same number of guavas.

How do we calculate and interpret a mean correctly?

Begin by deciding what each observation represents. It might be a day's flower count or the number of bounces in one attempt. The divisor, meaning the number used to divide the total, must count those observations.

What steps keep the calculation clear?

  1. Copy the full list and identify what one observation represents.
  2. Add every value, including repeated values.
  3. Count the observations in the original list.
  4. Divide the total by that count and state the meaning of the answer.

Worked example 3. Vaishnavi records 2, 7, 9, 4 and 3 hibiscus flowers blooming on five days. Calculate the average daily number.

Answer: Total flowers = 2 + 7 + 9 + 4 + 3 = 25. Mean = 25 ÷ 5 = 5 flowers per day. This is the daily count if the total were spread equally across the five days.

The calculated average is 5, although none of the five daily counts is 5. A mean therefore need not be one of the recorded observations. It summarises the whole collection rather than selecting a particular day's value.

Worked example 4. Shreyas records 6, 2, 9, 5, 4, 6, 3 and 5 bounces in eight attempts. Find the average number of bounces per attempt.

Answer: Total bounces = 6 + 2 + 9 + 5 + 4 + 6 + 3 + 5 = 40. Mean = 40 ÷ 8 = 5 bounces per attempt. Both occurrences of 6 and both occurrences of 5 are included.

Does zero count as an observation?

Zero is a recorded numerical result, not a missing observation. In match data, scoring zero still records a match played. “Did not play” has a different meaning: no score was made in that game, so it is not a played-game observation.

This distinction affects the divisor. Before calculating an average per game played, check whether an entry is a recorded zero or an absence from the game. Do not convert a missing performance into an invented score.

How do we find the median of a small data set?

Definition: The median is the middle value of sorted numerical data. When the number of observations is even, it is the mean of the two middle values.

An odd number of observations leaves one middle position. An even number leaves two middle positions. Sorting means arranging the values in numerical order. The middle of an unsorted list does not reliably locate the median.

Result: Find the middle after sorting

Write the observations in ascending order without removing repetitions. Count how many there are. With an odd count, select the middle observation. With an even count, add the two middle observations and divide their sum by two.

Worked example 5. Poovizhi's family members have heights 170, 173, 165, 118 and 175 centimetres. Find the median height. The abbreviation cm means centimetres.

Answer: Sorted heights are 118, 165, 170, 173 and 175 cm. There are five observations, so the third is the middle one. The median height is 170 cm.

There are two positions before the middle and two after it. Counting positions explains why the third observation is selected. Selecting the middle position is different from averaging the smallest and largest values.

Worked example 6. Yaangba's family members have heights 169, 173, 155, 165, 160 and 164 cm. Find the median height.

Answer: Sorted heights are 155, 160, 164, 165, 169 and 173 cm. The middle values are 164 and 165 cm. Median = (164 + 165) ÷ 2 = 164.5 cm. Parentheses show the addition to complete before dividing.

With six observations, neither of the two central positions can stand alone as the middle. Averaging their values supplies the median. Here 164.5 cm is not a recorded height, which is possible for an even-sized data set.

Keep the unit attached to the final result. The calculation concerns heights, so the median is a height measured in centimetres. Sorting the observations does not change what they measure.

How do mode, mean and median differ?

The mode is a value with the greatest frequency in a data set. It answers a question about what occurs most often. Unlike finding a mean, finding a mode requires counting occurrences rather than adding all the observed values.

How is a mode found?

List each distinct value and count how often it appears. “Distinct” means different: list a value once in the table, while its frequency records all its appearances. Compare the frequencies and select the value or values with the greatest count.

Worked example 7. A newspaper has 16, 18, 20, 22, 26, 16 and 10 pages from Monday to Sunday. Find the mode.

Answer: The value 16 occurs twice. Every other listed page count occurs once. The mode is therefore 16 pages. The greatest page count, 26, is not the mode.

More than one value may share the greatest frequency. For the bounce counts 6, 2, 9, 5, 4, 6, 3, 5, both 5 and 6 occur twice, while the other values occur once. Both are modes.

Table: Comparing representative values

MeasureHow it is foundWhat it represents
MeanAdd all values and divide by their countAn equal-share value
MedianSort and find the middle, averaging two middle values when neededThe centre by position
ModeCompare frequencies of valuesA most frequently occurring value

Can a mean give an unhelpful summary?

An outlier is a value that differs significantly from the rest. The mean may not always be an appropriate representative of data with outliers, because a very high or very low value can significantly affect the total.

In Poovizhi's height data, 118 cm is much lower than the other heights. The mean is 160.2 cm, while the median is 170 cm. The mean is below four of the five recorded heights and does not seem to represent this data very well.

This example does not make the median the best choice for every question. Each measure answers a different question. Interpret the selected measure in relation to the observations and the purpose of the comparison.

How do we construct a simple bar graph?

A bar graph represents numbers with rectangular bars. Their heights, for vertical bars, or lengths, for horizontal bars, correspond to the numbers represented. Bars have uniform widths and equal gaps so that their values can be compared clearly.

An axis is a reference line along which labels or numerical values are marked. The plural is axes. A vertical bar graph places categories along the horizontal axis and their numerical values along the vertical axis.

Result: Equal scale intervals represent equal numerical differences

The scale states how much one unit length represents. A unit length is one chosen equal interval on the graph. Equal physical intervals along the numerical axis must represent equal changes in value.

Choose a scale that fits the values on the page. For small counts, one unit length can represent one student. For larger scores, one unit length can represent ten runs. Write the scale so that someone else can read the bars correctly.

What is the construction method?

  1. Draw horizontal and vertical axes and label what each represents.
  2. Place the categories at equal spacing along the horizontal axis.
  3. Choose a suitable scale and mark the numerical axis at equal intervals, starting at zero for these graphs.
  4. Draw bars of equal width with equal gaps, reaching the required numerical heights.
  5. Add a title and check every bar against its original value.

Worked example 8. Sweet preferences are Jalebi 6, Gulab jamun 9, Gujiya 13, Barfi 3 and Rasgulla 7 students. Describe the bars using one unit length for one student.

Answer: Draw five equally spaced bars with heights 6, 9, 13, 3 and 7 units in the listed order. Label the horizontal axis “Sweets” and the vertical axis “Number of students”. Gujiya has the tallest bar.

What the figure shows

Sweet preferences of students

Five vertical bars show Jalebi 6, Gulab jamun 9, Gujia 13, Barfi 3 and Rasgulla 7 students. The vertical scale runs from 0 to 14 in steps of 1.

Reference: NCERT Class 6, page 90, unnumbered graph

The graph's Gujia label names the category written as Gujiya in the table. A bar graph makes the difference between the counts visible. It does not show which individual students supplied the responses.

How do we read bar graphs and compare their values?

Read the title, axis labels and scale before comparing bars. The unit tells what is being counted or measured, such as students or runs. A height on paper has meaning only after it is connected to the numerical scale.

How does the attendance graph work?

Table: Students absent in each class

ClassNumber absent
I3
II5
III4
IV2
V0
VI1
VII5
VIII7

The class labels I to VIII mean Classes 1 to 8. Read the row or bar belonging to the requested class. A zero absence count means full attendance for that class on that day; it does not mean the class has no students.

What the figure shows

Students absent

Classes 1 to 8 are labelled horizontally. The vertical axis counts students, with one unit length representing one student. Class 8 reaches 7; Class 5 has zero height.

Reference: NCERT Class 6, page 86, unnumbered graph

Worked example 9. In the absence data 3, 5, 4, 2, 0, 1, 5, 7 for Classes I to VIII respectively, identify the greatest absence and the classes with equal absence counts of 5.

Answer: Class VIII has the greatest absence count, 7 students. Classes II and VII each have 5 students absent. Their bars therefore have equal heights on the same graph.

How does a larger scale change bar heights?

Smriti's scores in Matches 1 to 8 are 80, 50, 10, 100, 90, 0, 90 and 50 runs. Using one unit length for ten runs keeps the graph manageable while retaining the values of the scores.

Worked example 10. At a scale of one unit length for 10 runs, how high are the bars for Smriti's scores of 100 and 50 runs?

Answer: A score of 100 runs requires 100 ÷ 10 = 10 units of height. A score of 50 runs requires 50 ÷ 10 = 5 units. Multiply a bar's height by 10 to read its score.

What the figure shows

Runs scored by Smriti

Matches 1 to 8 appear horizontally, with runs marked vertically in tens. Match 4 reaches 100 runs, Match 6 has zero height, and Matches 5 and 7 both reach 90.

Reference: NCERT Class 6, page 91, unnumbered graph

State comparisons using the values represented. Equal bars show equal scores; the tallest bar shows the greatest score in this set. The graph records these eight matches and does not establish what will happen in a later match.

How do coin and dice experiments help us understand chance?

Chance concerns whether something may happen when its result is uncertain. Probability describes likelihood. At this stage, experiments help us explore this idea by recording what happens repeatedly and comparing the counts of different results.

An experiment is an activity performed to observe a result. A trial is one repetition of it, such as one coin toss. An outcome is the result of a trial. An event is a specified outcome or set of outcomes being considered.

What can we record?

A coin toss can be recorded as heads or tails, the names for its two faces. An ordinary die is a cube with faces numbered 1 to 6; dice is the plural of die. Record the number on the upper face after each throw.

A sequence is a list kept in the order in which results occur. Preserve the sequence of tosses or throws as well as counting the frequencies. The sequence shows the order, while a frequency table shows how often each outcome occurred.

  1. Choose the experiment, such as tossing a coin or throwing a die.
  2. Record the outcome immediately after each trial, keeping the trials in order.
  3. Make a tally for heads and tails, or for each die number from 1 to 6.
  4. Count the tallies and check that all trials have been included.
  5. Compare the frequencies and examine the sequence for repetitions and changes.

Does randomness require alternation?

Randomness means that the result of an individual trial is not known with certainty beforehand. It does not require heads and tails to alternate. Repeated outcomes can occur, and a short experiment need not give equal counts for different outcomes.

For a fair coin, heads and tails are equally likely: neither is favoured. A fair die similarly favours none of its six faces. “Equally likely” describes chances, not a promise that a short recorded sequence will contain equal frequencies.

Compare coin results with dice results by recognising their different possible outcomes. Use the counts to discuss what happened in the experiment. Keep a prediction, meaning a statement about what may happen next, distinct from an already recorded observation.

Glossary

  • Data — A collection of facts, numbers, measurements or observations conveying information about something being investigated.
  • Observation — An individual recorded item or result belonging to a collection of data.
  • Hypothesis — A proposed statement that can be checked by collecting and examining relevant data.
  • Ungrouped data — Individual observations recorded without combining their values into numerical intervals.
  • Frequency — The number of times a particular value or category occurs in recorded data.
  • Arithmetic mean — The sum of all numerical observations divided by the number of observations.
  • Median — The middle value of sorted data, averaging the two middle values when necessary.
  • Mode — A data value whose frequency is greatest in the collection of observations.
  • Outlier — A value that differs significantly from the other values in a data set.
  • Bar graph — A representation using equal-width bars whose lengths or heights correspond to numerical values.
  • Scale — The rule stating the numerical amount represented by one unit length on a graph.
  • Trial — One repetition of an experiment, such as one coin toss or die throw.
  • Outcome — The result recorded when a particular trial of an experiment is completed.
  • Randomness — Uncertainty about an individual trial's result before that trial is carried out.

Common errors and misconceptions

  • Misconception: Divide the total by the number of different values to find the mean. Correct: Divide by the total number of observations, counting repeated values separately.
  • Misconception: The mean must appear in the data. Correct: The flower counts 2, 7, 9, 4 and 3 have mean 5, although 5 is not recorded.
  • Misconception: The middle entry in any list is its median. Correct: Sort the data first, keeping repetitions, then locate the middle position or positions.
  • Misconception: With an even number of observations, choose either central value. Correct: The median is the mean of the two middle values in the sorted list.
  • Misconception: The mode is the greatest number. Correct: It is a most frequent value. The greatest numerical value may appear less often than another value.
  • Misconception: A zero score means the player did not play. Correct: Zero is a recorded result from a played game; “did not play” supplies no played-game score.
  • Misconception: A bar five units high represents five items. Correct: Read the scale. At ten runs per unit length, that height represents fifty runs.
  • Misconception: Equally likely outcomes must occur equally often in a short experiment. Correct: Actual counts can differ, and random sequences can include repeated outcomes.

Exam-style questions with model answers

Q1. Five daily flower counts are 2, 7, 9, 4 and 3. Calculate the mean number of flowers per day. [2 marks]
  1. Add the five counts: 2 + 7 + 9 + 4 + 3 = 25 flowers.
  2. Divide by the five days: mean = 25 ÷ 5 = 5 flowers per day.
Q2. The heights of five family members are 170, 173, 165, 118 and 175 cm. Find the median and explain how you locate it. [3 marks]
  1. Arrange all five heights in ascending order: 118, 165, 170, 173 and 175 cm. Sorting is necessary before identifying the middle observation.
  2. There are five observations, an odd number, so there is one middle position. The third position has two observations before it and two after it.
  3. The third height is 170 cm. Therefore, the median height of the five family members is 170 cm.
Q3. Six heights are 169, 173, 155, 165, 160 and 164 cm. Find their median, showing the sorted data and the two middle values. [3 marks]
  1. Write the six observations in ascending order: 155, 160, 164, 165, 169 and 173 cm. No height is removed when the list is rearranged.
  2. There are six observations, so the third and fourth positions are central. Their values are 164 and 165 cm respectively.
  3. Take the mean of these two heights: median = (164 + 165) ÷ 2 = 164.5 cm. This calculated height need not occur in the original list.
Q4. Newspaper page counts from Monday to Sunday are 16, 18, 20, 22, 26, 16 and 10. Find the mode and distinguish it from the maximum. [2 marks]
  1. The mode is 16 pages, because it occurs twice while every other page count occurs once.
  2. The maximum is 26 pages. It is the greatest count, whereas the mode is the most frequent count.
Q5. Five people collect 3, 8, 10, 5 and 4 guavas; a second group of six people collects 5, 4, 6, 3, 4 and 8. Calculate both equal shares, compare them, and explain why the totals alone are insufficient. [5 marks]
  1. The first group's total is 3 + 8 + 10 + 5 + 4 = 30 guavas. All five people's collections are included.
  2. There are five people in the first group. Their equal share is the mean, calculated as 30 ÷ 5 = 6 guavas per person.
  3. The second group's total is 5 + 4 + 6 + 3 + 4 + 8 = 30 guavas. Both recorded collections of 4 are included.
  4. There are six people in the second group. Its equal share is 30 ÷ 6 = 5 guavas per person.
  5. The first group's share is one guava greater per person. Equal totals do not imply equal shares when the numbers of people are different.
Q6. Sweet preferences are Jalebi 6, Gulab jamun 9, Gujiya 13, Barfi 3 and Rasgulla 7 students. Describe a bar graph using one unit length for one student, and identify the greatest and least preferences. [5 marks]
  1. Give the graph the title “Sweet preferences of students”. Draw horizontal and vertical axes so that categories and their counts can be labelled clearly.
  2. Label the horizontal axis “Sweets” and place Jalebi, Gulab jamun, Gujiya, Barfi and Rasgulla along it with equal spacing.
  3. Label the vertical axis “Number of students”. Start at zero and mark equal intervals, using one unit length for one student.
  4. Draw equal-width bars with equal gaps. In the stated order, their heights must be 6, 9, 13, 3 and 7 units.
  5. Gujiya has the greatest preference, with 13 students, so its bar is tallest. Barfi has the least preference, with 3 students, so its bar is shortest.
Q7. Smriti scores 80, 50, 10, 100, 90, 0, 90 and 50 runs in Matches 1 to 8 respectively. A graph uses one unit length for 10 runs. Find the bar heights for Matches 4 and 8, identify the zero-score match, and name the matches scoring 90. [4 marks]
  1. Match 4 has 100 runs. Its bar is 100 ÷ 10 = 10 units high at the given scale.
  2. Match 8 has 50 runs. Its bar is 50 ÷ 10 = 5 units high.
  3. Match 6 records zero runs, so its bar has zero height. This still represents a recorded match.
  4. Matches 5 and 7 each record 90 runs. Their bars are equal in height because they represent equal scores on the same scale.
Q8. A fair coin has outcomes heads and tails; a fair die has faces numbered 1 to 6. Explain how to record repeated trials of each experiment, and whether equal chances require equal observed counts. [4 marks]
  1. Record heads or tails after every coin toss, retaining the order of the outcomes in a sequence.
  2. Record the upper-face number after every die throw. Tally the occurrences of each number from 1 to 6.
  3. Count the frequencies for each experiment and check that every recorded trial has been included in the counts.
  4. Equal chances do not require equal frequencies in a short experiment. Outcomes can repeat, and an individual future result remains uncertain.

Key takeaways

  • Choose data that answers the question being investigated, then organise the observations before drawing conclusions.
  • A frequency counts occurrences; a frequency table makes repeated values or categories easier to compare.
  • Calculate the mean by dividing the sum of all observations by their total count, including repetitions.
  • Find the median after sorting; with an even number of observations, average the two middle values.
  • The mode identifies a most frequent value, while the maximum identifies the greatest numerical value.
  • Bar graphs need clear labels, equal-width bars, equal gaps and a scale with consistent numerical intervals.
  • Keep both the sequence and frequencies of experimental outcomes because they show different features of the same data.
  • Equally likely outcomes can have unequal frequencies in a short experiment; chance does not guarantee alternation.

Test yourself

What is the frequency of a value?

It is the number of times that value occurs in the recorded collection of observations.

Why must repeated values be counted separately when calculating a mean?

Each repetition is a separate observation contributing to both the total and the number of observations.

Why is sorting essential before finding a median?

The median depends on the middle of the numerical order, not the order in which values were recorded.

What is the median of the sorted heights 155, 160, 164, 165, 169 and 173 cm?

The two middle heights are 164 and 165 cm. Their mean, and hence the median, is 164.5 cm.

What are the modes of 6, 2, 9, 5, 4, 6, 3 and 5?

The modes are 5 and 6. Each occurs twice, more frequently than the other values.

At ten runs per unit length, what does a five-unit bar represent?

It represents fifty runs, found by multiplying the height of five units by ten runs per unit.

Does zero students absent mean that a class has no students?

No. Zero students absent means that the class had full attendance on that day.

Must heads and tails alternate when a fair coin is tossed?

No. A random sequence may contain repeated outcomes; fairness does not require heads and tails to alternate.