Model G20 2027 at FLAME University, registrations now open

Statistics | ICSE Class 9 Maths Notes

26 min read

On this page

This note covers data collection, raw and arrayed data, discrete and continuous variables, tally marks, frequency tables, class limits and boundaries, continuous class intervals, class marks, mean and median of ungrouped data, and frequency polygons.

What are data, and how are they collected?

Data are collected information about a situation being studied. A variable is a characteristic whose value can change, such as height or the number of students in a class. An observation is an individual recorded value of a variable.

Data are usually collected in the context of a situation that we want to study. For example, to investigate the average height of a class, record the students' heights, organise the records and then interpret the information.

What distinguishes primary and secondary data?

Primary data are collected first-hand by the person conducting an enquiry. Asking students questions and recording their responses produces primary data for that investigator. Measuring their heights directly is another way of collecting information about the class.

Secondary data have already been collected and processed by somebody else. Published reports, books and other existing records can provide secondary data. The same information is primary for its original collector and secondary for somebody who later uses it.

Definition: Statistics involves collecting, organising, presenting, analysing and interpreting data. Organising makes individual records easier to compare; interpretation explains what the records show about the question being studied.

Why organise the records?

A long collection of unsorted observations can be difficult to interpret. Organisation brings related values together. A table displays information in rows and columns, while a graph presents it visually. Both help a reader recognise patterns in the collected information.

The purpose determines what information to collect and how to arrange it. Heights, marks and household sizes answer different questions. Keep the variable and its measurement unit clear, so that the values in a table or graph have an identifiable meaning.

How do raw, arrayed and grouped data differ?

Raw data are observations before they have been organised or classified. They are often large and cumbersome to handle. Their original order need not place similar values together, so locating the smallest, largest or middle value can require repeated searching.

Arrayed data are observations arranged in numerical order. Ascending order means smallest to largest; descending order means largest to smallest. Arranging the values changes their positions, but retains the observations, including each occurrence of a repeated value.

Worked example 1. Arrange the observations 5, 7, 6, 1, 8, 10, 12, 4, 3 in ascending order.

Answer: The array is 1, 3, 4, 5, 6, 7, 8, 10, 12. There are nine observations before and after arranging. The smallest value is 1 and the largest is 12.

What changes when data are grouped?

Grouped data collect observations into classes, which are stated ranges of values. The number of observations in each class is recorded. This gives a shorter summary than listing all the individual values, especially when the original collection is large.

For example, the observations 25, 25, 20, 22, 25 and 28 can be represented by the class 20 to 30 with six observations, using a convention that includes 20 and excludes 30. The class records a range and a count.

FormWhat the reader seesEffect on individual values
Raw dataThe collected observations before organisationIndividual recorded values remain available
Arrayed dataThe observations in numerical orderIndividual values and repetitions remain available
Grouped dataClasses with their numbers of observationsIndividual values within a class are no longer shown

This last difference is called loss of information. Knowing that six observations lie in a class does not identify all six original values. Retain the original list when an exact calculation from individual observations is required.

How can discrete and continuous variables be recognised?

A discrete variable takes separate permitted values, with gaps between them. A continuous variable can take any value within its possible range. The distinction concerns the values a variable can take, rather than whether the displayed records happen to contain decimal numbers.

Which quantities are counted?

The number of students in a class is discrete. It can be 25 or 26, but it cannot be 25.5. A change in the number of students happens in separate steps. The intermediate fractional values do not represent possible student counts.

Do not turn this example into a rule that every discrete variable must use whole numbers. A variable restricted to values such as 1/8, 1/16, 1/32 and 1/64 is also discrete: it cannot take every intermediate value between adjacent permitted fractions.

Which quantities are measured?

Height, weight, time and distance are examples of continuous variables. A height can pass through the intermediate values between 90 centimetres and 150 centimetres. A centimetre, written cm, is a unit used here for measuring height.

VariableTypeReason
Number of studentsDiscreteFractional student counts are not possible
HeightContinuousIntermediate measured values are possible
TimeContinuousIt can vary continuously within a range

Note: Grouping observations does not itself determine whether the variable is discrete or continuous. First consider the quantity being recorded, then consider how its values have been presented.

For example, heights may be displayed as individual observations or collected into classes. Both presentations concern the same continuous variable. Similarly, a table of counts remains a presentation of a discrete variable even when several counts are grouped together.

How do tally marks produce a frequency table?

The frequency of a value is the number of times it occurs. In calculations below, + means addition and = means equality. A frequency table places each value, or each class, beside its count. For grouped data, class frequency means the number of observations belonging to that class.

A tally mark is a counting stroke. Record one stroke for each observation. After four strokes, draw the fifth across them to complete a group of five. Counting completed groups and remaining strokes reduces the difficulty of counting a long list.

How is each observation counted?

  1. Write the values or class intervals that will form the rows of the table.
  2. Read the raw observations one at a time, retaining repeated values.
  3. Place one tally against the appropriate row for each observation.
  4. Count the tallies in every row and write the corresponding frequency.
  5. Add the frequencies and check that their sum matches the number of observations counted.

Worked example 2. Tabulate the observations 25, 25, 20, 22, 25, 28 using tallies.

Answer: The value 20 occurs once, 22 occurs once, 25 occurs three times and 28 occurs once. The total frequency is 1 + 1 + 3 + 1 = 6.

ValueTally marksFrequency
20/1
22/1
25///3
28/1
TotalSix strokes6

What must be checked?

The frequency is a count, not the value being counted. In this example, 25 is an observation value and 3 is its frequency. Writing 25 as its frequency would confuse the measurement with the number of occurrences.

For a grouped table, use the stated endpoint convention consistently. If a class includes its lower endpoint but excludes its upper endpoint, a value at the shared endpoint belongs to the next class. This prevents an observation from being counted in two rows.

What do class intervals, limits and class size mean?

A class interval is a range used to group observations. Its stated endpoints are its class limits. The smaller endpoint is the lower class limit, and the larger endpoint is the upper class limit. Both identify the class being discussed.

For the class 60 to 70, the lower limit is 60 and the upper limit is 70. The sign − means subtraction. Whether an observation equal to an endpoint is included depends on the convention used for the distribution. Read that convention before entering tally marks.

How do inclusive and exclusive intervals differ?

An inclusive interval includes observations equal to both stated limits. In an exclusive interval, one endpoint is excluded. In the examples here, continuous intervals use the convention that the lower endpoint is included and the upper endpoint is excluded.

Under this convention, 30 to 40 means 30 and above but below 40. The next class, 40 to 50, includes 40. The number 40 therefore has one place in the table, even though it is printed at the junction of two classes.

Worked example 3. Find the limits and class size of the continuous interval 60 to 70, with its upper endpoint excluded.

Answer: The lower limit is 60 and the upper limit is 70. The class size is 70 − 60 = 10. An observation equal to 70 belongs to the next interval under the stated convention.

How is class size measured?

Class size, also called class width, measures the span between the class boundaries. For continuous intervals such as 60 to 70, it is the upper endpoint minus the lower endpoint. A class boundary is an endpoint marking the actual division between adjoining continuous classes.

Discontinuous inclusive classes need adjustment before their continuous widths are measured. Do not automatically subtract the printed limits of an inclusive class and call that its continuous width. First establish its boundaries, then subtract the lower boundary from the upper boundary.

How are discontinuous intervals converted into continuous intervals?

Inclusive classes such as 800 to 899 and 900 to 999 leave a numerical gap between the upper limit of one class and the lower limit of the next. Converting them into continuous intervals makes adjacent classes meet at a common boundary.

The adjustment is half the gap between adjacent printed limits. Subtract it from each lower limit and add it to each upper limit. This places the new common boundary halfway between the two original limits on either side of the gap.

What is the conversion procedure?

  1. Subtract the upper limit of a class from the lower limit of the next class.
  2. Divide this gap by two to obtain the adjustment.
  3. Subtract the adjustment from every lower class limit.
  4. Add the adjustment to every upper class limit.
  5. Retain the original class frequencies and check that consecutive boundaries meet.

Worked example 4. Convert the income classes 800 to 899 and 900 to 999 into continuous intervals. Income is expressed in rupees, written ₹.

Answer: The gap is 900 − 899 = 1, and half the gap is 0.5. The intervals become 799.5 to 899.5 and 899.5 to 999.5. Both have class size 100.

The complete income distribution shows the adjustment without changing the number of employees in any class. Each frequency is a count of employees, while each interval gives their income range.

Printed income class (₹)Continuous income interval (₹)Number of employees
800 to 899799.5 to 899.550
900 to 999899.5 to 999.5100
1000 to 1099999.5 to 1099.5200
1100 to 11991099.5 to 1199.5150
1200 to 12991199.5 to 1299.540
1300 to 13991299.5 to 1399.510
TotalAll six intervals550

Note: The adjustment is 0.5 here because the gap is 1. Calculate half the actual gap instead of treating 0.5 as a universal correction. The conversion changes the boundaries, not the observed frequencies.

How are class marks calculated and used?

A class mark, or class midpoint, is the value halfway between the two endpoints of a class. It supplies a representative position for that class. On a frequency polygon, each class frequency is plotted above its class mark.

To calculate it, add the endpoints and divide their sum by two. Parentheses group the quantities to be calculated together. For an adjusted continuous class, use its lower and upper boundaries. Adding the same adjustment at one end and subtracting it at the other leaves their sum unchanged.

Result: The class midpoint

Class mark = (lower boundary + upper boundary) ÷ 2. The sign ÷ means division. For a class already written with continuous endpoints, the same calculation uses the stated lower and upper limits.

Worked example 5. Find the class mark of the continuous income interval 799.5 to 899.5.

Answer: Add the two boundaries: 799.5 + 899.5 = 1699. Divide by two: 1699 ÷ 2 = 849.5. The class mark is ₹849.5. The original limits also give (800 + 899) ÷ 2 = 849.5.

How does a midpoint differ from a width?

QuantityCalculation for 799.5 to 899.5Meaning
Class mark(799.5 + 899.5) ÷ 2 = 849.5The central position of the class
Class size899.5 − 799.5 = 100The width of the class

These quantities answer different questions. The midpoint locates a class along a numerical axis; the width measures how far the class extends. The frequency is different again: it counts the observations within the class.

A class mark is a representative value, not a list of the original observations. The class 20 to 30 has midpoint 25, but its observations can include 20, 22 and 28 as well as 25. Grouping does not make those original observations identical.

How is the mean of ungrouped data calculated?

The arithmetic mean, called the mean here, is the sum of all observations divided by their number. It is a measure of central tendency, meaning a single value used to summarise the centre or typical level of a collection of data.

Result: Mean of ungrouped observations

Mean = sum of observations ÷ number of observations. Use every observation, including repeated values. The number of observations is the number of entries in the list, not the number of distinct values that happen to appear.

  1. Read the complete list and identify the measurement unit.
  2. Count every observation, including repetitions.
  3. Add all the observation values to obtain their sum.
  4. Divide the sum by the number of observations and state the answer with its unit.

Worked example 6. Find the mean height of eight senior badminton trainees whose heights, in cm, are 165, 169, 164, 167, 170, 159, 164, 166.

Answer: The sum is 165 + 169 + 164 + 167 + 170 + 159 + 164 + 166 = 1324 cm. There are eight heights, so their mean is 1324 ÷ 8 = 165.5 cm.

How are all members of a collection included?

The height 164 appears twice, so it contributes twice to both the recorded collection and the calculation. A repeated height still belongs to another observation. Omitting its second occurrence changes the collection whose mean is being found.

Worked example 7. Find the mean height of all eleven trainees with heights, in cm, 165, 169, 164, 167, 170, 159, 164, 166, 146, 149, 153.

Answer: The eight senior heights total 1324 cm and the three junior heights total 448 cm. The complete total is 1772 cm. Dividing by eleven gives 1772 ÷ 11, approximately 161.09 cm to two decimal places.

Adding two subgroup means and dividing by two does not generally give the mean of the whole collection when the groups contain different numbers of observations. The direct method avoids this problem by adding all eleven original heights and dividing by eleven.

Keep enough precision during a calculation. Here the final decimal is rounded to two decimal places; the fraction 1772/11 retains the exact quotient. The slash in that fraction means division, just as the division sign does.

How is the median found for odd and even numbers of observations?

The median is the middle value of an ordered collection. It is a positional measure: arrange the observations before locating its position. For an even number of observations, the median is the arithmetic mean of the two middle values.

An odd count is not divisible by two; an even count is divisible by two. Let n denote the number of observations. A position counts where an item stands in the array; it is not necessarily the value written in that position. Keeping these two meanings separate is essential when applying a median rule.

Result: Median when n is odd

When n is odd, the median is the value in position (n + 1) ÷ 2 of the ordered array. An odd number of entries has a single central position, with equally many positions before and after it.

Worked example 8. Find the median of 5, 7, 6, 1, 8, 10, 12, 4, 3.

Answer: In ascending order the data are 1, 3, 4, 5, 6, 7, 8, 10, 12. Here n = 9, so the middle position is (9 + 1) ÷ 2 = 5. The fifth value is 6; therefore the median is 6.

Result: Median when n is even

When n is even, use the values in positions n ÷ 2 and (n ÷ 2) + 1. Add these two values and divide by two. Do not average their position numbers and report that result as the median.

Worked example 9. Find the median of these twenty marks: 25, 72, 28, 65, 29, 60, 30, 54, 32, 53, 33, 52, 35, 51, 42, 48, 45, 47, 46, 33.

Answer: The ascending array is 25, 28, 29, 30, 32, 33, 33, 35, 42, 45, 46, 47, 48, 51, 52, 53, 54, 60, 65, 72. The tenth and eleventh values are 45 and 46. Their mean, (45 + 46) ÷ 2, is 45.5 marks.

How does median differ from mean?

The mean uses the numerical size of every observation in a total. The median uses the central position or positions after sorting. The twenty-mark example also shows that a median need not be one of the original recorded values.

Repeated values must remain in the array. The two occurrences of 33 represent two observations, so both affect position counting. Removing one would change the number of entries and could change which positions are regarded as central.

For the nine-value example, four observations lie below 6 and four above it. In datasets containing repeated central values, some observations may equal the median. Do not insist that exactly half the observations must be strictly smaller in every dataset.

How is a frequency polygon drawn and interpreted?

A frequency polygon represents a grouped frequency distribution by joining plotted points with straight line segments. Each point pairs a class mark with its frequency. The horizontal line of reference is the x-axis; the vertical line is the y-axis.

An ordered pair writes the horizontal position first and the vertical position second. Here, (class mark, frequency) identifies a point. Use a consistent numerical scale on each axis, with class marks horizontally and frequencies vertically.

What steps produce the graph?

  1. Convert discontinuous classes to continuous intervals where necessary.
  2. Calculate each class mark from its two endpoints.
  3. Label the axes with the variable, its unit and the frequency being counted.
  4. Plot each class mark against its corresponding frequency.
  5. Join consecutive points with straight line segments.
  6. Complete the ends using zero-frequency classes immediately before and after the distribution.

The zero-frequency classes are adjoining classes with no observations, added to complete the polygon at the horizontal axis. For the equal-width income distribution, extend the sequence by one class of the same width at each end.

Worked example 10. Prepare the plotting points for income classes 800 to 899, 900 to 999, 1000 to 1099, 1100 to 1199, 1200 to 1299 and 1300 to 1399, with frequencies 50, 100, 200, 150, 40 and 10 respectively.

Answer: The class marks are 849.5, 949.5, 1049.5, 1149.5, 1249.5 and 1349.5. Pair them with the given frequencies. The adjoining zero-frequency class marks are 749.5 and 1449.5, each one class width, 100, beyond the nearest actual class mark.

Position along income axis (₹)FrequencyPoint to plot
749.50(749.5, 0)
849.550(849.5, 50)
949.5100(949.5, 100)
1049.5200(1049.5, 200)
1149.5150(1149.5, 150)
1249.540(1249.5, 40)
1349.510(1349.5, 10)
1449.50(1449.5, 0)

Draw and label

Income frequency polygon

Label the horizontal axis Income (₹) and the vertical axis Number of employees. Plot all eight points in the table, retaining equal horizontal spacing for successive class marks. Join them in order with straight segments, starting and ending on the horizontal axis.

What does a plotted point mean?

The point (1049.5, 200) represents 200 employees in the income class 1000 to 1099. It does not establish that each employee earns exactly ₹1049.5. The midpoint identifies the class; the vertical coordinate gives its frequency.

A histogram is a graph of adjoining rectangles with class intervals as their bases and areas proportional to frequency. For equal-width classes, joining the midpoints of the tops of these rectangles is another way to construct the frequency polygon.

What the figure shows

Daily-wage frequency polygon

Straight segments connect points across the tops of shaded adjacent rectangles. The horizontal axis is labelled Daily wage (in Rs), with continuous classes shown below it. The vertical axis is labelled No. of wage-earners.

See Fig. 4.6 in your NCERT textbook

The shape shows where frequencies rise or fall across the classes. When comparing distributions on the same axes, a frequency polygon is likely to be more useful because the horizontal and vertical lines of multiple histograms may coincide.

Glossary

  • Data — Collected information used to study a situation and answer a question about it.
  • Observation — One recorded value of the variable being studied in a collection of data.
  • Raw data — Observations in their original, unclassified form before numerical ordering or grouping.
  • Arrayed data — Individual observations arranged in ascending or descending numerical order, retaining repeated values.
  • Discrete variable — A variable taking separate permitted values, without taking every value between them.
  • Continuous variable — A variable capable of taking any value within its possible range.
  • Frequency — The number of occurrences of a value or observations belonging to a class.
  • Tally marks — Counting strokes recording observations, grouped in fives to make counting easier.
  • Class limits — The stated lower and upper endpoints used to describe a class interval.
  • Class boundaries — Endpoints defining the actual divisions between adjacent continuous class intervals.
  • Class size — The width obtained by subtracting a class's lower boundary from its upper boundary.
  • Class mark — The midpoint found by adding a class's two endpoints and dividing by two.
  • Arithmetic mean — The sum of every observation divided by the total number of observations.
  • Median — The central value of ordered data, averaging the two central values for an even count.
  • Frequency polygon — A graph joining class-mark and frequency points by successive straight line segments.

Common errors and misconceptions

  • Misconception: Arranging data means writing each different value once. Correct: Retain every occurrence of a repeated value; each represents an observation and contributes to position counting.
  • Misconception: Every discrete variable must take whole-number values. Correct: Separate permitted fractional values can also form a discrete variable when intermediate values are not permitted.
  • Misconception: A shared endpoint belongs to both neighbouring classes. Correct: Follow the inclusion convention consistently so each observation is counted in one class.
  • Misconception: Every discontinuous interval needs an adjustment of 0.5. Correct: Calculate half the actual gap; 0.5 applies to the illustrated gap of 1.
  • Misconception: The median position is the median value. Correct: Locate the relevant position in the ordered array and read the observation there.
  • Misconception: For an even count, either middle value is an acceptable median. Correct: Add the two central values and divide by two.
  • Misconception: Polygon points are plotted above class boundaries. Correct: Plot each frequency above its class mark and join consecutive points with straight segments.

Exam-style questions with model answers

Q1. Distinguish between discrete and continuous variables, using the number of students in a class and student height as examples. [2 marks]
  1. The number of students is discrete because it takes separate whole-number values and cannot include a fractional student.
  2. Height is continuous because it can take intermediate measured values within its possible range, including fractional values.
Q2. For the data 25, 25, 20, 22, 25, 28, state the frequencies of the distinct values, explain how tallies record them, and check the total. [3 marks]
  1. The distinct values are 20, 22, 25 and 28. Their respective frequencies are 1, 1, 3 and 1 because frequency counts how many times each value occurs.
  2. Enter one tally for each occurrence: one beside 20, one beside 22, three beside 25 and one beside 28. The repeated value 25 must be counted three times.
  3. The frequency total is 1 + 1 + 3 + 1 = 6, matching the six observations supplied.
Q3. Convert the inclusive income classes ₹800 to ₹899 and ₹900 to ₹999, containing 50 and 100 employees respectively, into continuous intervals. State the adjustment, new intervals, class sizes and frequencies. [4 marks]
  1. The gap between the classes is 900 − 899 = 1. Half of this gap is 0.5, which is the required adjustment.
  2. Subtract 0.5 from each lower limit and add 0.5 to each upper limit. The continuous intervals are ₹799.5 to ₹899.5 and ₹899.5 to ₹999.5.
  3. The first width is 899.5 − 799.5 = 100; the second width is 999.5 − 899.5 = 100. Both classes have size ₹100.
  4. The frequencies remain 50 and 100 employees respectively. Changing the boundaries does not change the observations counted in either class.
Q4. Calculate the mean of these eight heights in cm: 165, 169, 164, 167, 170, 159, 164, 166. Show the count, total and division. [3 marks]
  1. There are eight observations. The height 164 cm occurs twice, and both occurrences must be included because the list contains two observations with that value.
  2. Add every height: 165 + 169 + 164 + 167 + 170 + 159 + 164 + 166 = 1324 cm.
  3. Divide the total by the observation count. The arithmetic mean is 1324 ÷ 8 = 165.5 cm, expressed in the same unit as the given heights.
Q5. Find the median of 5, 7, 6, 1, 8, 10, 12, 4, 3, showing the ordered array and the middle position. [3 marks]
  1. Arrange all the observations in ascending order: 1, 3, 4, 5, 6, 7, 8, 10, 12. Sorting is required before locating a positional average.
  2. There are nine observations. The middle position is (9 + 1) ÷ 2 = 5, so the fifth observation in the array is needed.
  3. The fifth observation is 6. Therefore the median is 6, with four entries before it and four after it in this array.
Q6. Find the median of these twenty marks: 25, 72, 28, 65, 29, 60, 30, 54, 32, 53, 33, 52, 35, 51, 42, 48, 45, 47, 46, 33. Show the ordered data and both middle values. [4 marks]
  1. The ascending array is 25, 28, 29, 30, 32, 33, 33, 35, 42, 45, 46, 47, 48, 51, 52, 53, 54, 60, 65, 72.
  2. The count is twenty, an even number. The required central positions are 20 ÷ 2 = 10 and 10 + 1 = 11.
  3. The tenth value is 45 marks and the eleventh value is 46 marks. Both values are needed for an even count.
  4. The median is their arithmetic mean: (45 + 46) ÷ 2 = 45.5 marks. It need not equal one of the original marks.
Q7. Give complete plotting instructions for a frequency polygon of income classes ₹800 to ₹899, ₹900 to ₹999, ₹1000 to ₹1099, ₹1100 to ₹1199, ₹1200 to ₹1299 and ₹1300 to ₹1399, with respective frequencies 50, 100, 200, 150, 40 and 10 employees. Include continuous boundaries, points, end completion and the meaning of the highest point. [6 marks]
  1. Half the gap between adjacent classes is 0.5. The continuous intervals are 799.5 to 899.5, 899.5 to 999.5, 999.5 to 1099.5, 1099.5 to 1199.5, 1199.5 to 1299.5 and 1299.5 to 1399.5.
  2. The class marks, obtained by averaging each pair of boundaries, are 849.5, 949.5, 1049.5, 1149.5, 1249.5 and 1349.5.
  3. Label the horizontal axis Income (₹) and the vertical axis Number of employees. Use consistent scales and plot (849.5, 50), (949.5, 100), (1049.5, 200), (1149.5, 150), (1249.5, 40) and (1349.5, 10).
  4. Add zero-frequency points (749.5, 0) and (1449.5, 0). Their class marks continue the equal-width sequence one class before and after the actual distribution.
  5. Join all eight plotted points in increasing order of class mark using straight line segments. The polygon starts and finishes on the horizontal axis.
  6. The highest point is (1049.5, 200). It represents 200 employees in the original ₹1000 to ₹1099 income class, rather than establishing an identical income for those employees.
Q8. The observations 25, 25, 20, 22, 25, 28 are grouped into the class 20 to 30, including 20 but excluding 30. Explain what this grouped entry preserves, what it loses, and how its frequency, class size and class mark differ. [5 marks]
  1. The class frequency is 6 because all six supplied observations are at least 20 and below 30. It preserves the count of observations in that range.
  2. The class size is 30 − 20 = 10. This measures the width of the continuous interval, rather than counting how many observations it contains.
  3. The class mark is (20 + 30) ÷ 2 = 25. It is the midpoint used to represent the class, including when plotting a frequency polygon.
  4. The grouped entry no longer displays the individual observations 25, 25, 20, 22, 25 and 28. This is the loss of information caused by replacing the list with a class count.
  5. A class mark of 25 does not mean that all six observations equal 25. Frequency, width and midpoint describe different features of the same grouped entry.

Key takeaways

  • Organise raw observations into an ordered array when individual positions matter, keeping every repeated value in the collection.
  • Identify discrete and continuous variables from their possible values, rather than simply looking for decimals in the recorded data.
  • Tally each observation once in its correct row, then check that all frequencies add to the observation count.
  • Read the endpoint convention carefully so a value at a shared class boundary is assigned consistently.
  • Convert discontinuous intervals by using half the gap, adjusting both limits while retaining every original class frequency.
  • Calculate class size by subtracting boundaries, and calculate class mark by averaging the two endpoints.
  • Find mean by dividing the complete total by the observation count; find median from the central ordered positions.
  • Plot class marks against frequencies, join successive points with straight segments, and complete the polygon with adjoining zero-frequency classes.

Test yourself

Why does an ordered array retain repeated observations?

Each occurrence is a separate observation, so removing repetitions would change the collection and its position count.

A value of 40 is recorded. With lower endpoints included and upper endpoints excluded, does it belong to 30 to 40 or 40 to 50?

It belongs to 40 to 50 because that class includes its lower endpoint, while 30 to 40 excludes 40.

What adjustment makes 800 to 899 and 900 to 999 continuous?

The gap is 1, so subtract 0.5 from lower limits and add 0.5 to upper limits.

For continuous boundaries 799.5 and 899.5, what are the width and midpoint?

The width is 100, found by subtraction; the midpoint is 849.5, found by averaging the boundaries.

Eight heights total 1324 cm. What is their arithmetic mean?

Divide the total by eight observations: 1324 ÷ 8 = 165.5 cm.

What is the median of the ordered array 1, 3, 4, 5, 6, 7, 8, 10, 12?

There are nine values, so the fifth value is central. The median is therefore 6.

The middle values of twenty ordered marks are 45 and 46. What is the median?

Average the two middle values: (45 + 46) ÷ 2 = 45.5 marks.

What do the two coordinates of a frequency polygon point represent?

The first coordinate is the class mark; the second is the frequency of that class.