Model G20 2027 at FLAME University, registrations now open

Probability | ISC Class 12 Maths Notes

26 min read

On this page

This note covers events and probability laws, conditional probability, multiplication of probabilities, independent and dependent events, total probability, Bayes’ theorem, random variables, probability distributions and the mean of a random variable.

How do events and probability laws describe a random experiment?

A random experiment has possible results, called outcomes, but its particular result is uncertain before it is performed. The sample space, denoted by S, is the set of all possible outcomes. An event is a subset of S. An elementary outcome is one individual possible result.

Let A and B denote events. The notation P(A) means the probability that A occurs. An event occurs when the observed outcome belongs to it. Probabilities lie between 0 and 1, inclusive. The whole sample space is a sure event, with P(S) = 1.

What do “and”, “or” and “not” mean?

The intersection A ∩ B contains outcomes belonging to both events. The union A ∪ B contains outcomes belonging to at least one event, including their common outcomes. The complement A′ contains outcomes in S which do not belong to A.

The symbol ∅ denotes the empty set, containing no outcomes. It represents an impossible event and has probability zero. Mutually exclusive, or disjoint, events have no common outcome. Exhaustive events together cover the whole sample space.

Event expressionMeaningProbability rule
A′A does not occurP(A′) = 1 − P(A)
A ∪ BAt least one of A and B occursP(A ∪ B) = P(A) + P(B) − P(A ∩ B)
A ∪ B, when A ∩ B = ∅One of two disjoint events occursP(A ∪ B) = P(A) + P(B)

The addition theorem subtracts the intersection because adding the two individual probabilities counts that common part twice. The subtraction is unnecessary when the events are disjoint. “At least one” includes the possibility that both events occur; it does not mean exactly one.

For a finite sample space with equally likely outcomes, each outcome has the same probability. An event’s probability is then the number of favourable outcomes divided by the total number of outcomes. This counting rule needs the equal-likelihood assumption.

For coin notation, H denotes a head and T a tail; a string such as HT records the outcomes in order. A fair coin has probability 1/2 for each face. A fair die has six equally likely faces numbered 1 to 6.

What changes when a probability is conditional on another event?

Conditional probability measures the probability of an event when another event is known to have occurred. Let E and F denote events in the same sample space. The notation P(E|F) reads “the probability of E given F”; the vertical bar means “given”.

Definition: P(E|F) = P(E ∩ F)/P(F), provided P(F) ≠ 0. The condition F supplies the denominator, while the numerator measures outcomes satisfying both the event and the condition.

Knowing F restricts the possible outcomes to those in F. Among these, the outcomes favourable to E are precisely those in E ∩ F. Thus the denominator changes from the whole experiment to the event known to have occurred.

How does restricting the sample space help?

For three fair, independent coin tosses, S = {HHH, HHT, HTH, THH, HTT, THT, TTH, TTT}. Independence here means that one toss does not change another toss’s probabilities. Each listed outcome has probability 1/8.

Let E mean at least two heads, and F mean a tail on the first toss. Then F = {THH, THT, TTH, TTT}, and E ∩ F = {THH}. Consequently P(E|F) = 1/4, whereas P(E) = 1/2.

Worked example 1. Ten cards numbered 1 to 10 are mixed thoroughly, and one is drawn with every card equally likely. Given that its number exceeds 3, find the probability that it is even.

Answer: Let A mean an even number and B mean a number greater than 3. The possible numbers under B are {4, 5, 6, 7, 8, 9, 10}. The favourable numbers are {4, 6, 8, 10}.

Therefore P(A|B) = (4/10)/(7/10) = 4/7. The denominator counts the seven cards permitted by the condition, rather than all ten cards.

Write the condition in words before substituting. P(E|F) and P(F|E) have the same intersection in their numerators but generally different denominators. The information supplied by the question determines which of these two probabilities is required.

Which properties remain true for conditional probability?

With the condition held fixed, conditional probability obeys the familiar probability rules. Throughout this section F is the conditioning event, and P(F) must be positive. This requirement makes division by P(F) meaningful.

Property: Certainty within the given condition

P(S|F) = P(F|F) = 1. Since F is part of S, the intersection S ∩ F equals F. Also F ∩ F equals F. In both cases the conditional probability is P(F)/P(F).

Property: Addition and complements under a fixed condition

For any events A and B, P((A ∪ B)|F) = P(A|F) + P(B|F) − P((A ∩ B)|F). If A and B are disjoint, their intersection contributes zero. The same F must condition every term in this equation.

The conditional complement rule is P(E′|F) = 1 − P(E|F). Inside F, either E occurs or E does not occur. These possibilities are disjoint and cover F, so their conditional probabilities add to one.

Worked example 2. A fair die is thrown twice independently. Given that the sum is 6, find the probability that 4 appeared at least once.

Answer: Let F mean sum 6 and E mean at least one 4. The ordered pairs in F are {(1,5), (2,4), (3,3), (4,2), (5,1)}. The first entry records the first throw.

Only (2,4) and (4,2) also belong to E. Therefore P(F) = 5/36, P(E ∩ F) = 2/36 and P(E|F) = (2/36)/(5/36) = 2/5.

Order matters when throws are recorded successively: (2,4) and (4,2) are distinct outcomes. Their sums are equal, but the experiment records different results on its first and second throws. Keep both when listing the conditional sample space.

Note: P(E|F) is undefined by this formula when P(F) = 0. A zero denominator does not give a conditional probability of zero, even when the intersection also has probability zero.

How should unequal outcome probabilities be handled?

The conditional probability formula also applies when outcomes are not equally likely. In that case, calculate the probabilities of the relevant events by adding the probabilities of their outcomes. A ratio of outcome counts is not generally valid.

Consider tossing a fair coin. After a head, toss the coin again; after a tail, throw a fair die. Assume each later trial has the stated fair probabilities, whatever the first result. The eight recorded outcomes are HH, HT, T1, T2, T3, T4, T5 and T6.

Why are these eight outcomes not equiprobable?

“Equiprobable” means equally likely. HH and HT each have probability 1/4. Each outcome beginning with T followed by a die number has probability 1/12. There are different numbers of possible continuations after the first head and the first tail.

What the figure shows

Coin and die tree

The first split is labelled Head (H) and Tail (T). The head branch splits into (H,H) and (H,T). The tail branch splits into (T,1), (T,2), (T,3), (T,4), (T,5) and (T,6).

See Fig. 13.1 in your NCERT textbook

What the figure shows

Probabilities on the tree

The first head and tail branches each carry 1/2. The two final coin outcomes are labelled 1/4 each, and the six final die outcomes are labelled 1/12 each.

See Fig. 13.2 in your NCERT textbook

Worked example 3. In the fair coin and fair die experiment just described, find the probability that the die shows a number greater than 4, given that at least one tail occurs.

Answer: Let E mean a die result greater than 4 and F mean at least one tail. Then E ∩ F = {T5, T6}, with probability 1/12 + 1/12 = 1/6.

F includes HT and all six outcomes beginning with T. Hence P(F) = 1/4 + 6 × (1/12) = 3/4. Therefore P(E|F) = (1/6)/(3/4) = 2/9.

The two favourable outcomes and seven outcomes in F cannot be treated as a favourable-count ratio. Their probabilities differ. The event probability, rather than the number of written entries in its set, is what the conditional formula requires.

How does the multiplication theorem handle successive events?

The multiplication theorem finds the probability that events occur together. It follows directly by multiplying the conditional probability formula by its denominator. It is particularly useful for successive selections, where an earlier selection can alter the probabilities of later ones.

Theorem: Multiplication of probabilities

P(E ∩ F) = P(E)P(F|E), provided P(E) ≠ 0. Equivalently, P(E ∩ F) = P(F)P(E|F), provided P(F) ≠ 0. The two forms describe the same intersection, although they use different conditioning events.

Without replacement means that the selected object is not returned before the next draw. The contents then change. With replacement means returning it before the next draw; probabilities must still be interpreted using the random-selection conditions specified in the question.

Worked example 4. An urn contains 10 black and 5 white balls. Two balls are drawn randomly, one after the other without replacement. Find the probability that both are black.

Answer: Let E mean that the first ball is black and F that the second is black. P(E) = 10/15. Given E, there are 9 black balls among 14 remaining balls, so P(F|E) = 9/14.

Thus P(E ∩ F) = (10/15) × (9/14) = 3/7. Both the favourable count and the total count change after the first black ball is removed.

How does the rule extend to three events?

For a third event G, P(E ∩ F ∩ G) = P(E)P(F|E)P(G|(E ∩ F)), where the conditioning events have positive probabilities. The final factor conditions on both earlier events together, rather than on just the immediately preceding event.

Worked example 5. Three cards are drawn successively without replacement from a well-shuffled pack of 52 cards containing four kings and four aces, with no card both a king and an ace. Find the probability of two kings followed by an ace.

Answer: The first king has probability 4/52. Given that king, the second king has probability 3/51. Given both kings, four aces remain among 50 cards, so the final conditional probability is 4/50.

The required probability is (4/52) × (3/51) × (4/50) = 2/5525. This calculation follows the specified order; it does not include other placements of the ace.

How can independent events be distinguished from mutually exclusive events?

Events E and F are independent when P(E ∩ F) = P(E)P(F). For positive conditioning probabilities, this is equivalent to P(E|F) = P(E) and P(F|E) = P(F). Knowing that one occurred does not alter the probability of the other.

Events are dependent when the product equality fails. The test compares the actual intersection probability with the product of the separate probabilities. Events need not arise from different physical experiments to be independent.

What is the difference between independence and exclusion?

FeatureIndependent eventsMutually exclusive events
Defining testP(E ∩ F) = P(E)P(F)E ∩ F = ∅
Common outcomesMay have common outcomesHave no common outcomes
When both probabilities are positiveIntersection probability is positiveIntersection probability is zero

Two mutually exclusive events with nonzero probabilities cannot be independent. Their intersection probability is zero, while the product of their positive probabilities is greater than zero. The qualification “with nonzero probabilities” matters in this comparison.

Worked example 6. A fair die is thrown. Let E mean a multiple of 3 and F mean an even number. Determine whether E and F are independent.

Answer: E = {3,6}, F = {2,4,6} and E ∩ F = {6}. Therefore P(E) = 1/3, P(F) = 1/2 and P(E ∩ F) = 1/6.

Since (1/3) × (1/2) = 1/6, the events are independent. They are not mutually exclusive because the outcome 6 belongs to both.

Result: Independence survives taking complements

If E and F are independent, so are E and F′, E′ and F, and E′ and F′. For instance, P(E ∩ F′) = P(E) − P(E ∩ F) = P(E)[1 − P(F)] = P(E)P(F′).

For independent A and B, the probability of at least one is 1 − P(A′)P(B′). For three events A, B and C, mutual independence requires all three pairwise product equalities and P(A ∩ B ∩ C) = P(A)P(B)P(C). Pairwise checks alone are insufficient.

How does the theorem of total probability combine different cases?

A partition divides the sample space into events which are pairwise disjoint and exhaustive. Pairwise disjoint means that every two different events have no common outcome. In the total-probability formula, each partition event must also have nonzero probability. Exactly one of these alternative cases occurs in any trial.

Write the partition events as E₁, E₂, …, Eₙ. Here n is the number of events, and a subscript such as i identifies one of them, with i running from 1 to n. Let A be the event whose overall probability is required.

Theorem: Total probability

P(A) = P(E₁)P(A|E₁) + P(E₂)P(A|E₂) + … + P(Eₙ)P(A|Eₙ). Each product describes A occurring through one particular case. Adding these disjoint contributions gives the probability of A across all cases.

  1. Since the partition covers S, every outcome belonging to A belongs to one of the partition events.
  2. Therefore A is the union of A ∩ E₁, A ∩ E₂, and so on through A ∩ Eₙ.
  3. These intersections are disjoint, so their probabilities can be added without subtracting overlaps.
  4. Apply the multiplication theorem to each intersection, obtaining the weighted sum of conditional probabilities above.

What the figure shows

A partition of the sample space

A rectangle labelled S is divided into regions labelled E₁, E₂, E₃, E₄ and further regions indicated by dots. A shaded oval labelled A crosses the partition regions.

See Fig. 13.4 in your NCERT textbook

Worked example 7. For a construction job, the probability of a strike is 0.65. The probability of completion on time is 0.32 if there is a strike and 0.80 if there is no strike. Find the overall probability of timely completion.

Answer: Let A mean completion on time and B mean a strike. Then P(B′) = 1 − 0.65 = 0.35. B and B′ are disjoint and exhaustive, so they provide the required two cases.

P(A) = P(B)P(A|B) + P(B′)P(A|B′) = 0.65 × 0.32 + 0.35 × 0.80 = 0.208 + 0.28 = 0.488.

The weights are the probabilities of the cases themselves. An unweighted average of the conditional probabilities would ignore how likely each case is. Check that all cases are included and that no outcome has been counted in two cases.

How does Bayes’ theorem revise the probability of a possible cause?

Bayes’ theorem calculates the probability of a particular case after an associated event is observed. The possible cases E₁, E₂, …, Eₙ are called hypotheses. They form a partition, with positive probabilities, while the observed event A must also have positive probability.

P(Eᵢ) is the prior probability of hypothesis Eᵢ, before using the information A. P(Eᵢ|A) is its posterior probability, after using that information. The probability P(A|Eᵢ) describes the observation under that particular hypothesis.

Theorem: Bayes’ formula

P(Eᵢ|A) = P(Eᵢ)P(A|Eᵢ)/P(A). Calculate the denominator P(A) by total probability, adding P(E₁)P(A|E₁) through P(Eₙ)P(A|Eₙ). The numerator is one contribution to that same sum.

The proof combines two previous results: conditional probability gives P(Eᵢ|A) = P(Eᵢ ∩ A)/P(A), and multiplication gives P(Eᵢ ∩ A) = P(Eᵢ)P(A|Eᵢ). Thus the posterior measures the selected case’s share of the total probability of the observation.

Worked example 8. Bag I contains 3 red and 4 black balls; Bag II contains 5 red and 6 black balls. Choose either bag with probability 1/2, then choose a ball uniformly from it. Given that the ball is red, find the probability of Bag II.

Answer: Let E₁ and E₂ mean selection of Bags I and II, respectively, and A mean red. Then P(A|E₁) = 3/7 and P(A|E₂) = 5/11.

P(E₂|A) = [(1/2) × (5/11)]/[(1/2) × (3/7) + (1/2) × (5/11)] = 35/68. The red fraction within Bag II alone does not answer the reverse question.

How are unequal prior probabilities handled?

Worked example 9. Machines A, B and C produce 25%, 35% and 40% of a factory’s bolts respectively. Their defective percentages are 5%, 4% and 2% respectively. A randomly chosen bolt is defective. Find the probability it came from machine B.

Answer: Let D mean defective. The joint contributions from A, B and C are 0.25 × 0.05 = 0.0125, 0.35 × 0.04 = 0.0140 and 0.40 × 0.02 = 0.0080.

Therefore P(D) = 0.0345, and P(from B|D) = 0.0140/0.0345 = 28/69. Each defective proportion is weighted by its machine’s share of total production.

In this example A, B and C label machines, while “from B” labels an event. Keep physical labels and event definitions distinct. Before calculating, identify what is observed and which earlier case the question asks you to infer.

What is a random variable and how is its probability distribution formed?

A random variable is a real-valued function whose domain is the sample space of a random experiment. A function assigns exactly one value to each input; its domain is the set of permitted inputs. Thus each outcome receives one real number.

Let X denote the number of heads in two successive coin tosses. With S = {HH, HT, TH, TT}, the values are X(HH) = 2, X(HT) = 1, X(TH) = 1 and X(TT) = 0.

How do outcomes differ from values of X?

An outcome records the experiment’s result; a value of X records the chosen numerical feature. HT and TH are distinct outcomes, but both give X = 1. To find the probability of a value, combine all outcomes producing it.

A discrete random variable has a finite or countable set of possible values. Countable values can be listed in a sequence, even if the list does not end. Its probability distribution assigns probabilities to those values. For a possible value x, define the probability mass function p(x) = P(X = x).

For two fair, independent tosses, the four outcomes above each have probability 1/4. The resulting distribution is:

Value xOutcomes giving X = xp(x)
0TT1/4
1HT, TH1/2
2HH1/4

Every assigned probability must be nonnegative, and all the probabilities must add to 1. These conditions express that exactly one possible value occurs. Here 1/4 + 1/2 + 1/4 = 1. Equal probabilities for outcomes do not imply equal probabilities for values.

Different random variables can describe the same experiment. If Y means the number of heads minus the number of tails, then Y(HH) = 2, Y(HT) = Y(TH) = 0 and Y(TT) = −2. Negative values are allowed; negative probabilities are not.

What does a cumulative distribution function mean?

The cumulative distribution function is Fₓ(t) = P(X ≤ t), where t is a real-number threshold. For a discrete distribution, add p(x) over values x at or below t. This differs from p(x), which gives the probability at one particular value.

How is the mean of a random variable calculated and interpreted?

The mean, or expected value, of a discrete random variable is its probability-weighted average. It combines the possible values with how likely they are. Write μ, the Greek letter mu, or E(X), for this mean; E(X) here denotes expectation, not an event.

Result: Mean of a finite probability distribution

If X takes values x₁, x₂, …, xₘ with probabilities p₁, p₂, …, pₘ, then μ = E(X) = x₁p₁ + x₂p₂ + … + xₘpₘ. Here m is the number of possible values, and pᵢ = P(X = xᵢ) for each index i.

Each product is one value’s contribution to the weighted average. Using the ordinary arithmetic mean of the listed values ignores their probabilities. The probability weights, rather than the number of distinct values, determine the correct average.

  1. Identify the numerical quantity represented by X and list all its possible values.
  2. Find the probability attached to each value, combining different outcomes when they give the same value.
  3. Check that the probabilities are nonnegative and have total one.
  4. Multiply each value by its probability and add all the products to obtain E(X).

Worked example 10. Two fair coins are tossed independently. Let X be the number of heads. Find its mean using P(X = 0) = 1/4, P(X = 1) = 1/2 and P(X = 2) = 1/4.

Answer: E(X) = 0 × (1/4) + 1 × (1/2) + 2 × (1/4) = 1. The expected number of heads is one. Both the zero-head and two-head outcomes remain possible.

Expected value does not predict a guaranteed result of the next trial. It is a summary of the whole distribution, and it need not be a possible value of the variable. Keep the distinction between a realised value and a probability-weighted average.

A random variable, its distribution and its mean answer different questions. The variable specifies what is measured; the distribution gives the probability of each possible value; the mean summarises that distribution by weighting values with their probabilities.

Glossary

  • Sample space — The set containing all possible outcomes of the random experiment being considered.
  • Event — A subset of the sample space, occurring when the observed outcome belongs to that subset.
  • Complement — The event containing all outcomes in the sample space that do not belong to the given event.
  • Conditional probability — The probability of one event given that another event of positive probability has occurred.
  • Independent events — Events whose intersection probability equals the product of their individual probabilities.
  • Dependent events — Events whose intersection probability differs from the product of their individual probabilities.
  • Mutually exclusive events — Events with no common outcome, so they cannot occur together in the same trial.
  • Partition — Disjoint events covering the sample space, taken with positive probabilities when applying total probability.
  • Prior probability — The probability assigned to a hypothesis before using the information supplied by an observed event.
  • Posterior probability — The conditional probability of a hypothesis after the specified observation is taken into account.
  • Random variable — A real-valued function assigning one numerical value to each outcome in a sample space.
  • Probability distribution — The assignment of probabilities to the possible values of a random variable.
  • Expected value — The mean obtained by multiplying each possible value by its probability and adding the products.

Common errors and misconceptions

  • Misconception: P(E|F) and P(F|E) mean the same thing. Correct: They condition on different events and generally use different denominators.
  • Misconception: A zero-probability condition gives conditional probability zero. Correct: The elementary conditional formula is undefined when its denominator is zero.
  • Misconception: Every listed outcome can be counted with equal weight. Correct: Counting favourable outcomes works directly only when the relevant elementary outcomes are equally likely.
  • Misconception: “Independent” means the events cannot occur together. Correct: That describes mutual exclusion; independent events of positive probability have a positive intersection probability.
  • Misconception: Successive draws without replacement keep the same probabilities. Correct: Recalculate the contents after the earlier specified selections before using the next conditional probability.
  • Misconception: Bayes’ denominator contains only the desired cause. Correct: It includes the observation’s probability through every case in the partition.
  • Misconception: Distinct random-variable values are equally likely whenever outcomes are equally likely. Correct: Several outcomes may produce the same value, so their probabilities must be combined.
  • Misconception: Expected value is a guaranteed outcome. Correct: It is a probability-weighted average of possible values, rather than a prediction of a particular trial.

Exam-style questions with model answers

Q1. Events A and B satisfy P(A) = 7/13, P(B) = 9/13 and P(A ∩ B) = 4/13. Find P(A|B), stating the condition needed for the formula. [2 marks]
  1. The condition is P(B) ≠ 0; here P(B) = 9/13 is positive, so conditioning on B is permitted.
  2. P(A|B) = P(A ∩ B)/P(B) = (4/13)/(9/13) = 4/9.
Q2. A fair six-sided die, numbered 1 to 6, is thrown twice independently. Given that the sum is 6, find the probability that 4 appears at least once. [3 marks]
  1. Let F mean sum 6 and E mean at least one 4. The possible ordered pairs in F are (1,5), (2,4), (3,3), (4,2) and (5,1), so P(F) = 5/36.
  2. Of these pairs, (2,4) and (4,2) contain a 4. Thus P(E ∩ F) = 2/36.
  3. Apply conditional probability: P(E|F) = (2/36)/(5/36) = 2/5. The given sum restricts the calculation to five equally likely ordered pairs.
Q3. An urn contains 10 black and 5 white balls. Two balls are selected uniformly at random, successively and without replacement. Find the probability that both are black. [3 marks]
  1. Let E mean a black ball on the first draw and F mean a black ball on the second. There are 15 balls initially, so P(E) = 10/15.
  2. After a first black ball, 9 black balls remain among 14 balls. Therefore P(F|E) = 9/14.
  3. By the multiplication theorem, P(E ∩ F) = (10/15) × (9/14) = 3/7. The second factor uses the changed contents because the first ball is not returned.
Q4. A fair six-sided die numbered 1 to 6 is rolled. Let E mean a multiple of 3 and F mean an even number. Test independence and state whether the events are mutually exclusive. [4 marks]
  1. The events are E = {3,6} and F = {2,4,6}; hence P(E) = 2/6 = 1/3 and P(F) = 3/6 = 1/2.
  2. The common outcome is 6, so E ∩ F = {6} and P(E ∩ F) = 1/6.
  3. P(E)P(F) = (1/3) × (1/2) = 1/6, equal to the intersection probability. Therefore E and F are independent.
  4. They are not mutually exclusive, because the outcome 6 satisfies both event descriptions in the same roll.
Q5. For a construction job, the probability of a strike is 0.65. Timely completion has probability 0.32 given a strike and 0.80 given no strike. Find the overall probability of timely completion. [4 marks]
  1. Let A mean timely completion and B mean a strike. The two cases B and B′ are mutually exclusive and exhaustive.
  2. The no-strike probability is P(B′) = 1 − 0.65 = 0.35. Both cases have positive probabilities, so total probability applies.
  3. P(A) = P(B)P(A|B) + P(B′)P(A|B′) = 0.65 × 0.32 + 0.35 × 0.80.
  4. The two contributions are 0.208 and 0.28. Adding them gives P(A) = 0.488, the probability of completion on time across both possible cases.
Q6. Machines A, B and C supply 25%, 35% and 40% respectively of all bolts in a factory. Of each machine’s output, 5%, 4% and 2% respectively are defective. One bolt is chosen uniformly from all bolts and found defective. Find the probability it came from B. [5 marks]
  1. Let D denote a defective bolt, and E₁, E₂ and E₃ denote origin from machines A, B and C respectively. These origins are mutually exclusive and exhaustive.
  2. The prior probabilities are P(E₁) = 0.25, P(E₂) = 0.35 and P(E₃) = 0.40. The corresponding conditional defective probabilities are 0.05, 0.04 and 0.02.
  3. The joint probabilities of each origin and a defect are 0.0125, 0.0140 and 0.0080 respectively, found by multiplying each prior by its conditional probability.
  4. Total probability gives P(D) = 0.0125 + 0.0140 + 0.0080 = 0.0345. This includes defective bolts from all three origins.
  5. Bayes’ theorem gives P(E₂|D) = P(E₂)P(D|E₂)/P(D) = 0.0140/0.0345 = 28/69. This is the required probability of machine B after observing the defect.
Q7. Two fair coins are tossed independently. Let X be the number of heads. Construct its probability distribution, verify the total probability, and find E(X), its mean. [5 marks]
  1. The sample space is {HH, HT, TH, TT}, where H means head and T means tail. Fairness and independence make each ordered outcome have probability 1/4.
  2. The possible values of X are 0, 1 and 2. TT gives zero heads, HT and TH each give one head, and HH gives two heads.
  3. Combine probabilities of outcomes with the same value: P(X = 0) = 1/4, P(X = 1) = 1/2 and P(X = 2) = 1/4.
  4. All probabilities are nonnegative, and their sum is 1/4 + 1/2 + 1/4 = 1, verifying the required distribution conditions.
  5. The probability-weighted mean is E(X) = 0 × (1/4) + 1 × (1/2) + 2 × (1/4) = 1 head. This mean does not guarantee one head in every trial.
Q8. Bag I has 3 red and 4 black balls; Bag II has 5 red and 6 black balls. A bag is chosen with probability 1/2 each, then a ball uniformly from that bag. It is red. Find the probability that Bag II was selected. [4 marks]
  1. Let E₁ and E₂ mean selection of Bags I and II, and R mean a red ball. The priors are P(E₁) = P(E₂) = 1/2.
  2. The conditional red probabilities are P(R|E₁) = 3/7 and P(R|E₂) = 5/11, using each bag’s own total number of balls.
  3. Total probability gives P(R) = (1/2) × (3/7) + (1/2) × (5/11).
  4. Bayes’ theorem gives P(E₂|R) = [(1/2) × (5/11)]/[(1/2) × (3/7) + (1/2) × (5/11)] = 35/68.

Key takeaways

  • Define the event and the given condition separately before choosing the numerator and denominator of a conditional probability.
  • Conditional probability divides the intersection probability by the conditioning event’s probability, which must be positive.
  • The multiplication theorem combines successive events using conditional probabilities that reflect all earlier information required by the problem.
  • Independence is a probability product condition; mutual exclusion is the absence of common outcomes in the sample space.
  • Total probability adds weighted contributions across disjoint, exhaustive cases, giving the overall probability of the event.
  • Bayes’ theorem divides one case’s joint contribution by the observation’s total probability to obtain a posterior probability.
  • A random variable assigns numbers to outcomes; its distribution combines probabilities of outcomes that give the same value.
  • The mean is a probability-weighted average of possible values and does not guarantee the result of a particular trial.

Test yourself

What does the denominator in P(E|F) represent?

It is P(F), the probability of the event known to have occurred. It must be positive for the conditional formula to apply.

When may probabilities of A and B be added without subtracting an intersection term?

When the events are mutually exclusive, their intersection is empty and has probability zero, so P(A ∪ B) = P(A) + P(B).

Why does the second draw without replacement require a fresh count?

The first selected object has been removed, changing the total number available and possibly the number favourable to the second event.

What equality tests whether two events E and F are independent?

Check whether P(E ∩ F) = P(E)P(F). If this equality fails, the two events are dependent.

What conditions must the cases satisfy when applying the theorem of total probability?

They must be pairwise disjoint and exhaustive, and each must have positive probability so its conditional probability is defined.

How does a posterior probability differ from a prior probability?

A prior describes a hypothesis before using the specified observation. A posterior is the hypothesis’s conditional probability after incorporating that observation.

For two fair independent coin tosses, why does the number of heads equal 1 with probability 1/2?

The distinct outcomes HT and TH both give one head. Each has probability 1/4, so their probabilities add to 1/2.

How do you calculate the mean from a finite probability distribution?

Multiply each possible value by its probability and add the products, after checking that the probabilities are nonnegative and total one.