Probability begins with an experiment and a model
Probability quantifies uncertainty inside a stated model. The same outcome labels can carry different probabilities, and conditioning changes the population against which an event is compared. This lesson builds exact finite models, proves elementary identities, derives Bayes’ rule, and tests independence and selection claims. The binary positive-result example is synthetic mathematics with specified rates, rather than guidance about any actual diagnostic system.
Retrieval check: Module 03 supplies sets, intersections and complements; Module 05 supplies counting with explicit uniformity assumptions. No calculus is required. The labs use only the Python standard library, exact rational arithmetic and a locally seeded pseudorandom generator.
Sample spaces events and probability axioms
An experiment has a sample space Ω of possible elementary outcomes. An outcome ω is one element, while an event A is a set of outcomes answering a yes/no question. A probability model specifies Ω, an event collection F and a probability function P. In a finite model all subsets can be events, so F is the power set. In more general spaces an event collection is a sigma-algebra: it contains Ω and is closed under complements and countable unions. We state this structure now without requiring measure-theory machinery for our finite computations.
The axioms are P(A)≥0, P(Ω)=1, and countable additivity on pairwise disjoint events: P(∪A_i)=ΣP(A_i). Finite additivity follows by padding a finite disjoint collection with empty events. P(∅)=0 follows from P(Ω)=P(Ω)+P(∅). The complement identity P(Aᶜ)=1−P(A) follows from the disjoint decomposition Ω=A∪Aᶜ. Therefore every probability lies between zero and one. These are model rules; an arbitrary table of scores is not a probability table until nonnegativity and normalisation are established.
In finite Ω, assign masses p_ω≥0 summing to one and define P(A)=Σ_{ω∈A}p_ω. This automatically satisfies the axioms. Raw nonnegative weights w_ω with positive finite total W can be normalised as p_ω=w_ω/W. Uniformity is the special choice p_ω=1/|Ω|, not a consequence of having finitely many outcomes. If some outcomes are impossible in the model, their mass can be zero; their presence in Ω does not force a positive probability. Probabilities express the specified mechanism or belief model, whose adequacy must be evaluated separately.
Take Ω={HH,HT,TH,TT} with weights 1,2,3,4, total 10. Let A mean the first symbol is H and B the second is H. Then P(A)=3/10, P(B)=4/10 and P(A∩B)=1/10. The outcome HH has mass 1/10, not 1/4. These labels do not describe two independent fair coins unless that additional mechanism is explicitly assumed. The finite weighted table is itself a complete valid model.
For a countably infinite example, Ω={1,2,…} with p_n=2^{−n} is normalised because its geometric sum is one. There is no uniform positive mass on every positive integer: any fixed mass c>0 makes the partial sums eventually exceed one, while c=0 gives total zero by countable additivity. Thus “choose a random integer uniformly from all integers” does not specify such a probability distribution. Choosing uniformly from a finite range is a different, well-defined experiment with a stated range.
Countable additivity also gives continuity of probabilities. If A₁⊂A₂⊂… increases to A, write A as the disjoint union of A₁ and differences A_k\A_{k−1}; its probability is the limit of the partial sums P(A_k). For decreasing events use complements to obtain P(∩A_k)=lim P(A_k). This connects an infinite event to finite approximations without replacing a mathematical limit by a finite simulation. General uncountable unions need not be covered by the same axiom; a continuous outcome can have zero point mass while a whole interval has positive probability, as the next module develops.
Set identities determine event operations. A∩B means both, A∪B means at least one, and Aᶜ means the event does not occur. “Exactly one” is (A\B)∪(B\A), not the ordinary union. A partition consists of disjoint events covering Ω and is useful for conditioning by cases. Translate the verbal experiment first: selecting two records with replacement and selecting without replacement use different outcome spaces and dependence mechanisms even if the result labels look similar.
An event selects outcomes from a normalised model. Four labels need not imply four equal masses or two independent fair draws.
What separates an outcome from an event? Can a finite sample space be nonuniform? Why is a uniform distribution on all positive integers impossible under these axioms?
Show answer
An outcome is an element and an event is a set of elements. Arbitrary nonnegative masses summing to one give a finite model. Equal positive integer masses would sum beyond one, whereas equal zero masses sum to zero; neither normalises countably many integers.
Uniformity unions inclusion–exclusion and event bounds
Counting gives P(A)=|A|/|Ω| only when elementary outcomes are equally likely. The choice of elementary outcomes matters: sums of two fair independent dice range from 2 to 12, but those eleven sums are not equally likely. The ordered die pairs are the 36 equiprobable elementary outcomes. Sum seven has six pairs and sum two only one. Grouping elementary outcomes creates a new representation with masses equal to the sums of the grouped probabilities, rather than resetting each group to equal mass.
If A⊂B, decompose B into A and B\A to obtain P(B)=P(A)+P(B\A)≥P(A). This monotonicity lets probability bounds follow from set inclusions. For arbitrary A,B, decompose their union into disjoint A\B, A∩B and B\A. Adding masses gives P(A∪B)=P(A)+P(B)−P(A∩B). The subtraction corrects double counting; it does not assume independence. For three events add singleton probabilities, subtract all three pairwise intersections, and add the triple intersection once.
The union bound P(∪_{i=1}^m A_i)≤Σ_i P(A_i) follows by discarding nonnegative overlap corrections, or by replacing each A_i with its part not already counted. The same disjointification and countable additivity extend it to countably many events. This bound requires no independence and can be loose. Since probability is at most one, report the tighter of one and the sum when that improves presentation. A bound above one is still an algebraically valid upper bound, but supplies no useful additional restriction.
Suppose three components have failure probabilities .01, .02 and .03 in a specified operation. Without a joint model, failure of at least one has probability at most .06 by the union bound and at least .03 by inclusion of the third failure event. If their failures are mutually independent, the exact probability is 1−(.99)(.98)(.97)=.058906. That exact expression needs independence, whereas the interval [.03,.06] does not. Common power loss could make the independence assumption unrealistic.
Disjointness and independence describe different relations. Disjoint events cannot occur together, so P(A∩B)=0. Independence asks for P(A∩B)=P(A)P(B). If both probabilities are positive, disjoint events are necessarily dependent: knowing A occurred excludes B. If one event has probability zero, the product condition can hold even for disjoint events, though conditioning on that null event remains undefined by the elementary ratio. Avoid defining independence solely by a conditional statement with an unavailable denominator.
For the weighted four-symbol table, union mass is .3+.4−.1=.6 and exactly-one mass is .3+.4−2(.1)=.5. Under independent fair symbols, those corresponding masses would be .75 and .5. The exactly-one equality across two different models is coincidental and does not establish that the models are identical. A few matching event probabilities cannot recover every elementary mass unless the chosen events determine the complete distribution.
Event bounds also clarify algorithm claims. If each of m checks fails with probability at most ε, the probability any fails is at most mε even when the failures are dependent. A target total failure bound δ can therefore be obtained by making each bound at most δ/m. This is a sufficient allocation of error budgets, not proof that exactly mε failures occur. Randomised algorithms need a model for their random bits and often condition on a fixed input; randomness of the input is an additional assumption, not supplied by the algorithm’s seed.
For repeated sampling without replacement, the denominator changes with the draw. From a box of three red and two blue tokens, probability of two reds is (3/5)(2/4)=3/10 for ordered draws without replacement. With replacement and independently repeated draws it is (3/5)²=9/25. Both answers use the same initial proportions but different experiments. The counting version agrees: without replacement choose two tokens uniformly from the ten unordered pairs; three pairs are both red. Agreement checks the complete model rather than legitimising an unspecified drawing procedure.
Which assumption justifies favourable-count divided by total-count? Does the union bound need independence? Can two disjoint positive-probability events be independent?
Show answer
Equiprobable elementary outcomes justify counting ratios. The union bound is valid without independence. Positive disjoint events have joint mass zero but a positive product, so they are dependent.
Conditioning chain rules and table denominators
For P(B)>0 define P(A|B)=P(A∩B)/P(B). Conditioning restricts attention to B and renormalises its probabilities. On a finite model, each outcome in B gets conditional mass p_ω/P(B), while outcomes outside B get zero. This produces a valid probability model on the restricted population. The conditioning event need not occur earlier in time and does not automatically describe an intervention. Learning that an event happened and externally forcing a mechanism are different operations unless a causal model justifies equating them.
The denominator is the population after the vertical bar. P(A|B) asks what fraction of B outcomes also satisfy A; P(B|A) asks the reverse fraction. They share the same intersection numerator but generally different denominators. In the weighted example P(A|B)=(1/10)/(4/10)=1/4, while P(B|A)=(1/10)/(3/10)=1/3. The numerical values are conditional probabilities of sets, not informal translations of the same sentence.
In a synthetic population of 100,000 records, label D has prevalence 1%, a positive system result has rate 90% within D, and rate 5% outside D. Expected table counts are D+:900, D−:100, not-D+:4,950, not-D−:94,050. Sensitivity P(+|D)=900/1,000=.9, while P(D|+)=900/5,850=2/13≈.153846. The positive-result population is larger than the D population, so reversing the denominator changes the answer dramatically. These are model-expected counts, not observed data.
Rearranging the conditional definition gives the multiplication rule P(A∩B)=P(B)P(A|B)=P(A)P(B|A), where the relevant denominators are positive. For a sequence, P(A₁∩…∩A_k)=P(A₁)P(A₂|A₁)…P(A_k|A₁∩…∩A_{k−1}), provided each conditioning prefix has positive probability. This chain rule does not assume independence. Independence is the extra simplification that replaces conditional factors by corresponding marginal factors. A probability tree encodes exactly these conditional branches, whose children must sum to one under each positive parent.
If a prefix has zero probability, its intersection with later events also has zero mass. One can compute that branch as zero without pretending a ratio conditioned on the zero prefix is defined. Some models assign branch probabilities to unreachable states as part of a generative specification; these assignments do not determine unique conditional ratios at those states. For continuous observations, conditioning on exact zero-mass points requires a more advanced conditional-density construction later. The elementary event formula cannot be extended by dividing zero by zero.
The law of total probability uses a finite or countable partition B_i with positive-probability cases: P(A)=Σ_i P(A|B_i)P(B_i). It follows by decomposing A into disjoint A∩B_i. Null cases contribute zero through intersection masses without defined conditional ratios. Partition cases must be mutually exclusive and cover the population; overlapping groups cannot be inserted into the sum without correcting double counting. This law combines within-group rates using group proportions, so different mixture proportions can change a pooled rate even if every within-group mechanism stays fixed.
For two-stage selection, suppose a record is from group H with probability .2 and L with probability .8, and event E has respective probabilities .8 and .1. Then P(E)=.2(.8)+.8(.1)=.24. The H-and-E mass is .16, hence P(H|E)=2/3, while the initial H mass was .2. This shift follows from explicit event filtering and does not establish a causal effect of group membership. The joint table or tree provides an auditable calculation when verbal claims obscure the base population.
Chain-rule factorizations also appear in sequence models. P(tokens t₁,…,t_k)=Π_i P(t_i|t₁,…,t_{i−1}) is a general identity for positive-probability prefixes, not an assumption that tokens are independent. A language model approximates those conditional distributions using its own parameterisation and training data. The mathematical factorisation does not certify the quality of those approximations; it separates the probability identity from the modelling choice.
The same joint table supplies sensitivity and posterior, but each uses a different denominator. The positive rate is .0585, not .058.
What makes elementary conditioning defined? Does the chain rule need independence? What weights combine partition-specific conditional rates?
Show answer
The conditioning event must have positive probability. Chain factors are conditional and need no independence. Total probability weights each within-case rate by that case’s probability, with disjoint exhaustive cases.
Bayes’ rule likelihood ratios and prior odds
Bayes’ rule follows by using the same joint mass in both conditional directions: P(A|B)=P(B|A)P(A)/P(B), with P(B)>0 and the right-side conditional defined. A partition gives the denominator as Σ_i P(B|A_i)P(A_i). The prior P(A) specifies mass before learning B; the likelihood P(B|A) describes how likely that evidence is under the case; the posterior renormalises their product after learning B. These are functions of different arguments and cannot be interchanged merely because both are probabilities.
For a binary case D with prior π, sensitivity s=P(+|D), and false-positive rate r=P(+|Dᶜ), total positive probability is πs+(1−π)r. Provided this is positive, posterior is πs/[πs+(1−π)r]. With π=.01,s=.9,r=.05, the true-positive mass is .009 and false-positive mass .0495, giving .0585 total and 2/13 posterior. A rare case can be outnumbered by false positives despite a large sensitivity. The expression makes that effect a countable comparison rather than a slogan about system accuracy.
Prior odds are P(D)/P(Dᶜ)=1/99. The positive likelihood ratio is s/r=.9/.05=18. Posterior odds are 18/99=2/11, giving posterior probability odds/(1+odds)=2/13. For a negative result, likelihood ratio is (1−s)/(1−r)=.1/.95=2/19; posterior odds become 2/1881 and probability 2/1883. The negative posterior is small because negative evidence and the low prior both favour not-D in this model.
The odds derivation divides P(D|B) by P(Dᶜ|B); the common evidence denominator cancels, leaving likelihood ratio times prior odds. It requires nonzero relevant denominators, or a careful limiting/support treatment when evidence is impossible under one case. If r=0 and πs>0, a positive result has posterior one. If πs+(1−π)r=0, positives have zero probability and the elementary posterior is undefined. Distinguish a genuinely impossible-evidence case from a numerical zero caused by rounding or an unsupported model parameter.
When prior π changes but conditional rates s,r remain stable, the posterior changes. At π=.5 with the same rates, positive posterior is .9/(.9+.05)=18/19≈.947368. This is a calculation under a transport assumption: the within-case response rates must remain applicable in the new population. Real dataset changes can alter those rates too, so Bayes’ formula does not promise that simply replacing the prevalence repairs every deployment shift. State which part of the model is assumed invariant and which is re-estimated.
Repeated evidence cannot be multiplied independently without a conditional-independence assumption. If two result events E₁,E₂ are conditionally independent given D and also given Dᶜ, their combined likelihood ratio factors into the product of the two separate ratios. Without both conditions use P(E₁∩E₂|D)/P(E₁∩E₂|Dᶜ) from the joint model. A duplicate of the same observation is perfectly dependent and provides no new information: conditioning twice on B is just conditioning on B once. Counting copied evidence twice can manufacture unjustified confidence.
Bayesian event updating here assumes the model probabilities are supplied. Later modules estimate unknown parameters and introduce priors over parameters; those are additional levels of uncertainty. A known-rate event posterior is neither a confidence interval nor a guarantee about an individual realised outcome. Its meaning is conditional probability within the model. To evaluate a system empirically, record the population, outcomes and sampling protocol, then assess whether the assumptions and estimated rates remain defensible.
For computational implementation form nonnegative joint masses first and then normalise. Check that their sum is positive. With many factors, products may underflow; logarithms and stable normalisation later help preserve meaningful arithmetic, but they do not fix a wrong conditional model. An explanation should report the numerator’s joint event and denominator’s evidence event in words. That check catches reversed conditionals more reliably than memorising the order of letters in a formula.
The explorer changes the prior, sensitivity and false-positive rate for this synthetic binary model. It shows expected counts per 100,000 records, the evidence probability and both conditional directions. Endpoint settings include impossible positive evidence; the posterior must then be labelled undefined rather than shown as zero or left at its previous value. At zero D or not-D prior mass, the generative mechanism can still specify a branch parameter, while its conditional ratio from the joint table is undefined.
Which event does the Bayes denominator describe? Why do copied positive results not supply two independent likelihood ratios? What makes the posterior undefined?
Show answer
The denominator is the probability of the evidence event, here all positives. Copied evidence is the same event, not a conditionally independent new draw. A zero evidence probability makes the elementary conditional ratio undefined.
Pairwise mutual and conditional independence
Events A,B are independent when P(A∩B)=P(A)P(B). With positive denominators this is equivalent to P(A|B)=P(A) and its reverse. Complementation preserves independence: P(A∩Bᶜ)=P(A)−P(A∩B)=P(A)(1−P(B)), and the other complemented forms follow similarly. This exact product identity belongs to the model. Finite sample frequencies will usually fail exact equality even under an independent generating mechanism, so small numerical deviations need statistical interpretation rather than automatic rejection of the mechanism.
For k events, mutual independence requires factorisation for every subcollection, not only every pair and not only the full intersection. Pairwise independence checks the two-event subcollections; it can coexist with a higher-order dependence. For fair independent bits X,Y, define events A={X=1}, B={Y=1}, C={X xor Y=1}. Each has probability 1/2 and each pair has intersection 1/4, yet the triple intersection is empty because X=Y=1 makes their xor zero. The triple product is 1/8, disproving mutual independence.
The four equiprobable outcomes (0,0),(0,1),(1,0),(1,1) give A∩B={(1,1)}, A∩C={(1,0)} and B∩C={(0,1)}. All three pair masses are 1/4, matching (1/2)². But A∩B∩C=∅ has mass zero. A reliability calculation multiplying three marginal success probabilities would therefore be wrong even though all pair checks pass.
Conditional independence given E with P(E)>0 is independence inside the conditional model: P(A∩B|E)=P(A|E)P(B|E). Conditioning can create or destroy independence. Given each value of a latent group, two measurements may be independent, while mixing groups creates marginal association. Conversely selecting records by a condition involving two originally independent events can create dependence among the selected records. Neither conditional nor marginal independence automatically implies the other.
For a latent binary group Z with equally probable values, let A and B be conditionally independent with rate .9 when Z=1 and .1 when Z=0. Marginal P(A)=P(B)=.5, but P(A∩B)=.5(.81)+.5(.01)=.41, exceeding .25. Shared group variation explains the association despite within-group independence. This example gives a complete joint mechanism: draw Z then draw A and B independently using that group’s rate. Declaring independence merely from the existence of distinct variables would erase the common source.
For the opposite selection effect, start with independent fair bits A,B and retain only S={A or B}. The selected outcomes are (1,0),(0,1),(1,1), each with conditional mass 1/3. Thus P(A|S)=P(B|S)=2/3 and P(A∩B|S)=1/3, which differs from 4/9. Within the selected records, observing A=0 forces B=1. The original independent model remains correct; selection creates a new conditional population in which it no longer factors.
The xor table passes every pair test but fails the triple test. Independence is a claim about specified collections and the specified conditioning population.
Mutual independence of generated outcomes should also be distinguished from reproducibility. A fixed seed makes a pseudorandom simulation repeatable; it does not magically create independent experiments if the same generated sequence is reused and treated as fresh data. Mathematical analyses usually idealise random draws with specified independence; the implementation should state its generator, local seed and reuse pattern. Correlated draws and repeated evidence affect probability products, and later modules quantify their effect on uncertainty estimates.
Does pairwise independence imply mutual independence? Does conditional independence imply marginal independence? Can retaining A or B create dependence from independent fair bits?
Show answer
The xor construction disproves the first claim. A latent-group mixture disproves the second. Conditioning on the union removes the (0,0) outcome and gives joint 1/3 versus product 4/9, so the selected population is dependent.
Base populations selection effects and model verification
Every rate has a denominator and a population. A positive-only sample estimates a fraction conditional on positivity, rather than the population prevalence. In the synthetic table, the selected D fraction is 2/13 while population prevalence is 1/100. A selection rule depending on outcomes therefore changes what a reported frequency estimates. Enlarging that selected sample can estimate the wrong target ever more precisely; size does not remove a denominator mismatch. State the target event and population before collecting or interpreting counts.
Pooled comparisons also depend on group mixtures. Suppose system A succeeds on 9/10 easy records and 30/100 hard records, while B succeeds on 80/100 easy and 2/10 hard records. A has higher within-group success in both groups: .9>.8 and .3>.2. Yet A’s pooled success is 39/110≈.354545, below B’s 82/110≈.745455, because A was evaluated mostly on hard records. This Simpson reversal is an arithmetic mixture effect, not a contradiction in probability.
With a specified target mixture of half easy and half hard, A’s standardised success is .5(.9)+.5(.3)=.6 and B’s is .5(.8)+.5(.2)=.5. This answers a different, explicitly shared evaluation question from the original pooled samples. It assumes the within-group rates transport to that target. Neither standardisation nor the raw observational comparison alone proves that changing the system causes the difference; causal identification needs a suitable design and assumptions developed later.
An auditable finite model has a complete outcome list or generative tree, nonnegative masses, a normalised total, precise events and explicit conditioning. Check that complementary branches sum to one and a partition really covers the space. Test identities with exact fractions where practical. A finite enumeration can prove an identity for that finite model, while a simulation only samples its mechanism. Neither establishes that the chosen mechanism adequately represents every real application.
Lab 1 enumerates weighted outcomes and the xor construction exactly. Lab 2 generates a million locally seeded synthetic records and compares conditional frequencies at increasing checkpoints with the derived posterior. A trace can fluctuate, and its error need not decrease at every checkpoint. This lesson does not infer a convergence theorem from one run. The law of large numbers, concentration and Monte Carlo error statements arrive in Module 24 with their assumptions. Report zero-denominator frequencies as undefined, especially in small or rare-event samples.
Lab 3 separates sensitivity from posterior, demonstrates a selected-population dependency and computes the Simpson reversal. Its examples use known mathematical masses or explicit finite tables. When a real dataset is substituted, rates become estimates of an unknown model and need uncertainty analysis. A synthetic expected table can contain fractional expected counts and is not a promise that a finite realised sample has those exact counts. Keep model probabilities, expected populations and observed frequencies labelled distinctly.
Within-group success and pooled success answer different questions when evaluation mixtures differ. A shared mixture makes the comparison’s denominator explicit.
The same checks matter in CS and AI. Reliability needs joint failure assumptions; randomised algorithms need stated randomness and input conditions; classifiers need a target population and class-conditional rates; sequence factorizations need the right conditional prefixes. Bayes identities connect events within a model but do not manufacture model validity, independent evidence or causal interpretation. Your exit result is a complete finite model and a correctly explained conditional calculation, with uncertainty about the model itself kept visible.
Does a large selected sample fix a wrong target population? Can simulation alone prove the posterior formula? What does a common-mixture comparison assume?
Show answer
No: selection changes the target even at large size. The formula follows from the axioms and conditional definition; simulation illustrates a specified mechanism. Standardisation assumes within-group rates apply to the shared target mixture and does not alone establish causation.
Common misconceptions
| Claim | Repair |
|---|---|
| Finite outcomes are automatically uniform. | Specify elementary masses or a mechanism establishing uniformity. |
| Disjoint events are independent. | Positive disjoint events cannot factor. |
| P(A | B)=P(B |
| The chain rule assumes independence. | Its conditional factors are general. |
| High sensitivity means a high positive posterior. | The prior and false-positive mass also matter. |
| Pairwise independence licences every product. | Mutual independence checks all subcollections. |
| Conditioning preserves independence. | Mixture and selection can alter it. |
| More selected data fixes selection bias. | More data can estimate a different target more precisely. |
Three reproducible labs
Lab 1 · Exact weighted events and independence
Download lab1_exact_weighted_events.py
"""Exact finite probability models and pairwise versus mutual independence."""
from fractions import Fraction as F
from itertools import product
if __name__ == "__main__":
weights = {"HH": 1, "HT": 2, "TH": 3, "TT": 4}
total = sum(weights.values())
probability = lambda event: sum((F(weights[w], total) for w in event), F(0))
A = {w for w in weights if w[0] == "H"}
B = {w for w in weights if w[1] == "H"}
print("normalised masses:", {w: str(F(n, total)) for w, n in weights.items()})
print("P(A)=%s P(B)=%s P(intersection)=%s P(union)=%s" % (
probability(A), probability(B), probability(A & B), probability(A | B)))
print("P(A|B)=%s P(B|A)=%s" % (
probability(A & B)/probability(B), probability(A & B)/probability(A)))
assert probability(A | B) == probability(A)+probability(B)-probability(A & B)
print("independent:", probability(A & B) == probability(A)*probability(B))
omega = list(product((0, 1), repeat=2))
events = [set(w for w in omega if w[0] == 1),
set(w for w in omega if w[1] == 1),
set(w for w in omega if w[0] ^ w[1] == 1)]
p = lambda event: F(len(event), len(omega))
for i, j in ((0, 1), (0, 2), (1, 2)):
joint, factor = p(events[i] & events[j]), p(events[i])*p(events[j])
print("pair %d,%d: joint=%s product=%s" % (i+1, j+1, joint, factor))
assert joint == factor
triple = p(events[0] & events[1] & events[2])
factor = p(events[0])*p(events[1])*p(events[2])
print("triple: joint=%s product=%s" % (triple, factor))
assert triple != factor
print("PASS: weighting matters; pairwise independence is not mutual independence")
normalised masses: {'HH': '1/10', 'HT': '1/5', 'TH': '3/10', 'TT': '2/5'}
P(A)=3/10 P(B)=2/5 P(intersection)=1/10 P(union)=3/5
P(A|B)=1/4 P(B|A)=1/3
independent: False
pair 1,2: joint=1/4 product=1/4
pair 1,3: joint=1/4 product=1/4
pair 2,3: joint=1/4 product=1/4
triple: joint=0 product=1/8
PASS: weighting matters; pairwise independence is not mutual independence
Write each event set before running and predict both conditional denominators. Verify normalisation, inclusion–exclusion and the pair/triple xor distinction using exact fractions.
Lab 2 · Synthetic base-rate simulation
Download lab2_base_rate_simulation.py
"""Seeded synthetic positive-result model; frequencies are not medical advice."""
from fractions import Fraction as F
import random
if __name__ == "__main__":
prevalence, sensitivity, false_positive = F(1, 100), F(9, 10), F(1, 20)
positive_mass = prevalence*sensitivity+(1-prevalence)*false_positive
posterior = prevalence*sensitivity/positive_mass
print("synthetic model: P(D)=1/100 P(+|D)=9/10 P(+|not D)=1/20")
print("exact P(+)=%s P(D|+)=%s = %.8f" % (positive_mass, posterior, float(posterior)))
generator = random.Random(20261004)
tp = fp = fn = tn = 0
checkpoints = {1000, 10000, 100000, 1000000}
for n in range(1, max(checkpoints)+1):
d = generator.random() < float(prevalence)
positive = generator.random() < float(sensitivity if d else false_positive)
if d and positive:
tp += 1
elif d:
fn += 1
elif positive:
fp += 1
else:
tn += 1
if n in checkpoints:
estimate = tp/(tp+fp) if tp+fp else None
sensitivity_estimate = tp/(tp+fn) if tp+fn else None
print("N=%7d TP=%5d FP=%5d FN=%4d TN=%6d posterior=%s sensitivity=%s" % (
n, tp, fp, fn, tn,
"undefined" if estimate is None else "%.6f" % estimate,
"undefined" if sensitivity_estimate is None else "%.6f" % sensitivity_estimate))
assert tp+fp+fn+tn == n
print("absolute posterior error:", "%.8f" % abs(estimate-float(posterior)))
print("One seeded trace is a model illustration, not a convergence proof or fitted parameter claim.")
synthetic model: P(D)=1/100 P(+|D)=9/10 P(+|not D)=1/20
exact P(+)=117/2000 P(D|+)=2/13 = 0.15384615
N= 1000 TP= 8 FP= 47 FN= 1 TN= 944 posterior=0.145455 sensitivity=0.888889
N= 10000 TP= 90 FP= 488 FN= 16 TN= 9406 posterior=0.155709 sensitivity=0.849057
N= 100000 TP= 914 FP= 5036 FN= 105 TN= 93945 posterior=0.153613 sensitivity=0.896958
N=1000000 TP= 9019 FP=49418 FN= 989 TN=940574 posterior=0.154337 sensitivity=0.901179
absolute posterior error: 0.00049100
One seeded trace is a model illustration, not a convergence proof or fitted parameter claim.
Predict the posterior 2/13 and explain every counter. Check that cumulative counts sum to N. Distinguish conditional empirical sensitivity from conditional empirical posterior, and explain why a single seeded trace is illustrative rather than a convergence proof.
Lab 3 · Reversed conditioning and selected populations
Download lab3_conditioning_and_selection_faults.py
"""Exact denominators, selected populations, and a Simpson reversal."""
from fractions import Fraction as F
from itertools import product
if __name__ == "__main__":
# Model masses represented by expected counts in a synthetic population.
tp, fp, fn, tn = 900, 4950, 100, 94050
sensitivity = F(tp, tp+fn)
posterior = F(tp, tp+fp)
print("P(+|D)=%s P(D|+)=%s" % (sensitivity, posterior))
print("positive-only sample disease fraction:", posterior)
print("population disease fraction:", F(tp+fn, tp+fp+fn+tn))
assert posterior != sensitivity
omega = list(product((0, 1), repeat=2))
selected = [w for w in omega if w[0] or w[1]]
pa = F(sum(w[0] for w in selected), len(selected))
pb = F(sum(w[1] for w in selected), len(selected))
joint = F(sum(w[0] and w[1] for w in selected), len(selected))
print("independent source fair bits; keep A or B")
print("selected P(A)=%s P(B)=%s P(A and B)=%s product=%s" % (pa, pb, joint, pa*pb))
assert joint != pa*pb
table = {"A": {"easy": (9, 10), "hard": (30, 100)},
"B": {"easy": (80, 100), "hard": (2, 10)}}
for name, groups in table.items():
easy = F(*groups["easy"])
hard = F(*groups["hard"])
pooled = F(sum(r[0] for r in groups.values()), sum(r[1] for r in groups.values()))
standardised = (easy+hard)/2
print("system %s: easy=%s hard=%s pooled=%s common-half-mixture=%s" % (
name, easy, hard, pooled, standardised))
print("A is better in each stratum, worse pooled; the evaluation mixtures differ.")
print("PASS: conditioning changes denominators and selection changes the target population")
P(+|D)=9/10 P(D|+)=2/13
positive-only sample disease fraction: 2/13
population disease fraction: 1/100
independent source fair bits; keep A or B
selected P(A)=2/3 P(B)=2/3 P(A and B)=1/3 product=4/9
system A: easy=9/10 hard=3/10 pooled=39/110 common-half-mixture=3/5
system B: easy=4/5 hard=1/5 pooled=41/55 common-half-mixture=1/2
A is better in each stratum, worse pooled; the evaluation mixtures differ.
PASS: conditioning changes denominators and selection changes the target population
Repair the denominator reversal, derive the selected union table and compare raw pooled rates with a specified common mixture. Describe exactly which target each quantity answers.
Fourteen exercises with full solutions
Exercises 1–12 are required. Optional Exercises 13–14 add 35 minutes beyond the ten-hour core schedule.
For weights HH:1, HT:2, TH:3, TT:4, compute P(first H), P(second H) and their union.
Show solution
Total 10 gives .3 and .4; joint HH has .1. Union .3+.4−.1=.6. Equal label counts are irrelevant because masses are unequal.
Use the same table to compute P(first H|second H) and the reverse conditional.
Show solution
The shared numerator is .1; denominators are .4 and .3 respectively, giving 1/4 and 1/3. Both conditioning events have positive mass, so both ratios are defined.
Three event probabilities are .02,.03,.04. Give an independence-free upper bound on their union and a lower bound.
Show solution
Union bound gives at most .09; inclusion of the largest event gives at least .04. No exact value follows without intersections or another joint assumption.
For π=.01, s=.9, r=.05, compute P(+), P(D|+) and the positive likelihood ratio.
Show solution
P(+)=.009+.0495=.0585=117/2000. Posterior .009/.0585=2/13≈.153846. Likelihood ratio s/r=18, which multiplies prior odds 1/99 to give posterior odds 2/11.
Prove the complement identity and event monotonicity from disjoint additivity.
Show solution
Ω=A∪Aᶜ disjoint gives 1=P(A)+P(Aᶜ). If A⊂B, disjoint B=A∪(B\A) gives P(B)=P(A)+P(B\A)≥P(A). Nonnegativity is used in the last step.
Prove total probability for a finite partition and derive Bayes’ rule with its domain conditions.
Show solution
A is the disjoint union of A∩B_i, so P(A)=ΣP(A∩B_i)=ΣP(A|B_i)P(B_i) for positive cases; null cases contribute zero via intersections. The shared mass P(A∩B)=P(B|A)P(A)=P(A|B)P(B) gives Bayes when the relevant denominators are positive. Division by a null conditioning event is unavailable.
Prove independence survives replacing B by its complement and explain the null-event distinction.
Show solution
P(A∩Bᶜ)=P(A)−P(A∩B)=P(A)(1−P(B))=P(A)P(Bᶜ). The product identity works even if a marginal is zero. Its equivalent conditional expression requires a positive conditioning denominator, so one cannot use P(A|B) when P(B)=0.
A box has three red and two blue tokens. Compare the probability of two red draws with and without replacement.
Show solution
Without replacement it is (3/5)(2/4)=3/10; with independent replacement it is (3/5)²=9/25. Alternatively, without replacement there are ten equiprobable unordered pairs and three red pairs. The drawing mechanism determines the appropriate calculation.
Groups H,L have masses .2,.8 and E rates .8,.1. Compute P(E) and P(H|E), then explain the denominator.
Show solution
P(E)=.16+.08=.24. H-and-E mass .16 gives P(H|E)=.16/.24=2/3. The denominator includes E records from both groups, not all H records. This is conditional selection, with no causal conclusion supplied.
In the synthetic binary model, change π to .5 while holding s=.9,r=.05 fixed. Compute the new positive posterior and state the transport assumption.
Show solution
Joint masses .45 and .025 give posterior .45/.475=18/19≈.947368. The calculation assumes both class-conditional rates remain valid in the new population; a prevalence change need not be the only real dataset change.
Three xor events pass every pair independence check. A report multiplies their three marginal masses. Repair it.
Show solution
Each event has mass 1/2 and each pair joint 1/4, but the triple event is empty, mass zero versus product 1/8. Pairwise independence is insufficient for a three-factor product; mutual independence requires all subcollections.
A positive-only sample reports disease fraction 2/13 as population prevalence, and a copied test result is treated as new independent evidence. Repair both.
Show solution
The selected fraction estimates P(D|+), whereas the model prevalence is P(D)=1/100. Selection changes the denominator. A copied result is the same event, so conditioning again gives no new evidence; multiplying its likelihood ratio twice unjustifiably assumes conditional independence.
For independent fair A,B, retain only A or B. Derive the selected joint table and disprove conditional independence.
Show solution
Selection mass is 3/4; retained (1,0),(0,1),(1,1) each have conditional mass 1/3. Each event has conditional marginal 2/3 but joint 1/3≠4/9. Original independence is preserved as a statement about the original population, not the selected one.
Compute pooled and half-easy half-hard success for A with 9/10 easy and 30/100 hard, and B with 80/100 easy and 2/10 hard. Explain the reversal.
Show solution
Pooled A=39/110 and B=82/110, so B appears better. Within groups A=.9,.3 exceeds B=.8,.2. The common half mixture gives A=.6>B=.5. Different group weights cause the reversal. Standardisation specifies a target and assumes within-group transport; it does not alone identify a causal effect.
Ten-question self-check
Show answer
Masses D+ .009, D− .001, not-D+ .0495, not-D− .9405 sum to one. Positive mass .0585 gives posterior .009/.0585=2/13. Prior odds 1/99 times likelihood ratio 18 give 2/11 posterior odds and 2/13 probability. Duplicate evidence is not a new conditionally independent event; a positive-only sample estimates P(D|+) rather than prevalence .01. These statements belong to the specified synthetic model.
Reading with a purpose
Use the probability-model and conditioning materials in MIT 6.041SC and the discrete probability selections in MIT Mathematics for Computer Science. All worked tables and demonstrations here are original constructions with explicit assumptions.
| When | Selection and question |
|---|---|
| Session 1 · 15 minutes | Models and axioms: where does uniformity enter a counting calculation? |
| Session 4 · 15 minutes | Conditioning and independence: what denominator and factorisation does each claim require? |
Retrieval exit task and next step
Create a complete weighted model, calculate a union and two reversed conditionals, derive Bayes in masses and odds, and disprove a pairwise-to-mutual claim. Identify the population before interpreting a selected rate.
Exit task: reproduce the four-cell base-rate table and posterior 2/13, explain why .9 sensitivity answers a different question, and show why copying evidence does not justify a second independent update.
Ready to move on: you can specify probabilistic assumptions and conditional denominators. Random variables next map outcomes to numerical observations and distinguish discrete masses, continuous densities and distribution functions. See the course overview.
Notation and bilingual terminology
| Term or notation | Meaning | 中文 |
|---|---|---|
| Ω / ω / A | Sample space / outcome / event | 样本空间、结果、事件 |
| P(A | B) | Conditional probability on a positive-mass B |
| Partition / union bound | Exhaustive disjoint cases / overlap-safe upper bound | 划分、并集界 |
| Prior / likelihood / posterior | Before-evidence mass / evidence rate / updated mass | 先验、似然、后验 |
| Odds / likelihood ratio | p/(1−p) / evidence-rate ratio | 几率、似然比 |
| Pairwise / mutual independence | Pair products / every subcollection product | 两两、相互独立 |
| Selection / target mixture | Conditional population / specified group weights | 选择、目标混合 |