Ran Wei/Maths Series
中文
Mathematical Foundations for Computer Science and AI — Ran Wei

Vectors, geometry, and array notation

Compute with coordinate vectors, justify lengths and angles, choose meaningful feature scales, and make NumPy array shapes part of the mathematical specification.

8 hours4 sessions3 labs12 exercises + 2 extensions10 quiz questions

By the end you can

  • Distinguish points, displacements, coordinate tuples, and storage arrays.
  • Compute linear and affine combinations with their interpretation.
  • Derive Cauchy–Schwarz and Euclidean triangle inequalities from dot products.
  • Compare norms, unit directions, cosine similarity, and feature-scale choices.
  • Predict flat/row/column shapes and repair accidental broadcasting.

Before you start

Modules 01 and 03: real numbers, functions, indexed sums, and Cartesian products. Complete NumPy array preparation before the labs; no earlier linear algebra is assumed.

Contents

Study plan

8 hours

Times include practice and are estimates. Split a session when useful. Optional extension exercises add 35 minutes. Progress is stored in this browser and shared between language editions.

1

Numbers become geometry only after choosing coordinates

A data record can be written as a list of numbers. Comparing two records by a dot product or distance looks straightforward, but the result depends on feature order, units, and the operation chosen. Converting centimetres to metres can change an unweighted nearest neighbour even though the physical records are unchanged. Mixing a column array with a flat array can produce every pairwise difference instead of the intended paired differences.

This module develops the mathematics and the implementation contract together. You will distinguish points from displacements, derive lengths and angle bounds, compare norms, and inspect shapes before computing. A successful numerical call establishes that an operation was executable; a mathematical interpretation needs its own assumptions about what the coordinates represent.

Retrieval check: work with real tuples and finite sums, and describe a Cartesian product. Review Module 01 and Module 03 if needed. There is no assumed earlier linear algebra. Complete NumPy array preparation before the labs; its roughly thirty-minute orientation is additional preparation when you need it. The numerical baseline uses Python 3.11 and the pinned NumPy environment described there.

2

Scalars coordinate vectors points and displacements

A scalar is a single number used to scale a vector; here scalars are real numbers. A coordinate vector in Rd\mathbb R^d is an ordered d-tuple of real entries. The positive integer d is the number of coordinates. The Cartesian product construction from Module 03 supplies the set of all such tuples. The coordinate order is part of the specification: (height,age) and (age,height) use different feature conventions even if both arrays have length two.

In geometry, a vector can describe a displacement independent of where it is drawn. Its arrow records both magnitude and direction. Sliding the arrow without rotating or stretching it preserves that displacement. In data work, a tuple can instead collect feature measurements, some of which have different units. Writing those measurements as a vector makes coordinate algebra possible, but it does not automatically make a Euclidean length physically meaningful.

A point is a location, whereas a displacement connects locations. Given points p,q in a chosen coordinate frame, q−p is the displacement from p to q. Translating the origin adds the same coordinate offset to both points and leaves their difference unchanged. Position coordinates and displacement coordinates may use the same tuple notation, so the surrounding interpretation must identify which object is being added or compared.

Worked example
A displacement is independent of the origin

Let p=(1,2),q=(4,6). The displacement q−p is (3,4). Moving the coordinate origin so both displayed points gain (10,−5) gives p′=(11,−3),q′=(14,1), and q′−p′ remains (3,4). An arrow from (0,0) to (3,4) represents that same displacement without being the original point q. Its Euclidean length, developed below, is five.

Translation of points and displacementsPoints p=(1,2), q=(4,6) have difference (3,4). New points (11,−3), (14,1) have the same difference.A common offset changes point coordinates, not displacementp = (1,2)q = (4,6)(3,4)p′ = (11,−3)q′ = (14,1)(3,4)Both new tuples add (10,−5); the difference stays (3,4).
Figure 9.1

The point-to-point arrow and its translated copy have the same coordinate difference. Coordinates of individual points change when the frame changes.

A zero vector has every entry zero and describes no displacement. It is a valid vector with length zero, not a missing observation. A record with unknown age also needs a separate missing-data representation; substituting zero would assert a numerical value and change the geometry. Similarly, a vector of category codes such as 1,2,3 does not acquire meaningful distances merely because the codes are numeric. The encoding determines which comparisons make sense.

Coordinate unit vectors eᵢ have one in position i and zero elsewhere. A tuple x can be written ∑i=1dxiei\sum_{i=1}^d x_i e_i: each coordinate contributes along its own axis. In two dimensions this gives (3,4)=3(1,0)+4(0,1). These coordinate directions have unit length and are mutually perpendicular under the ordinary dot product. Module 11 develops the more general ideas of basis, dimension, and coordinate changes.

Two vectors to be added or compared must have the same coordinate dimension and corresponding meanings. A three-feature vector cannot be paired entrywise with a two-feature vector by silently dropping a coordinate. Even equal lengths are insufficient if one record uses a different feature order. The lab checks lengths before using Python zip because zip otherwise stops at the shorter input and could hide a modelling error.

The mathematical dimension d differs from storage axes. A single three-coordinate vector may be stored in a NumPy array with one axis (3,) or two axes (3,1). A table of one hundred such vectors has shape (100,3), still representing observations with three coordinates each. Keeping these counts separate prevents confusion between the number of examples, the number of features, and the number of axes in a container.

Check your understanding

What is the displacement from (−1,3) to (2,−1)? Does a zero vector mean missing data? Does a (100,3) table make each observation a hundred-dimensional vector?

Show answer

The displacement is (3,−4). A zero vector contains known zeros; missingness is a different modelling condition. Each row of the table has three coordinates, while one hundred counts observations and the array has two axes.

3

Addition scaling linear and affine combinations

Vector addition and scalar multiplication act coordinatewise: (u+v)i=ui+vi(u+v)_i=u_i+v_i and (αu)i=αui(\alpha u)_i=\alpha u_i. The result stays in the same coordinate dimension. Adding displacement arrows corresponds to placing the second arrow’s tail at the first arrow’s head; the combined displacement joins the starting point to the final endpoint. Negative scaling reverses direction, positive scaling preserves it, and scaling by zero gives the zero vector.

Coordinatewise real arithmetic implies commutativity and associativity of vector addition, additive identity zero, additive inverse −u, and distributive scaling rules. For example α(u+v)=αu+αv because the equality holds in every coordinate. These laws explain why finite sums of vectors can be rearranged or grouped under the stated real-number model. In floating-point software, different grouping can change small rounding errors, which Module 28 studies separately.

A linear combination has the form ∑j=1kαjvj\sum_{j=1}^k\alpha_jv_j with real coefficients and equal-dimensional vectors. Coefficients need not be positive or sum to one. The vector 2(3,4)+3(4,−3)=(18,−1) is a valid combination. Its coordinates may lie beyond either original vector’s endpoints. A linear combination is an algebraic construction, not necessarily an interpolation or a probability-weighted average.

A line through a point p in nonzero direction u is p+tu for real t. Restricting t to [0,1] gives the segment from p to p+u; other t values continue along the same line. If u is zero, this expression produces only p, so calling it a full line would be misleading. An affine line allows an offset p, whereas a line consisting only of tu passes through the chosen origin.

An affine combination of points has coefficients summing to one. For two points it is (1−t)p+tq. When 0≤t≤1 it interpolates the segment; when t lies outside that interval it extrapolates. Translating each point by b changes the combination by exactly b, since the coefficients sum to one. A coefficient sum other than one would scale that translation, making the interpreted result depend incorrectly on the point-coordinate origin.

Worked example
Interpolation keeps the point interpretation

For p=(1,2),q=(4,6), take t=3/4. The affine point is (1/4)p+(3/4)q=(3.25,5), lying three quarters of the way from p to q. Its displacement from p is (3/4)(q−p)=(2.25,3). At t=2 the expression gives (7,10), a point beyond q on the same line. The coefficients still sum to one, although one is negative, so this is affine extrapolation rather than a convex interpolation.

A convex combination is an affine combination with nonnegative coefficients. The coefficient conditions are algebraic; whether they represent actual probabilities requires a separate model. In data processing, a convex mixture of feature vectors may be meaningful for continuous measurements but not for every category encoding or constrained object. A mixture of two valid numerical records need not correspond to a physically valid or observed individual.

Feature operations must retain units. Doubling a displacement measured in metres preserves its unit; adding metres to metres is meaningful. Summing coordinates measured in seconds and metres as a single physical length has no automatic interpretation. You may deliberately choose a weighted numerical geometry for learning, but document that choice and its scale factors. Algebraic compatibility and application meaning are both necessary parts of a useful calculation.

Coordinate rearrangement must be applied consistently to all compared objects. Reordering both vectors by the same permutation preserves their dot product and the standard norms later defined. Reordering only one mixes feature positions and changes the result. This distinction is a simple example of separating a change of representation from a change of the represented relationship; shape checks alone cannot detect a feature-order mismatch.

For example, one record (height,age)=(170,20) and another stored as (age,height)=(30,180) have matching array lengths but incompatible coordinate meanings. A computation pairing their first entries multiplies height by age. Reordering the second to (180,30) repairs correspondence, while choosing feature scales is still a separate task. A reliable data interface documents names and units alongside the shape tuple.

Check your understanding

Is 2p+3q an affine point combination? When does p+tu describe a genuine line? Why does (1−t)p+tq respond correctly to shifting the coordinate origin?

Show answer

The first coefficient sum is five, so it is not affine. A genuine line requires u≠0. The affine coefficients sum to one, so adding offset b to both points adds exactly b to their combination.

4

Dot products orthogonality angles and Cauchy–Schwarz

The ordinary real dot product is u⋅v=∑i=1duiviu\cdot v=\sum_{i=1}^d u_iv_i. It is a scalar, unlike entrywise multiplication, which retains a vector of products. It is symmetric and linear in each argument: u·(αv+βw)=α(u·v)+β(u·w). Its self-product u·u is a sum of squares, nonnegative and zero exactly when every coordinate is zero. Define Euclidean length as ∥u∥2=u⋅u\|u\|_2=\sqrt{u\cdot u}.

Vectors are orthogonal when u·v=0. For nonzero vectors in ordinary geometry this means perpendicular directions. The zero vector is algebraically orthogonal to every vector but has no direction or ordinary angle. That boundary shows why the dot-product predicate and a geometric angle statement have different domain conditions. In the plane, (3,4) and (4,−3) are orthogonal because twelve minus twelve is zero.

To introduce angles, an acute right-triangle cosine is the adjacent side divided by the hypotenuse. Extend it to signed directional alignment: cosine is one for the same direction, zero at a right angle, and minus one for opposite directions. Angles between nonzero vectors run from zero to 180 degrees, equivalently zero to π radians; π radians denotes a half-turn. The inverse cosine selects the angle in that range for an alignment value between minus one and one.

For two nonzero coordinate vectors, their geometric relation is u⋅v=∥u∥2∥v∥2cos⁡θu\cdot v=\|u\|_2\|v\|_2\cos\theta. In the plane, choose unit direction û=u/||u||₂ and its perpendicular (−u₂,u₁)/||u||₂. The signed component of v along û is v·û; dividing it by ||v||₂ is the cosine by the right-triangle interpretation. Multiplying back gives the formula. This is not a claim that every tuple of unrelated feature measurements already has a useful physical angle.

Worked example
A unit-axis component and a perpendicular direction

For w=(2,3) and e₁=(1,0), w·e₁=2, so its component along that unit axis is 2e₁=(2,0). The remainder (0,3) is orthogonal to e₁. For u=(3,4),v=(4,−3), both lengths are five and the dot product is zero, so their angle is 90 degrees. Projection onto a general subspace is developed in Module 12; this example needs only a coordinate unit direction.

A component and orthogonal remainderVector (2,3) has first-axis component (2,0) and perpendicular remainder (0,3). The dot product with the first unit axis is two.The component along the first unit axis is twow = (2,3)2e₁ = (2,0)(0,3)0w·e₁ = 2(0,3)·e₁ = 0
Figure 9.2

The coordinate component and perpendicular remainder form a right triangle. A dot product selects the signed component along a unit direction.

The Cauchy–Schwarz inequality states ∣u⋅v∣≤∥u∥2∥v∥2|u\cdot v|\leq\|u\|_2\|v\|_2. Give a proof before dividing by lengths to form a cosine. If v=0, both sides are zero. Otherwise let t=(u·v)/(v·v). Expanding the nonnegative square length ||u−tv||₂² gives ||u||₂²−(u·v)²/||v||₂²≥0. Multiply by the positive denominator and take nonnegative square roots to get the claimed inequality.

With both vectors nonzero, equality holds exactly when u−tv=0 in that proof, meaning one is a scalar multiple of the other. Positive multiples give cosine one and negative multiples give minus one. If either vector is zero, equality also holds in the inequality, while a cosine remains undefined. Consequently a zero self-product should never be used as a divisor simply because the inequality itself includes zero vectors.

Cauchy–Schwarz guarantees the nonzero-vector cosine quotient lies in [−1,1]. In floating-point calculations a rounding error can push a value slightly outside that interval, so angle display may clamp a numerically computed quotient after the nonzero and shape checks. This is a local numerical display policy, not a new mathematical inequality. Large magnitudes or underflow can require additional numerical handling beyond our small-coordinate labs.

The dot product depends on the chosen coordinate geometry. Scaling one feature changes its contribution to products and angles; a weighted inner product can model a different geometry, but its assumptions must be stated. In AI, attention and many similarity calculations use dot products of learned feature coordinates. Their application meaning depends on how those coordinates were constructed, not only on the arithmetic fact that corresponding entries can be multiplied.

Check your understanding

Prove the dot product of (3,4) with (4,−3) is zero. Which condition makes the cosine quotient defined, and which inequality bounds it?

Show answer

The product is 3·4+4·(−3)=0. Both vectors must be nonzero for division by their lengths. Cauchy–Schwarz bounds the quotient between minus one and one; zero vectors still satisfy the inequality but have no defined quotient.

5

Norms metrics and the effect of feature units

A norm assigns a nonnegative length to each vector, zero only at zero, satisfies ||αu||=|α| ||u||, and obeys the triangle inequality ||u+v||≤||u||+||v||. Nonnegativity alone is insufficient: a proposed length must satisfy every law. These properties ensure that scaling, cancellation, and path-length arguments behave consistently. The Euclidean formula just introduced is one norm among several useful coordinate choices.

The Manhattan norm is ∥u∥1=∑i∣ui∣\|u\|_1=\sum_i|u_i|. The maximum norm is ∥u∥∞=max⁡i∣ui∣\|u\|_\infty=\max_i|u_i|. Both have positivity and scaling directly from absolute values. For Manhattan’s triangle law, apply |uᵢ+vᵢ|≤|uᵢ|+|vᵢ| and sum. For the maximum norm, every coordinate sum is at most ||u||∞+||v||∞ in absolute value, so taking the maximum proves its triangle law.

For Euclidean length, expand ||u+v||₂²=||u||₂²+2u·v+||v||₂². Cauchy–Schwarz bounds the middle term by 2||u||₂||v||₂, making the result at most (||u||₂+||v||₂)². Both sides are nonnegative, so taking square roots proves the triangle inequality. This links a dot-product theorem to a norm axiom rather than assuming the ordinary diagram proves the law in every coordinate dimension.

The squared Euclidean length is useful in objectives but is not itself a norm. Scaling a vector by α scales its squared length by α², rather than |α|. For u=v=(1,0), the squared length of u+v is four, greater than the sum one plus one, so the norm triangle inequality fails. Calling an optimisation penalty a squared norm describes how it was constructed; it does not transfer all norm laws to that new function.

Squaring a nonnegative distance preserves ranking, which explains why nearest-neighbour code can omit a square root when only the ordering is needed. It changes the numeric distance and its metric properties, however: on the real line the squared distance from zero to two is four, exceeding the squared zero-to-one plus one-to-two distances of two. A ranking shortcut is therefore different from substituting a new metric into a proof that needs the triangle inequality.

Worked example
The same displacement has several valid lengths

For u=(3,4), the Manhattan length is seven, Euclidean length five, and maximum length four. Thus the distance from p=(1,2) to q=(4,6) is five under Euclidean geometry. A movement restricted to horizontal and vertical segments has shortest total coordinate travel seven, a different model. The maximum length measures the largest individual coordinate change. Each answers a different operational question rather than correcting the other two.

Unit balls for three normsThe maximum-norm boundary is a square, Euclidean a circle, Manhattan a diamond. All meet the coordinate unit-axis endpoints.Two-dimensional unit balls: each boundary has its norm one01−1Manhattan: diamondEuclidean: circleMaximum: square
Figure 9.3

In two coordinates the unit balls are a diamond, a circle, and a square for the Manhattan, Euclidean, and maximum norms. They contain different directions at different displayed radii.

A metric d(p,q) is nonnegative, zero exactly when p=q, symmetric, and satisfies d(p,r)≤d(p,q)+d(q,r). A norm induces a metric d(p,q)=||q−p||. Norm positivity gives the first two conditions, homogeneity at −1 gives symmetry, and the norm triangle law applied to (q−p)+(r−q) gives the metric triangle law. The mathematical object is now a distance between points, rather than a length of one displacement vector.

Consider a query (0,0) and candidates P=(1,100),Q=(3,0), with coordinates measured in hours and centimetres. An unweighted Euclidean formula gives distances approximately 100.005 and three, selecting Q. Converting centimetres to metres changes P to (1,1), giving √2 while Q remains three, selecting P. The physical measurements have not changed; leaving the numeric metric weights fixed changed the chosen geometry.

An explicit scale vector s with positive entries defines dₛ(p,q)=sqrt(∑ᵢsᵢ²(qᵢ−pᵢ)²). This is a metric because it applies an invertible coordinate scaling then Euclidean distance. If you convert units and want to preserve an existing metric, change its scales inversely to the coordinate conversion. If a scale is zero, differences on that coordinate can be ignored even for distinct points, so the identity-of-indiscernibles metric condition can fail.

For fixed finite dimension d, the standard norms satisfy ||u||∞≤||u||₂≤||u||₁≤√d ||u||₂. The first inequalities compare a largest squared entry, the sum of squares, and the square of the absolute-entry sum. The final inequality is Cauchy–Schwarz applied to (|uᵢ|) and the all-ones vector. These bounds compare lengths, but do not make neighbour rankings identical, and their constants depend on dimension.

Choosing a norm or feature scale therefore requires a modelling reason. Similar record lengths, local paths in a grid, and maximum permitted component error can favour different measures. A numerical result is meaningful only relative to that choice and the feature semantics. Later statistical preprocessing will introduce learned scale estimates and training-only fitting; this module uses declared unit conversions so that no unlearned statistics is required.

Check your understanding

Compute the three norms of (−3,4). Why can centimetre-to-metre conversion change an unweighted neighbour, and what happens if a metric’s scale is zero on one coordinate?

Show answer

The norms are seven, five, and four. Fixed numeric weights give the second coordinate a different contribution after conversion; preserving a chosen metric requires adjusting its weights consistently. A zero scale may assign zero distance to distinct points differing only in that coordinate, violating the metric condition.

6

Unit directions normalisation and cosine similarity

For a nonzero vector u, normalisation divides by its Euclidean length: û=u/||u||₂. Norm homogeneity gives ||û||₂=1, and the divisor is positive, so the direction is preserved. For (3,4) the unit direction is (0.6,0.8). Dividing by the length is not the same as dividing by the sum of entries; negative coordinates and cancellation make that sum an unsuitable general length.

The zero vector cannot be normalised by this formula because its length is zero and it has no direction. An application must declare its handling: reject the operation, report an undefined score, or use a separately justified convention. Silently adding a small constant to the denominator changes the function being computed; it should be explained as a policy rather than presented as the original geometric definition.

For two nonzero vectors, cosine similarity is cosim⁡(u,v)=(u⋅v)/(∥u∥2∥v∥2)=u^⋅v^\operatorname{cosim}(u,v)=(u\cdot v)/(\|u\|_2\|v\|_2)=\hat u\cdot\hat v. It measures alignment while ignoring positive magnitude changes: replacing u with αu for α>0 leaves the score unchanged. A negative scale reverses direction and changes its sign. A score of one means parallel positive directions, not that the original vectors have equal coordinates or lengths.

Worked example
Alignment is not a probability or an equality test

Vectors (1,0) and (100,0) have cosine one and Euclidean distance ninety-nine. Vectors (1,0) and (−1,0) have cosine minus one, which cannot be an ordinary probability. Orthogonal nonzero vectors have cosine zero; that does not mean a missing vector or statistical independence. The result is a geometrical score under the chosen coordinates, not a calibrated confidence in a label.

For unit vectors u,v, squared Euclidean distance is ||u−v||₂²=2−2u·v=2−2cosim(u,v). Therefore among unit candidates compared with a unit query, maximising cosine and minimising Euclidean distance give the same ranking, including ties. The derivation uses both unit lengths. On raw vectors, varying magnitudes add other terms and the ranking equivalence need not hold.

Normalisation can discard information an application needs. If magnitude encodes the amount of activity or a physical displacement length, replacing each vector with its direction loses that quantity. If only direction is intended, the change may be appropriate. The important decision is which variation should count as similar; neither cosine nor Euclidean distance is universally the right choice for every embedding or measurement system.

Interactive

Drag the endpoints or use their coordinate inputs and arrow-key controls. Compare dot product, lengths, and angle, including a zero vector. The nonzero and orthogonal examples above remain usable without scripting.

Check your understanding

Normalise (3,4). What is cosim(u,−u) for nonzero u? When does cosine ranking equal Euclidean ranking?

Show answer

The unit direction is (0.6,0.8). Opposite nonzero directions have cosine −1. The ranking equivalence holds when the compared query and candidates are unit vectors; it does not follow for arbitrary raw lengths.

7

Array shapes broadcasting and explicit implementation contracts

A NumPy array has a shape tuple, a number of axes, and a dtype. Mathematical coordinates need a declared storage convention. We use flat arrays (d,) for standalone vectors, rows (1,d) and columns (d,1) when an operation requires those orientations, and tables (B,d) for B observations with d features. Module 10 develops matrix products; here identifying which axis represents features already prevents many faults.

For flat compatible vectors, u*v produces d entrywise products, while u @ v produces their summed dot product. Multiplying a (B,d) table by a length-d scale vector applies each feature scale to each row by broadcasting. np.linalg.norm(table, axis=1) returns B row lengths; omitting the intended axis can collapse an entire table to one aggregate quantity. The chosen axis is part of the mathematical question.

Broadcasting compares axis lengths from the right. Equal lengths match; length one can expand; a missing leading axis acts as one. Consequently (3,1) combined with (3,) produces (3,3). This is legal array arithmetic, but it pairs each column entry with every flat entry. A successful call cannot tell whether those pairwise combinations were intended. The correct output shape must be specified before running it.

Worked example
A wrong broadcast changes the number of comparisons

Predictions [1,2,3] have shape (3,), observations [1,2,4] have shape (3,1). Direct subtraction gives rows [0,1,2], [−1,0,1], [−3,−2,−1]. Their nine squared entries average to 21/9=7/3. Paired differences should be [0,0,−1] in a (3,1) column, whose three squared entries average to 1/3. The repaired subtraction makes both operands columns; the mathematical pairing, not merely the display, changes.

Broadcasting changes the pairing problemSubtracting flat [1,2,3] and column [1,2,4] generates nine differences; paired columns yield only [0,0,−1]. Mean squares are 7/3 and 1/3 respectively.Predictions (3,) − observations (3,1): all pairwise differencesWrong broadcast (3,3)Repaired pairs (3,1)0120-1010-3-2-1-1Nine squared entries: mean 7/3Three squared entries: mean 1/3
Figure 9.4

The incompatible intended orientations broadcast to a full grid of pairwise differences. Matching columns retain one comparison per observation.

Transposing a flat (d,) array does not make a column; it still has one axis. Use an explicit new axis, such as v[:,None], or a reshape whose entry count is checked. Turning (d,) into (d,1) changes the operation contract without changing entry values. Flattening an arbitrary table can lose the distinction between observation and feature axes, so use it only when the intended mathematical object is genuinely that one-dimensional sequence.

Arrays with more than two axes appear in images, batches, time sequences, and AI models. A shape such as (B,T,d) might mean batch, time, features; another application might assign different meanings. The shape tuple itself does not name those semantics. State the axis order, identify reductions and permitted broadcasts, and check output shape at the boundary of each operation. Tensor algebra and automatic differentiation later build on these explicit conventions.

Numerical checks should combine structure and values. Require matching expected shapes before np.allclose, since comparisons can broadcast too. Choose a tolerance appropriate to the small floating computations rather than interpreting approximate equality as an exact proof. Our loop-versus-NumPy lab compares norms at relative and absolute tolerance 10⁻¹² for small coordinates; the algebraic proofs establish the general formulas independently of that finite numerical agreement.

Check your understanding

Predict (4,2)*(2,) and (3,1)-(3,). Why is v.T insufficient for a column, and why check shapes before approximate equality?

Show answer

The first broadcasts to (4,2) with featurewise multiplication. The second broadcasts to (3,3), potentially the wrong pairing. A flat transpose keeps (d,). Approximate comparison can also broadcast and conceal a structural mismatch, so assert the intended shapes first.

8

Common misconceptions and failure cases

Claim Why it fails Repair
A numeric tuple automatically has physical distance Feature units and encodings may be incompatible State the geometry and units
A point and a displacement are interchangeable Origin shifts affect point coordinates Use differences and affine combinations appropriately
Dot product is entrywise multiplication It sums the products to a scalar Check the operator and output shape
A zero vector has a cosine direction Its length is zero Declare an undefined-case policy
Cosine one means equal vectors or certainty Positive multiples share alignment; scores can be negative Interpret the score’s actual definition
Norm choice cannot change neighbours Different balls and scales change comparisons Choose and justify a metric
A flat transpose makes a column It retains one array axis Add an explicit axis
No exception proves correct pairing Broadcasting can implement a different problem Specify inputs, axes, and output shape
9

Three CPU labs

Complete the NumPy preparation and install the pinned numerical requirements. These scripts use small CPU arrays, with no external data. Predict first and explain the captured output; the downloaded files are the executed sources.

Lab A Loop calculations and NumPy checks

Predict: orthogonal dot product, three norms, displacement length, and the unit direction of (3,4). Run: compare loop formulas with NumPy. Explain: the coordinate unit-axis component and the length guard before zip. Change: use a negative coordinate, keep dimensions equal, and predict which norm values remain unchanged. The comparisons use explicit numerical tolerances; the lesson’s proofs are separate from these finite checks.

Download lab1_vectors.py

"""Loop-based vector calculations checked against NumPy on small real coordinates."""
from math import sqrt
import numpy as np

def dot(u, v):
    if len(u) != len(v):
        raise ValueError("Coordinate lengths must agree")
    return sum(a * b for a, b in zip(u, v))

def norm(u, kind):
    if kind == 1:
        return sum(abs(x) for x in u)
    if kind == 2:
        return sqrt(dot(u, u))
    if kind == "inf":
        return max((abs(x) for x in u), default=0)
    raise ValueError("This lab supports norms 1, 2, inf")

u, v = (3, 4), (4, -3)
print("u, v:", u, v)
print("Loop / NumPy dot:", dot(u, v), float(np.array(u) @ np.array(v)))
for kind in (1, 2, "inf"):
    reference = np.linalg.norm(np.array(u, dtype=float), ord=np.inf if kind == "inf" else kind)
    assert np.isclose(norm(u, kind), reference, rtol=1e-12, atol=1e-12)
    print("Norm", kind, ":", norm(u, kind), "; NumPy:", float(reference))
p, q = (1, 2), (4, 6)
displacement = tuple(b - a for a, b in zip(p, q))
print("Displacement q-p / Euclidean distance:", displacement, norm(displacement, 2))
unit = tuple(x / norm(u, 2) for x in u)
assert np.isclose(norm(unit, 2), 1, rtol=1e-12, atol=1e-12)
print("Unit direction of u:", unit)
direction = (1, 0)
projection = tuple(dot((2, 3), direction) * x for x in direction)
print("Component of (2,3) along the first unit axis:", projection)
try:
    dot([1, 2], [3])
except ValueError as error:
    print("Rejected:", error)
print("Tolerances compare numerical values; the length check prevents silent zip truncation.")
Output
u, v: (3, 4) (4, -3)
Loop / NumPy dot: 0 0.0
Norm 1 : 7 ; NumPy: 7.0
Norm 2 : 5.0 ; NumPy: 5.0
Norm inf : 4 ; NumPy: 4.0
Displacement q-p / Euclidean distance: (3, 4) 5.0
Unit direction of u: (0.6, 0.8)
Component of (2,3) along the first unit axis: (2, 0)
Rejected: Coordinate lengths must agree
Tolerances compare numerical values; the length check prevents silent zip truncation.

Lab B Units and neighbour rankings

Predict: which candidate is closest to (0,0) before and after centimetres become metres. Run: read Euclidean and Manhattan orders. Explain: the changed geometry and the stable tie order in the scaled Manhattan result. Change: adjust the metric weights inversely to the unit conversion and verify the original distances return. Stability of sorting makes ties reproducible, not mathematically distinct distances.

Download lab2_feature_units.py

"""Neighbour rankings under an explicitly changed coordinate scale."""
import numpy as np

labels = np.array(["P", "Q", "R"])
# Feature 1 measured in hours; feature 2 measured in centimetres.
query = np.array([0.0, 0.0])
candidates = np.array([[1.0, 100.0], [3.0, 0.0], [0.0, 200.0]])

for name, scale in [("hours + centimetres, unweighted", np.array([1.0, 1.0])),
                    ("hours + metres, unweighted", np.array([1.0, 0.01]))]:
    differences = (candidates - query) * scale
    distances = np.linalg.norm(differences, axis=1)
    order = np.argsort(distances, kind="stable")
    print(name)
    print("Euclidean distances:", [round(float(d), 6) for d in distances])
    print("Rank order:", labels[order].tolist())
    l1 = np.abs(differences).sum(axis=1)
    print("Manhattan distances:", l1.tolist(), "; stable tie order:", labels[np.argsort(l1, kind="stable")].tolist())
assert labels[np.argmin(np.linalg.norm(candidates - query, axis=1))] == "Q"
assert labels[np.argmin(np.linalg.norm((candidates - query) * [1, 0.01], axis=1))] == "P"
print("Changing units while leaving numeric weights fixed changes the metric and can change neighbours.")
print("To preserve a chosen metric after conversion, transform its weights consistently.")
Output
hours + centimetres, unweighted
Euclidean distances: [100.005, 3.0, 200.0]
Rank order: ['Q', 'P', 'R']
Manhattan distances: [101.0, 3.0, 200.0] ; stable tie order: ['Q', 'P', 'R']
hours + metres, unweighted
Euclidean distances: [1.414214, 3.0, 2.0]
Rank order: ['P', 'R', 'Q']
Manhattan distances: [2.0, 3.0, 2.0] ; stable tie order: ['P', 'R', 'Q']
Changing units while leaving numeric weights fixed changes the metric and can change neighbours.
To preserve a chosen metric after conversion, transform its weights consistently.

Lab C Wrong pairing and undefined cosine

Predict: the wrong (3,3) difference grid, the repaired (3,1) discrepancies, and the zero-vector result. Run: compare the two mean squared differences and rejected cases. Explain: why NumPy can produce nan for unchecked zero division and why the guarded function reports undefined instead. Change: compare unit nonzero vectors and verify distance²=2−2cosim. The code is scoped to small finite coordinates; extreme floating magnitudes require further numerical handling.

Download lab3_shape_and_cosine_faults.py

"""Repair a successful wrong broadcast and define the zero-vector policy explicitly."""
import numpy as np

predicted = np.array([1.0, 2.0, 3.0])
observed = np.array([[1.0], [2.0], [4.0]])
wrong = predicted - observed
paired = predicted[:, None] - observed
assert wrong.shape == (3, 3) and paired.shape == (3, 1)
print("Wrong broadcast shape / mean squared difference:", wrong.shape, round(float(np.mean(wrong**2)), 6))
print("Paired shape / mean squared difference:", paired.shape, round(float(np.mean(paired**2)), 6))
assert np.isclose(np.mean(paired**2), 1 / 3, rtol=1e-12, atol=1e-12)

def cosine(u, v):
    u, v = np.asarray(u, dtype=float), np.asarray(v, dtype=float)
    if u.ndim != 1 or v.ndim != 1 or u.shape != v.shape:
        raise ValueError("Require matching flat vector shapes")
    length_u, length_v = np.linalg.norm(u), np.linalg.norm(v)
    if length_u == 0 or length_v == 0:
        raise ValueError("Cosine is undefined for a zero vector")
    return float((u @ v) / (length_u * length_v))

zero, other = np.array([0.0, 0.0]), np.array([1.0, 2.0])
with np.errstate(invalid="ignore", divide="ignore"):
    bad = (zero @ other) / (np.linalg.norm(zero) * np.linalg.norm(other))
print("Unchecked zero-vector cosine:", float(bad))
assert np.isnan(bad)
for u, v in [([3, 4], [4, -3]), ([1, 0], [-1, 0]), ([0, 0], [1, 2]), ([1, 2], [1])]:
    try:
        print("Cosine:", u, v, "=", cosine(u, v))
    except ValueError as error:
        print("Rejected:", u, v, str(error))
print("Scope: small finite coordinates. Extreme floating magnitudes need further numerical handling.")
Output
Wrong broadcast shape / mean squared difference: (3, 3) 2.333333
Paired shape / mean squared difference: (3, 1) 0.333333
Unchecked zero-vector cosine: nan
Cosine: [3, 4] [4, -3] = 0.0
Cosine: [1, 0] [-1, 0] = -1.0
Rejected: [0, 0] [1, 2] Cosine is undefined for a zero vector
Rejected: [1, 2] [1] Require matching flat vector shapes
Scope: small finite coordinates. Extreme floating magnitudes need further numerical handling.
10

Exercises with full solutions

Exercises 1–12 are required; 13–14 extend the geometry. Always annotate feature meaning and array shape when using a data interpretation.

Exercise 1★★★calculation5 min

For u=(3,4),v=(4,−3), compute u+v,2u−v, and −u.

Show solution

Coordinatewise results are (7,1),(2,11),(−3,−4). Each retains two coordinates. Negative scaling reverses the displacement direction.

Exercise 2★★★calculation5 min

For p=(1,2),q=(4,6), compute (1/4)p+(3/4)q and identify whether it is affine and convex.

Show solution

The point is (3.25,5). Coefficients sum to one, so it is affine; both are nonnegative, so it is convex and lies on the segment. Its displacement from p is (2.25,3).

Exercise 3★★★calculation5 min

Find q−p for p=(−1,3),q=(2,−1), then add offset (10,5) to both and recompute.

Show solution

The displacement is (3,−4). Shifted points (9,8),(12,4) give the same difference. Position coordinates change while a common translation cancels in the displacement.

Exercise 4★★★conceptual5 min

Explain why 2p+3q is not an origin-independent affine point construction. Does p+tu define a line when u=0?

Show solution

A common offset b changes 2p+3q by 5b, not b, because coefficients sum to five. The result lacks the affine point transformation rule. For u=0, p+tu is always p, a single point rather than a line.

Exercise 5★★★proof14 min

Prove u·(αv+βw)=α(u·v)+β(u·w) for equal-dimensional real vectors, and prove u·u=0 implies u=0.

Show solution

Expand the finite sum ∑uᵢ(αvᵢ+βwᵢ), distribute in each coordinate, and split into α∑uᵢvᵢ+β∑uᵢwᵢ. A self-product is ∑uᵢ²; all terms are nonnegative, so a zero sum forces each term zero and each coordinate zero. The result uses matching dimensions and real coordinates.

Exercise 6★★★proof14 min

Prove Cauchy–Schwarz using ||u−tv||₂². Include v=0 and the equality condition for both nonzero vectors.

Show solution

If v=0 both sides vanish. Otherwise choose t=(u·v)/||v||₂². Expansion gives 0≤||u||₂²−(u·v)²/||v||₂², hence (u·v)²≤||u||₂²||v||₂². Nonnegative square roots give the absolute-value inequality. With nonzero vectors equality requires ||u−tv||₂²=0, so u=tv; scalar dependence also directly gives equality.

Exercise 7★★★proof14 min

Derive the Euclidean norm’s triangle inequality from Cauchy–Schwarz, then explain why d(p,q)=||q−p||₂ is symmetric and satisfies a metric triangle law.

Show solution

Expand ||u+v||₂² and bound 2u·v by 2||u||₂||v||₂. This gives at most (||u||₂+||v||₂)², and nonnegative square roots give the norm law. Symmetry follows because ||−x||₂=||x||₂. Apply the norm triangle to r−p=(q−p)+(r−q) to obtain the metric law. Positivity and equality only at zero complete the metric conditions.

Exercise 8★★★application10 min

For query (0,0),P=(1,100),Q=(3,0), compare Euclidean neighbours before and after the second coordinate is multiplied by 0.01. How would you preserve the original metric?

Show solution

Initially distances √10001≈100.005 and three select Q. After scaling, P=(1,1) has √2 while Q remains three, selecting P. To preserve the original numerical metric, scale the new second coordinate by weight one hundred inside the norm, undoing the coordinate conversion. Choosing unweighted metres instead deliberately changes the geometry.

Exercise 9★★★application10 min

For a (4,2) table of four two-feature records, identify the shape of feature scaling by (2,) and row norms with axis=1. Explain why array ndim is not the feature dimension.

Show solution

Scaling broadcasts to (4,2), applying each of the two scale values to its feature column. Row norms reduce the feature axis and yield (4,). The table’s ndim is two because it has observation and feature axes; each record has two coordinates. These counts happen to agree here but refer to different things, as a (4,3) table shows.

Exercise 10★★★application10 min

Prove for unit u,v that ||u−v||₂²=2−2cosim(u,v). Explain the ranking equivalence and why cosine one is not a raw-vector equality test.

Show solution

Expand to ||u||₂²+||v||₂²−2u·v=2−2u·v. Unit lengths make u·v=cosim. Minimising the distance or its square therefore maximises cosine among unit candidates. Raw (1,0),(100,0) have cosine one but differ greatly; the equality test needs matching coordinates, not only alignment.

Exercise 11★★★diagnosis10 min

Predictions (3,) and observations (3,1) are subtracted without reshaping. Find the resulting shape and explain the displayed mean squared differences for [1,2,3] and [1,2,4].

Show solution

Right-aligned shapes broadcast to (3,3), producing rows [0,1,2],[−1,0,1],[−3,−2,−1]. Squared sum 21 over nine entries gives 7/3. Paired column differences [0,0,−1] give 1/3 over three entries. Make both operands (3,1) or both (3,) for the intended pairing and verify shape before comparing values.

Exercise 12★★★diagnosis10 min

A similarity function normalises every vector, calls cosine a probability, and treats a flat transpose as a column. Diagnose all three assumptions with concrete cases.

Show solution

Zero has zero norm and cannot be normalised by division, so specify an undefined-case policy. Opposite vectors have cosine −1, disproving probability interpretation; even positive scores are not automatically calibrated probabilities. A flat (d,) transpose remains (d,); use an explicit new axis to obtain (d,1) when required.

Exercise 13★★★proof15 min

Extension: prove ||u||∞≤||u||₂≤||u||₁≤√d ||u||₂ for real d-coordinate vectors. State why this does not imply equal nearest-neighbour rankings.

Show solution

A largest squared entry is at most the sum of squares, giving the first bound after square roots. The square of ∑|uᵢ| equals ∑uᵢ² plus nonnegative cross terms, giving the second. Cauchy–Schwarz between (|uᵢ|) and d ones gives the final bound. Bounds compare magnitudes but do not preserve every ordering: from the origin in two dimensions, (3,0) and (2,2) have Manhattan distances three,four but Euclidean distances three,√8, reversing their order.

Exercise 14★★★application20 min

Extension: a scale vector s=(1,0) defines dₛ(p,q)=||s*(q−p)||₂. Show the failed metric condition, and explain why positive coordinate scales avoid that failure.

Show solution

Points p=(0,0),q=(0,5) differ, but their scaled displacement is zero and the proposed distance is zero. Identity of indiscernibles fails. With every scale positive, the diagonal scaling is injective, so a zero scaled displacement implies every original difference is zero. The Euclidean positivity, symmetry, and triangle laws then transfer, yielding a genuine metric.

11

Self-check quiz

Questions 1–9 are scored automatically. Question 10 is written and self-reviewed.

1
What stays unchanged when the same offset is added to two points?
2
Which coefficients define a convex combination?
3
What is (3,4)·(4,−3)?
4
Which norm lengths does (3,4) have in order 1,2,∞?
5
When is the ordinary cosine quotient defined?
6
What does cosine one establish?
7
What can changing one feature’s units do to an unweighted distance?
8
What is the broadcast shape of (3,1)+(3,)?
9
Why check shapes before np.allclose?
Show answer

Unweighted distances from (0,0) to P=(1,100),Q=(3,0) choose Q. Scaling the second coordinate by 0.01 makes P distance √2 and Q distance three, choosing P because the numeric geometry changed. Preserve an existing metric by adjusting its feature weight inversely. Subtracting (3,) predictions and (3,1) observations produces (3,3) pairwise comparisons instead of three paired differences. Make both columns or both flat and check the output shape. Arithmetic can be correct for an unintended question, so the answer must specify units, axis semantics, and the intended comparison.

12

Reading with a purpose

Use analytic-geometry selections from Mathematics for Machine Learning, the authors’ book companion, to compare dot products, lengths, and projections. Consult NumPy’s basic array guide and broadcasting rules for the storage operations used by the labs. The lesson’s proofs and examples are original; numerical checks do not replace the general arguments.

When Selection and question
Session 1 · 20 minutes Coordinate vectors and geometry: which objects are points, and which are displacements?
Session 4 · 20 minutes NumPy shapes and broadcasting: how do right-aligned lengths determine the output axes?

Separate coordinate geometry from programming terminology such as array ndim. A word used by both subjects may refer to different counts.

13

Retrieval exit task and next step

Without notes, compute a vector combination, prove Cauchy–Schwarz, derive the Euclidean triangle law, and explain the zero-vector exception. Give a metric-changing unit conversion and a wrong-shaped operation that runs without an exception.

Exit task: use u=(−3,4),v=(3,4). Compute u·v=7, both Euclidean lengths five, cosine 7/25, and unit directions (−0.6,0.8),(0.6,0.8). A two-row table of these vectors has shape (2,2); row norms yield (2,). Explain why a three-row two-feature table would still contain two-coordinate vectors, not three-coordinate ones.

Ready to move on: you can state coordinate and shape assumptions before applying a formula. The next linear-algebra lesson treats matrices as maps and solves linear systems. See the course overview for availability.

14

Notation and bilingual terminology

Term or notation Meaning 中文
Scalar / vector / point Scale number / displacement or tuple / location 标量、向量、点
Linear / affine / convex combination Arbitrary coefficients / sum one / additionally nonnegative 线性、仿射、凸组合
u·v / orthogonal Sum of coordinate products / zero dot product 点积、正交
Cauchy–Schwarz Bound on a dot product by two lengths 柯西–施瓦茨不等式
Norm / metric Vector length / point distance satisfying laws 范数、度量
Unit direction / cosine similarity Normalised nonzero vector / directional alignment 单位方向、余弦相似度
Shape / axis / broadcasting Storage lengths / coordinate index dimension / compatible expansion 形状、轴、广播
Feature scale / pairing Declared geometry weights / corresponding-record comparison 特征尺度、配对