ROBOTECA FREE SAMPLE
Dashboard

A free lesson from Robotics & ROS 2: the whole module, nothing cut short.

LESSON · Autonomous Mobile Robotics

Rotations & homogeneous transformations

Turn 1 40 min LESSON

ALearning Material

To convert a position from one frame to another, you need the math of frames: how a frame can be rotated and translated relative to another, and how to apply that to a point's coordinates. Rotation handles ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍orientation (the frames point different ways), translation handles position (the origins are in different places), and together they form a rigid transformation. The operation that takes coordinates in one frame to coordinates in another. The elegant tool that packages rotation and translation into a single, composable operation is the homogeneous transformation matrix: the workhorse of robotics geometry. This lesson is the how behind the frames of the previous lesson.

Start with rotation. A frame can be turned relative to another (its axes point different ways), described by an angle (in 2D) or a rotation matrix (a compact, general description that works in 2D and 3D). Multiplying a point's coordinates by a rotation matrix gives its coordinates after the rotation, and composing rotations is just multiplying their matrices. Now add translation: the frames' origins are also in different places, so you shift by adding an offset vector. A full ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍rigid transformation is 'rotate, then translate', and because doing that as 'matrix-multiply then vector-add' is awkward to chain, robotics uses a clever trick: the homogeneous transformation matrix, a single (3x3 in 2D, 4x4 in 3D) matrix that packs both rotation and translation together, so that applying a transform is one matrix-vector multiply and composing transforms (sensor -> robot -> world) is one matrix-matrix multiply.

Rotation + translation = a rigid transform; homogeneous matrices make it composable:

  • Rotation: the frames point different ways, so a rotation matrix rotates a point's coordinates when you multiply by it. Compose rotations by multiplying matrices, ; order matters, because rotations don't commute.
  • Translation: the frames' origins differ, so add an offset vector to shift the point.
  • Rigid transform: rotate, then translate: ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍.
  • Homogeneous transform matrix (the trick): pack and into one matrix, 4x4 in 3D:

Applying a transform is then one matrix-vector multiply, using , and composing transforms (sensor to robot to world) is one matrix-matrix multiply, . That makes it the workhorse: one object represents a frame's full pose (position and orientation) and chains by multiplication.

The disciplines. Rotation (orientation: the frames point different ways) is captured by a rotation matrix R (multiply a point's coordinates by R to rotate them; compose rotations by multiplying matrices, and order matters: rotations don't commute). Translation (position: the origins differ) is an added offset t‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍. A rigid transformation is 'rotate then translate' (p' = Rp + t). The homogeneous transformation matrix T packs both R and t into a single matrix (4x4 in 3D, using coordinates written as [x, y, z, 1]), so that applying a transform is one matrix-vector multiply (p' = Tp) and composing transforms is one matrix-matrix multiply (): which is why one T can represent a frame's full pose (position + orientation) and chains of frames (sensor -> robot -> world) collapse to multiplying their matrices. The habits: use rotation matrices for orientation, respect that rotation order matters, and use homogeneous matrices to combine and compose transforms. This is the math that powers the transform tree.

Math you need here. A rotation or transform is a matrix-as-machine: its columns are where the base axes land, and chaining frames is matrix multiplication (order matters). See the linear-algebra foundation lesson (this topic, Turn 1).

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Formulas & method. The pointer above says a transform is a matrix-as-machine; these are the matrices, and the first three lines are the whole of section D:

Elementary rotations, with and , the angle in radians:

is the odd one out: its minus sits bottom-left, not top-right. Getting that wrong builds a rotation that still looks like a rotation and points the wrong way.

Is it a rotation at all: (orthonormal) and (+1, not -1, which would be a reflection).

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Rigid transform: , rotate then translate. is a different point: translating first rotates the offset too.

Homogeneous 4x4:

The subscripts must meet at every join ( gives ). Read off for free: the top of the last column (T[:3, 3] in NumPy) is the child frame's origin expressed in the parent. To invert:

not : the offset must be rotated into the new frame.

Euler angles to one matrix:

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍This is the usual robotics order, and one of twelve valid conventions. Angles never add except about a shared axis: , exactly.

How far apart two rotations are, as one number for a whole orientation error:

and between two unit vectors, .

Method for a chain. Write one per joint with its subscripts, multiply left to right so the subscripts cancel, then check the translation column against what you expect physically before trusting anything downstream. A camera 0.6 m up a mast should read about 0.6 m up in T[:3, 3]; if it does not, the error is in the chain and not in the target.

Why it exists. Converting coordinates between frames requires accounting for both how the frames are oriented (rotation) and where their origins are (translation), and chaining many such conversions (sensor to robot to world) must be efficient and reliable. Rotation matrices capture orientation and compose by multiplication; homogeneous transformation matrices fold rotation and translation into one object that both applies to points and composes with other transforms by a single multiplication: giving robotics a clean, uniform, composable representation of pose. This is precisely the machinery the frames of the previous lesson need, and it underlies the transform tree, kinematics, and every geometric computation in the topic.

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Mental model. A homogeneous transform is like a single instruction that says 'here is how this frame sits relative to that one: turned this way and shifted to here': one object holding both the facing (rotation) and the location (translation) of a frame. Applying it to a point is like re-describing that point from the other frame's point of view. And composing them is like following a chain of such instructions: 'the camera sits like this on the robot, and the robot sits like that in the world' multiply together into 'the camera sits like this-and-that in the world', one combined instruction. The matrix is just the bookkeeping that makes 'turned and shifted' chain cleanly by multiplication.

Common misunderstandings.

  • "Rotation is just adding angles." In 2D a single rotation is an angle, but general rotations are matrices (and in 3D you can't just add angles, orientation is 3D), and rotations don't commute: rotating then rotating again depends on the order (R2R1 != R1‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍R2 in general): 'turn left then up' differs from 'up then left'. Compose by multiplying matrices in the right order, not adding.
  • "Translation and rotation can be applied in any order." A rigid transform is rotate then translate (p' = Rp + t); swapping them gives a different result. The homogeneous matrix encodes the correct* combined operation, so use it rather than ad-hoc rotate/translate steps that are easy to get backwards.
  • "Homogeneous coordinates / the extra '1' are a needless complication." Appending a 1 (writing a point as [x, y, z, 1]) is the trick that lets a single matrix do both rotation and translation and lets transforms compose by multiplication. It's not clutter: it's exactly what makes pose one composable object, which is why all of robotics uses it.

Connections. This is the math behind the previous lesson's frames and the transform tree (map -> odom -> base_link -> sensor, the Nav2 SLAM lesson): each edge is a homogeneous transform, and tf2 composes them by multiplication; ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍homogeneous transforms reappear as the tool of forward kinematics (chaining joint transforms down a manipulator: Turn 2) and of converting sensor measurements into the map (the perception/costmaps lesson); rotation representations deepen in Turn 2 (rotation-matrix properties, Euler angles, quaternions); and Python's transforms/tf work is the same math in code. Rotation + translation, packaged as homogeneous matrices, is the geometric engine of the whole topic.

BImmediate Active Recall

QUERY

How is rotation represented, and what does it mean that rotations don't commute?

REVEAL
ANSWER

Rotation (a frame's orientation, which way its axes point) is represented by an angle (in 2D) or, generally, a rotation matrix R. Multiplying a point's coordinates by R gives its coordinates after the rotation, and this works in 2D and 3D. ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Rotations don't commute means the order of rotations matters: applying rotation R1 then R2 (R2R1) generally gives a different result from R2 then R1 (R1R2). Physically, 'turn left then tilt up' lands you facing somewhere different from 'tilt up then turn left'. So you compose rotations by multiplying their matrices in the correct order. You cannot just add angles (especially in 3D).

Did you recall it?
QUERY

What is a rigid transformation, and what are its two parts?

REVEAL
ANSWER

A rigid transformation converts a point's coordinates from one frame to another by accounting for the two ways the frames differ: ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍rotation (the frames are oriented differently: their axes point different ways, handled by a rotation matrix R) and translation (the frames' origins are in different places, handled by adding an offset vector t). The full operation is 'rotate then translate': p' = R*p + t. It's called rigid because it preserves distances and shapes (it only repositions and reorients, no stretching). Rotation handles orientation, translation handles position, together they relate any two frames.

Did you recall it?
QUERY

What is a homogeneous transformation matrix, and what two things does it make easy?

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍REVEAL
ANSWER

A homogeneous transformation matrix T packs both the rotation R and the translation t into a single matrix (4x4 in 3D), using points written with an extra 1 ([x, y, z, 1]). It makes two things easy: (1) applying a transform is one matrix-vector multiply (p' = Tp). It does the rotate-then-translate in one step; and (2) composing transforms is one matrix-matrix multiply (), so a chain of frames (sensor -> robot -> world) collapses to multiplying their matrices. That's why one T represents a frame's full pose (position + orientation) and chains of frames compose by multiplication*, the workhorse representation of robotics geometry.

Did you recall it?
‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍QUERY

Why is appending a '1' (using homogeneous coordinates) the key trick?

REVEAL
ANSWER

Because a plain rotation matrix can only rotate (multiply). It can't translate (add an offset) within a single matrix multiply. By writing a point with an extra coordinate of 1 ([x, y, z, 1]) and using a 4x4 matrix with R in the top-left and t in the last column, the single matrix multiply now does both Rp and + t at once (the 1 'pulls in' the translation column). This is what lets one matrix represent rotation and translation together, and (because it's now just matrix multiplication) lets transforms compose by multiplying (T2 * T1). So the extra 1 isn't clutter; it's exactly what turns 'rotate then translate' into one composable matrix operation*, making pose a single chainable object.

Did you recall it?

CConceptual Questions

Answer each in your own words in the box, then reveal the model answer to compare. These ask why, not how, and your answers are saved.

PROMPT

Why do rotation matrices and the fact that rotations don't commute matter so much in robotics, and what goes wrong if you treat orientation as if you could just add angles?

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍REVEAL MODEL ANSWER
MODEL ANSWER

Rotation matrices and the non-commutativity of rotations matter so much in robotics because orientation is a genuinely multi-dimensional, order-dependent quantity, and representing or composing it incorrectly produces wrong orientations, which, since everything in robotics is frame-relative, corrupts where the robot thinks things are and which way it should go. First, why rotation matrices: a rotation re-describes directions, and in general (certainly in 3D) you cannot capture 'which way a frame is turned' with a single number. In 2D a rotation is a single angle, but even there, applying it to a point's coordinates is a matrix operation (it mixes the x and y components); in 3D, orientation has three degrees of freedom and rotations about different axes interact, so a rotation must be represented by a rotation matrix (or an equivalent object like a quaternion). A structured thing that correctly transforms coordinates, not just a number you add. The rotation matrix is the honest representation of 'how this frame is oriented relative to that one', and multiplying a point's coordinates by it correctly gives the point's coordinates after the rotation. Second, why non-commutativity matters: rotations ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍do not commute. The order in which you apply them changes the result (R2 then R1 is generally not the same as R1 then R2). This is not a mathematical curiosity; it's physically true and easy to feel. Hold an object, rotate it 90 degrees about the vertical axis, then 90 degrees about a horizontal axis, and note its final orientation; now start over and do the same two rotations in the opposite order. The object ends up facing a different way. Because rotations compose by matrix multiplication, and matrix multiplication is not commutative, the math correctly captures this: R2R1 != R1R2. So in robotics, when you build up an orientation from several rotations (a sensor mounted at an angle on an arm segment that is itself rotated, etc.), you must multiply the rotation matrices in the correct order corresponding to how the rotations are physically applied. What goes wrong if you treat orientation as if you could just 'add angles': you get wrong orientations, and silently. If you naively add angles (especially in 3D, where 'the angles' aren't even a well-defined single thing), or compose rotations in the wrong order, the resulting orientation is simply incorrect. The frame is computed as facing a direction it isn't. And because orientation feeds into every frame transformation (a transform is rotate-then-translate, and a wrong rotation rotates points wrongly), a wrong orientation mis-places everything that transform touches: a sensor measurement gets rotated into the wrong direction before being placed on the map, so the obstacle appears in the wrong spot; a goal direction is computed wrongly, so the robot drives off-course. The errors are geometric and confident, no exception is thrown, the numbers just point the wrong way. This is why robotics insists on proper rotation representations (matrices, and later quaternions) and on respecting composition order: orientation is multi-dimensional and order-dependent, so it ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍must be handled with the right math. Getting it wrong doesn't fail loudly; it quietly rotates the robot's whole picture of the world, which is exactly the kind of bug frames-and-transforms discipline exists to prevent.

Compared to the model answer - did you get it?
PROMPT

Why is the homogeneous transformation matrix such a powerful and elegant representation, and what does packing rotation and translation into one composable object enable across robotics?

REVEAL MODEL ANSWER
‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍MODEL ANSWER

The homogeneous transformation matrix is powerful and elegant because it unifies the two distinct things that relate frames, rotation and translation, into a single object that both applies to points and composes with other transforms by the same operation (matrix multiplication), turning all of spatial geometry into uniform, chainable linear algebra. Consider the problem it solves. Relating two frames requires a rotation (they're oriented differently) and a translation (their origins differ), and the rigid transformation is 'rotate then translate': . Done literally, that's a matrix multiply plus a vector add, two different operations, and chaining several of them (sensor to robot to world) becomes an awkward nested sequence of multiplies and adds that's easy to get wrong and hard to manipulate as a whole. The homogeneous trick fixes this beautifully: by writing points with an extra coordinate of 1 (so a 3D point is [x, y, z, 1]) and building a 4x4 matrix with the rotation R in the top-left block and the translation t in the last column (and a bottom row of [0 0 0 1]), a ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍single matrix-vector multiply Tp now computes both Rp and the + t at once. The appended 1 multiplies the translation column and pulls it in. So rotation and translation are folded into one matrix and one operation. The elegance compounds when you compose: because applying a transform is now just matrix multiplication, composing transforms is also just matrix multiplication. If takes a point from the sensor frame to the robot frame, and takes it from the robot frame to the world frame, then takes it straight from sensor to world. One matrix that is the whole chain, obtained by multiplying the links. This single, uniform rule: 'transforms apply to points and compose with each other by matrix multiplication', is what makes the representation so powerful, and it enables a great deal across robotics. Pose as one object: a single T captures a frame's full pose (orientation and position), so 'where and how is the camera, the gripper, the robot?' is one matrix, not separate rotation and position handled differently. ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍The transform tree: the whole tree of frames (map -> odom -> base_link -> sensor) is just transforms on its edges, and asking 'where is this point in that frame?' is answered by multiplying the transforms along the path. Exactly what ROS 2's tf2 does. Forward kinematics: a manipulator's tip pose is the product of the homogeneous transforms across each joint, so chaining down the arm is repeated matrix multiplication (Turn 2's DH parameters formalise exactly this). Perception: putting a sensor measurement on the map is multiplying it by the sensor-to-map transform. Invertibility: a transform's inverse (world-to-sensor from sensor-to-world) is just the matrix inverse, so you can convert either direction. In short, by making pose a single composable object and all frame conversions a single operation, the homogeneous transformation matrix turns the messy bookkeeping of 'turned and shifted, chained many times' into clean linear algebra, which is why it is the workhorse of robotics geometry and the engine under frames, the transform tree, kinematics, and perception. Its power is precisely that one idea (pack rotation and translation, append a 1) buys a uniform, composable, invertible algebra of pose.

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Compared to the model answer - did you get it?

DPractice Problems

P1 (easy). Build the three elementary 3-D rotation matrices Rx, Ry, Rz as code, then answer with numbers:

apply and to the point (1, 0, 0), give both results and both matrices, and state the distance between the two answers. Then repeat for (0, 1, 0). Finally, write the rigid transform for , and apply it to (1, 0, 0), and say what and each did.

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍P2 (medium). Show with numbers that adding angles is wrong.

Compose a 30 degree yaw with a 40 degree pitch as R_pitch @ R_yaw, and separately build the single matrix for "roll 0, pitch 40, yaw 30" using the same Rz Ry Rx convention. Give both matrices, state whether they are equal, push the vector (1, 0, 0) through each, and compute both the angle between the two output vectors and the total rotation angle between the two matrices. Then state where this bites in practice.

P3 (harder). Build a real chain with 4x4 homogeneous matrices: a base at (3, 2, 0) yawed 30 degrees, a pan mast 0.10 m forward and 0.60 m up on the base panned 45 degrees, and a camera 0.05 m forward, 0.12 m up on the pan head, pitched 20 degrees down.

Produce T_world_cam, transform a target lying 4 m along the camera z-axis into world coordinates, invert the chain to get the target back, and read the camera's world position straight off the matrix. Then state the three properties that make one 4x4 object worth more than keeping R and t separately.

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Solutionsclick to reveal

P1. The three elementary rotations.

import numpy as np

def Rx(deg):
    t = np.radians(deg); c, s = np.cos(t), np.sin(t)
    return np.array([[1, 0, 0], [0, c, -s], [0, s, c]])

def Ry(deg):
    t = np.radians(deg); c, s = np.cos(t), np.sin(t)
    return np.array([[c, 0, s], [0, 1, 0], [-s, 0, c]])

def Rz(deg):
    t = np.radians(deg); c, s = np.cos(t), np.sin(t)
    return np.array([[c, -s, 0], [s, c, 0], [0, 0, 1]])

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍The check that they are rotations at all, worth running once when you write them:

assert np.allclose(R @ R.T, np.eye(3))      # orthonormal
assert np.isclose(np.linalg.det(R), 1.0)    # +1, not -1 (which would be a reflection)

The two orders.

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍The matrices are different, so the results are different.

Point Rz(90) @ Rx(90) Rx(90) @ Rz(90)
(1, 0, 0) (0, 1, 0) (0, 0, 1)
(0, 1, 0) (0, 0, 1) (-1, 0, 0)

Distance between the two answers for (1, 0, 0): on a unit vector, which is 90 degrees apart. Not a rounding difference, not a small discrepancy: two completely different directions from the same two rotations.

A physical check, worth doing with your hand. Point your thumb along +x. Rotate 90 degrees about z, then 90 about x, and note where the thumb ends up; start again and do x first, then z. They do not agree, and this is why the matrices do not either.

The rigid transform.

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍What each part did. changed the point's direction from the x-axis to the y-axis, leaving its distance from the origin unchanged at 1. moved it bodily by (2, 1, 0) with no change of direction at all.

The order in is not arbitrary, and reversing it gives a different answer:

Rotate then translate is the convention because it composes correctly: it is what the homogeneous matrix encodes, and it is what every library, URDF and tf tree means by a pose. Translating first rotates the offset too, which is the arithmetic behind a sensor that ends up in the wrong place when a mount is written back to front.

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍P1Compared to this solution - did you get it right?

P2. The convention, stated first, because half the confusion here is convention.

intrinsic Z-Y-X, the usual robotics order.

The two computations.

Composed, :

"Angles added", :

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍They are not equal. Only four of the nine entries match, and the differences are not small: -0.3830 against -0.5000, 0.6428 against 0.5567.

Through the vector (1, 0, 0):

The angle between them: .

And the total rotation between the two matrices, from :

Twenty degrees of orientation error from a step that felt like arithmetic. On an arm with a 1 m reach that is 35 cm at the end effector; on a camera it is most of the field of view.

Why it happens. The composed version applies the yaw first, so the pitch afterwards is about the already-rotated y-axis. The Euler-angle version applies pitch first and yaw second, about different axes. ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Rotations are not quantities that add, they are operators that compose, and composition depends on order because each rotation moves the axes the next one uses.

Angles do add in exactly one case: rotations about the same axis. Rz(30) @ Rz(40) = Rz(70), exactly. That single case is where the intuition comes from, and it is why the mistake is so natural: it is correct in 2D, where there is only one axis.

Where this bites in practice.

Averaging orientations. Taking the mean of two sets of roll-pitch-yaw values does not give the rotation halfway between them, and can give something wildly wrong near a wrap-around. The correct operation is a quaternion slerp or a rotation-matrix average, neither of which looks like taking a mean.

Interpolating a trajectory. Linearly interpolating Euler angles between two orientations produces a path that wobbles rather than turning smoothly, and near a gimbal-lock configuration it can swing violently through an intermediate orientation that was never requested.

Accumulating a gyro. Integrating body-frame angular rates by adding them to roll, pitch and yaw is wrong for the same reason, and it is a real source of drift. The correct integration composes a small rotation each step: R = R @ expm(skew(omega) dt)‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍, or the quaternion equivalent.

Adding an offset to a calibration. "The camera is mounted 5 degrees further round, so add 5 to the yaw" is only right when that axis is the last in the convention and nothing else is rotated. Compose the correction matrix instead, and the question never arises.

The rule: never add angles, always multiply matrices (or quaternions), and state the convention next to every set of Euler angles you store. The convention is not optional metadata; "roll 0, pitch 40, yaw 30" names a different orientation under Z-Y-X than under X-Y-Z, and there are twelve valid conventions.

P2Compared to this solution - did you get it right?

P3. Building the transforms.

def T(R, t):
    M = np.eye(4)
    M[:3, :3] = R
    M[:3,  3] = t
    return M

T_world_base = T(Rz(30),  [3.00, 2.00, 0.00])
T_base_pan   = T(Rz(45),  [0.10, 0.00, 0.60])     # mast: forward 0.1, up 0.6, panned 45
T_pan_cam    = T(Ry(-20), [0.05, 0.00, 0.12])     # camera on the head, pitched 20 down

T_world_cam = T_world_base @ T_base_pan @ T_pan_cam

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍The subscripts meet at every join: world_BASE x BASE_PAN x PAN_cam, leaving world_cam.

The result.

The target.

Sanity check on the z component. The camera sits 0.72 m up and is pitched 20 degrees down, so a target 4 m along its optical axis should be well above it: . Matches. (The pitch is negative about y, which tips the camera's z-axis upward in this convention; if you expected the target below the camera, that sign is the thing to check, and this arithmetic is how you catch it.)

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍The inverse, and the round trip.

Exactly the camera-frame point we started from.

The camera's world position, read off directly:

No arithmetic required. The last column of a homogeneous transform is the origin of the child frame expressed in the parent, which is the quickest available check on a chain: the camera should be about 0.72 m up and roughly 0.1 m from the base position of (3, 2), and it is.

And the rotation block is still a rotation:

Three rotations composed, and the result is exactly orthonormal, which is the property that lets the chain be extended indefinitely without drift.

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍The three properties that make one 4x4 object worth more than separate R and t.

One operation for both, so composition is associative and cannot be got wrong. With separate parts, chaining is R_total = R2 @ R1 and t_total = R2 @ t1 + t2, and that second expression is where the errors live: forgetting to rotate t1 by R2 is the single most common bug in hand-written transform code, and it produces exactly the "sensor in the wrong place" error of the frames lesson. With matrices it is T2 @ T1, and the rotation of the translation is done by the multiplication itself.

A pose and a transformation become the same object. T_world_cam is simultaneously "where the camera is" and "the operation that converts camera coordinates to world coordinates". That identity is why a tf tree can store one thing per edge and answer both questions, and why a URDF's joint definitions are the robot's kinematics rather than a description of them.

It composes with everything else in the same algebra. Projection matrices, scaling, and perspective all live in the same 4x4 form, so a camera model is K [R|t] and the whole pipeline from a world point to a pixel is one matrix product. ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Keeping R and t apart forces a special case at every one of those joins.

The one thing to watch: numerical drift over very long chains. Repeated multiplication accumulates floating-point error in the rotation block, so after thousands of compositions is no longer quite . The fix is to re-orthonormalise periodically (via SVD, or by normalising the equivalent quaternion), and the check is the assertion above. In a robot's tf tree this rarely matters, because each transform is recomputed from its source rather than accumulated; it matters in dead reckoning, where the composition really is cumulative.

P3Compared to this solution - did you get it right?

EFeynman Exercise

Explain to a beginner, using the idea of a single instruction that says 'here is how this frame sits relative to that one: turned this way and shifted to here': (1) why rotation (which way a frame faces) and translation (where its origin is) are the two parts of relating frames, (2) why rotation order matters (turning then tilting differs from tilting then turning), and (3) why packing both into one homogeneous matrix lets you re-describe a point from another frame's viewpoint and chain 'camera-on-robot' with 'robot-in-world' into 'camera-in-world' by multiplying.

REVEAL MODEL ANSWER
‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍MODEL ANSWER

Rotations and homogeneous transforms are best understood as a single instruction that says 'here is how this frame sits relative to that one, turned this way and shifted to here'. First, relating two frames has two parts: which way a frame faces (rotation) and where its origin is (translation). Two frames can point in different directions (one's 'forward' is the other's 'left': that's a rotation) and their starting points can be in different places (one origin is a metre over from the other: that's a translation). To re-describe a point from the other frame's viewpoint, you have to account for both: turn it to match the facing, and shift it to match the origin. Second, rotation order matters: turning then tilting is not the same as tilting then turning. This feels surprising but it's real: take your phone, turn it flat 90 degrees, then tilt it up 90 degrees, and notice where the screen points; now start over and tilt up first, then turn. The screen ends up pointing somewhere different. So when you combine rotations, you can't just add up angles or do them in any order: you have to apply them in the ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍right order (in the math, multiply the rotations in the right order). Third, packing both into one 'homogeneous' matrix lets you re-describe a point from another frame's viewpoint and chain instructions by multiplying. Instead of juggling 'turn, then shift' as two separate steps, robotics bundles the rotation and the translation into a single instruction (a matrix) that says exactly how one frame sits relative to another. Applying it to a point re-describes that point from the other frame's viewpoint in one go. Chaining is what makes it useful: if you have 'the camera sits like this on the robot' and 'the robot sits like that in the world', you multiply the two instructions to get 'the camera sits like this-and-that in the world', one combined instruction, built by multiplication. That's why this one tool (the homogeneous transformation matrix) is the geometric engine of robotics: it captures a frame's full 'turned and shifted' relationship as a single thing, applies it in one step, and chains down a whole sequence of frames (sensor to robot to world, or joint to joint down an arm) just by multiplying the instructions together.

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Compared to the model answer - did you get it?

FError Analysis Framework

  • Treating a general rotation as just adding angles. Why: a rotation is 'an angle'. Recognise: general/3D orientation is multi-dimensional, a rotation matrix, not a number. Avoid: use rotation matrices (or quaternions); don't add angles, multiply matrices.
  • Composing rotations in any order. Why: rotations seem like they'd add up. Recognise: rotations don't commute. Order changes the final orientation. Avoid: multiply rotations in the correct order matching how they're physically applied (R2*R1).
  • Applying translation before rotation (or in ad-hoc steps). Why: rotate and translate seem independent. Recognise: a rigid transform is rotate THEN translate; swapping gives a different result. ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍Avoid: use the homogeneous matrix, which encodes the correct combined operation.
  • Dropping the homogeneous '1' / keeping rotation and translation separate. Why: the extra coordinate looks pointless. Recognise: the 1 is what lets one matrix do both rotation and translation and compose by multiplication. Avoid: write points as [x,y,z,1] and use 4x4 homogeneous matrices so pose is one composable object.

GMini Challenge

Explain rotations and homogeneous transformations for a new roboticist: rotation matrices (and why order matters), translation, the rigid transform (rotate then translate), and the homogeneous transformation matrix: explaining why it packs rotation and translation into one composable object and what that enables (the transform tree, kinematics, perception).

‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍REVEAL MODEL ANSWER
MODEL ANSWER

Rotation: a frame's orientation (which way its axes point) is a rotation matrix R. Multiply a point's coordinates by R to rotate them (an angle in 2D; a matrix in general/3D, because orientation is multi-dimensional). Compose rotations by multiplying matrices, and order matters: rotations don't commute (R2R1 != R1R2) - 'turn then tilt' differs from 'tilt then turn'. You can't just add angles.

Translation: the frames' origins differ, so you add an offset vector t (shift the point).

Rigid transform (rotate then translate): p' = R*p + t. It accounts for both how the frames are oriented (R) and where their origins are (t), preserving distances/shapes. Order is rotate then translate (swapping gives a different result).

Homogeneous transformation matrix T: the trick that packs both R and t into ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍one matrix (4x4 in 3D, with R top-left, t in the last column), using points written with an extra 1 ([x, y, z, 1]). The appended 1 makes a single matrix-vector multiply do both Rp and + t. So: applying a transform = one multiply (p' = Tp); composing transforms = one multiply ().

Why it packs them into one composable object, and what it enables: because applying is now matrix multiplication, composing is too: so a chain of frames collapses to multiplying matrices. This makes pose a single object (T = a frame's full orientation + position) and enables: the transform tree (map -> odom -> base_link -> sensor are transforms on edges; querying a point's frame = multiplying along the path: exactly tf2, localization-slam-nav2 04); forward kinematics (a manipulator's tip pose = the product of joint transforms: Turn 2 DH parameters); ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍perception (a measurement onto the map = multiply by the sensor-to-map transform, the perception/costmaps lesson); and invertibility (the reverse transform = the matrix inverse). One idea (pack R and t, append a 1) buys a uniform, composable, invertible algebra of pose: the geometric engine of the whole topic (and the same math as Python's transforms/tf). Rotation representations deepen in Turn 2 (rotation-matrix properties, Euler angles, quaternions).

Compared to the model answer - did you get it?

Quiz Check

A quick auto-graded check, separate from the recall cards above. Your score is pooled with the recall cards into this module's Mastery score, and completing this lesson requires the quiz submitted with pooled mastery at 80% or above.

QUIZAuto-graded check · feeds your mastery score
  1. ‍​‌‌​​‌‌​​‌‌‌​​‌​​‌‌​​‌​‌​‌‌​​‌​‌​​‌​‌‌​‌​‌‌‌​​‌‌​‌‌​​​​‌​‌‌​‌‌​‌​‌‌‌​​​​​‌‌​‌‌​​​‌‌​​‌​‌‍A general rotation is represented by:

  2. A rigid transformation relating two frames is:

  3. A homogeneous transformation matrix is useful because:

  4. Composing transforms along a chain of frames (sensor -> robot -> world) is done by:

This is a free sample

Progress and the spaced-repetition reviews are part of the course. The full track continues from here.