ALearning Material
An embedded system is a computer built into a device to perform a specific, dedicated function: the controller in a robot, a washing machine, a car's engine, a thermostat. Before diving into programming microcontrollers, you need the big picture: what an embedded system is, how it differs from a general-purpose computer, and the architectural choices (MCU vs MPU, the memory model) that shape how you design for it.
The defining characteristics of an embedded system: it's dedicated (does one specific job, not general computing), resource-constrained (limited memory, processing, power: unlike a PC), often real-time (must respond within deadlines), and directly interfaces hardware (sensors, actuators, via GPIO/peripherals). This makes embedded design fundamentally different from PC/server programming.
The key architectural distinction is microcontroller (MCU) vs microprocessor (MPU):
MICROCONTROLLER (MCU): CPU + memory (RAM, Flash) + peripherals all on ONE chip
- self-contained, low-power, cheap, real-time, runs ONE program (often bare-metal or RTOS)
- for dedicated control: robots' controllers, sensors, appliances
MICROPROCESSOR (MPU): just the CPU; needs external RAM, storage, etc.; runs an OS (Linux)
- more powerful, for complex tasks (vision, networking); a small computer (e.g. a Raspberry Pi)
You also need the memory model: embedded systems distinguish Flash (non-volatile: holds the program) from RAM (volatile: working data), and many MCUs use a Harvard architecture (separate program and data buses) for speed. The disciplines: understand the embedded mindset (dedicated, constrained, real-time, hardware-interfacing), choose MCU vs MPU for the job (dedicated control vs complex computing), and know the memory model (Flash vs RAM, the constraints). This framing shapes every embedded design decision.
Formulas & method (budgeting RAM, and knowing where each object lives).
RAM BUDGET sum every buffer, every static, and the worst-case stack:
ring_buffer_samples * bytes_per_sample
+ each fixed buffer
+ reported statics (.data + .bss)
+ measured worst-case stack depth
and leave HEADROOM: 20 to 30 percent, because the worst case you measured
is the worst case you happened to hit.
Where each thing lives, which is the whole of the Flash/RAM split:
| Object | Lives in | Why |
|---|---|---|
const lookup table |
Flash, read in place | it never changes, so copying it to RAM wastes RAM |
char buf[256] inside a function |
stack (RAM), transiently | created on entry, gone on return; it is why stack depth is a budget item |
global int counter = 5 |
RAM (.data), with the initial value stored in Flash |
startup code copies the 5 from Flash into RAM before main |
global int counter; (no initialiser) |
RAM (.bss), zeroed at startup |
zeros do not need to be stored, only written |
| a string literal | Flash | it is const data; passing it passes a pointer into Flash |
So a global initialised variable costs both Flash and RAM, and a const costs only Flash. The
practical rules that follow: mark read-only tables const (on some toolchains, plus a
PROGMEM-style attribute) so they stay out of RAM, prefer a fixed-size buffer to anything that grows,
and measure the stack rather than estimating it, by filling RAM with a pattern at startup and
looking at how much of it has been overwritten after a long run.
Why it exists. Embedded systems are dedicated, resource-constrained, often real-time computers that directly control hardware, which makes designing them fundamentally different from general-purpose programming. Understanding what an embedded system is, the MCU-vs-MPU choice, and the memory model frames the whole approach, so you design for the constraints and the dedicated, hardware-interfacing nature, rather than as if it were a PC.
Mental model. A microcontroller is a whole tiny self-contained appliance-brain on one chip (CPU, memory, and the I/O all built in) dedicated to one job, like the single-purpose brain inside a thermostat. A microprocessor is more like a bare computer CPU that needs external memory and storage bolted on to become a small general computer (a Raspberry Pi). You pick the self-contained appliance-brain for dedicated control, the small computer for heavy, complex work.
Common misunderstandings.
- "An embedded system is just a small PC." It's dedicated, resource-constrained, real-time, and hardware- interfacing: designed under tight constraints for one job, not general computing; the mindset is different.
- "MCU and MPU are the same / interchangeable." An MCU is self-contained (CPU+memory+peripherals on one chip) for dedicated low-power real-time control; an MPU is just a CPU needing external memory/OS for complex computing. Choose by the job.
- "Memory is just memory." Embedded distinguishes Flash (non-volatile program storage) from RAM (volatile working data), both tightly limited, and the Harvard model (separate buses) affects how it works; the memory model matters.
Connections. This frames the whole embedded topic: the MCU is what the GPIO/peripherals/interrupts/timers lessons program (the MCU/GPIO and interrupts/timers lessons), the constraints drive the RTOS/memory/low-power lessons (Pass 2), the Flash/RAM model underlies the memory-management lesson, and the MCU-vs-MPU choice connects to Embedded Linux (MPU) and the robotics platform.
BImmediate Active Recall
QUERYWhat defines an embedded system, and how does it differ from a general-purpose computer?
REVEAL
An embedded system is a computer built into a device for a specific, dedicated function. It differs from a PC by being dedicated (one job), resource-constrained (limited memory/processing/power), often real-time (must meet deadlines), and directly hardware-interfacing (sensors/actuators via GPIO/peripherals). These make embedded design fundamentally different from general-purpose programming.
QUERYWhat is the difference between a microcontroller (MCU) and a microprocessor (MPU)?
REVEAL
An MCU has the CPU, memory (RAM/Flash), and peripherals all on one chip (self-contained, low-power, cheap, real-time, running one program (bare-metal or RTOS)) for dedicated control. An MPU is just the CPU, needing external RAM/storage and running an OS (Linux), more powerful, for complex tasks (a small computer like a Raspberry Pi). MCU for dedicated control; MPU for heavy computing.
QUERYWhat is the embedded memory model (Flash vs RAM), and what is Harvard architecture?
REVEAL
Embedded systems distinguish Flash (non-volatile: holds the program, retained when powered off) from RAM (volatile: working data, lost on power-off), both tightly limited. Harvard architecture uses separate buses for program and data (vs von Neumann's shared bus), allowing the CPU to fetch instructions and data simultaneously. Common in MCUs for speed. Knowing Flash vs RAM and the limited sizes is essential to embedded design.
QUERYWhen would you choose an MCU vs an MPU?
REVEAL
Choose an MCU for dedicated, low-power, real-time control where the job is specific and constrained (a robot's motor/sensor controller, an appliance): it's self-contained, cheap, and deterministic. Choose an MPU (running Linux) for complex, powerful tasks (vision, networking, heavy computation, rich software) where you need a small general computer. Match the choice to whether the job is dedicated control or complex computing.
CConceptual Questions
Answer each in your own words in the box, then reveal the model answer to compare. These ask why, not how, and your answers are saved.
Why does the embedded mindset (dedicated, resource-constrained, real-time, hardware-interfacing) make embedded design fundamentally different from general-purpose programming, and why does framing it this way matter?
REVEAL MODEL ANSWER
Embedded design is fundamentally different because each of those characteristics inverts an assumption that general-purpose programming takes for granted, and together they demand a different way of thinking. Dedicated: an embedded system does one specific job forever, so you design the whole system, hardware and software, around that single function, rather than writing general software to run on someone else's general machine; the program, the chip, and the I/O are co-designed for the task. Resource-constrained: where a PC programmer assumes effectively unlimited memory and processing, an embedded developer works with kilobytes of RAM, limited Flash, a slow clock, and a tight power budget, so efficiency isn't a nice-to-have but a hard requirement: every byte and cycle counts, data structures and algorithms are chosen for smallness, and wasteful abstractions are unaffordable. Real-time: many embedded systems must respond within strict deadlines (a motor control loop, a safety cutoff), so correctness includes timing (being right too late is wrong) which forces attention to determinism, interrupt latency, and worst-case timing that general programming rarely considers. Hardware-interfacing: the program directly manipulates physical hardware through registers, GPIO, and peripherals, reading sensors and driving actuators, so the developer must understand the electronics and the chip's registers, not just abstract data. Software and hardware are inseparable. Framing embedded work by these traits matters because it sets the defaults and priorities of every design decision: you assume scarcity and design for efficiency, you assume deadlines and design for determinism, you assume direct hardware control and design around registers and peripherals, and you assume a single dedicated purpose and co-design the whole system for it. A developer who approaches an embedded system as 'a small PC' brings the wrong assumptions (expecting abundant resources, ignoring timing, abstracting away the hardware) and produces software that doesn't fit, misses deadlines, or can't talk to the hardware. So the mindset isn't a label; it's the lens that makes you design appropriately for what an embedded system actually is, which is exactly why understanding it first frames the entire topic.
Why is the MCU-vs-MPU distinction a foundational design choice, and how does the self-contained-vs-needs-external-support difference determine which fits a given job?
REVEAL MODEL ANSWER
The MCU-vs-MPU distinction is foundational because it's one of the first and most consequential decisions in an embedded design: it determines the system's capability, cost, power, complexity, real-time behaviour, and software model, and it follows directly from the structural difference between the two. A microcontroller integrates the CPU, memory (RAM and Flash), and peripherals onto a single chip, making it self-contained: it needs little external support, boots and runs its one program immediately (bare-metal or under an RTOS), draws little power, costs little, and behaves deterministically. Ideal for dedicated, real-time control where the job is specific and the constraints are tight. A microprocessor is essentially just the CPU; it requires external RAM, external storage, and supporting circuitry, and it runs a full operating system (typically Linux): making it far more powerful and capable of complex software (vision, networking, multitasking, rich libraries) but also bigger, more power-hungry, more expensive, more complex to design around, and less inherently real-time. So the structural difference (everything-on-one-chip versus a-CPU-that-needs-a-computer-built-around-it) maps onto exactly the trade-offs that decide the fit. If the job is dedicated control with hard real-time needs, low power, low cost, and modest computation (a robot's motor/sensor controller, an appliance, a sensor node) the MCU's self-containment, determinism, and efficiency make it the right choice; an MPU would be overkill, power-hungry, and harder to make real-time. If the job demands heavy computation, complex software, networking, or a rich OS (computer vision, a high-level robot brain, connectivity-heavy applications) the MPU's power and OS support are necessary, and the MCU couldn't do it. Real systems often use both: an MPU for the complex high-level work and one or more MCUs for the dedicated real-time control, each playing to its strength. Recognizing that the choice flows from this self-contained-vs-needs-support structure (and the capability/power/cost/real-time trade-offs it implies) is what lets you pick the right processor for the job: a decision that shapes the entire rest of the design, which is why it's foundational.
DPractice Problems
P1 (easy). A sensor node has 20 KB of RAM and 128 KB of Flash. Its firmware needs: a 512-sample ring buffer of 16-bit ADC readings, a 256-byte UART receive buffer, a 1 KB transmit buffer, a 2 KB filesystem scratch area, and the compiler reports 1.8 KB of other statics. Measured worst-case stack depth is 3.2 KB.
Compute the RAM budget and say whether it fits, with how much margin. Then say which single item you would attack first if it did not, and what happens at run time when a RAM budget is exceeded by 200 bytes.
P2 (medium). Choose MCU or MPU for three robotics subsystems, and justify each on specific numbers rather than a general preference.
(a) A motor controller closing a current loop at 20 kHz with a jitter budget of 2 microseconds. (b) A vision node running an object detector at 10 frames per second on 640 by 480 images. (c) A battery-management node reading eight cell voltages once a second and lasting five years on a coin cell.
For each, name the choice, the deciding number, and the failure that would result from the other choice.
P3 (harder). Explain the Flash and RAM split by working through what happens to four specific things: a const lookup table of 4 KB, a char buf[256] inside a function, a global int counter = 5, and a string literal passed to a logging call.
For each, say where it lives at rest, where it lives at run time, and what it costs. Then state the two programming habits that follow directly from this memory model, and what each one prevents.
Solutionsclick to reveal
P1. The budget.
| Item | Size |
|---|---|
| ADC ring buffer, 512 samples times 2 bytes | 1024 B |
| UART receive buffer | 256 B |
| Transmit buffer | 1024 B |
| Filesystem scratch | 2048 B |
| Other statics | 1843 B |
| Static total | 6195 B |
| Worst-case stack | 3277 B |
| Total | 9472 B, about 9.3 KB |
It fits, using 46 percent of 20 KB, leaving 10.7 KB of margin. Comfortable, and comfortable is the right target: a device that fits in 95 percent of its RAM has no room for the next feature, no room for a deeper call path in an error handler, and no margin for the stack being deeper than measured.
Which item to attack first if it did not fit: the filesystem scratch area, 2 KB.
Not because it is the largest (the stack is), but because it is the one whose size is a choice rather than a consequence. A scratch buffer is usually sized to a convenient round number rather than to a requirement, and halving it costs nothing but a few more read-modify-write cycles.
The order after that: the transmit buffer (1 KB is often more than the link needs, and reducing it costs throughput, not correctness), then the ADC ring buffer (its depth is set by how long the consumer may be delayed, which is a measurable number rather than a guess), and the stack last, because reducing the stack means restructuring code.
What happens at run time when RAM is exceeded by 200 bytes.
This is the part worth being precise about, because it is not what a PC programmer expects.
The linker usually does not tell you. Statics are allocated at link time and the linker will report an overflow of the data section: that case is caught. The dangerous case is the stack, which is allocated at run time and typically grows downward from the top of RAM toward the statics growing up. Nothing checks that they have met.
So the 200 bytes are consumed silently, and the deepest call path writes 200 bytes over the top of your statics. The symptom is:
- A variable somewhere in the program changes value for no reason, at a moment unrelated to any code that touches it
- The corruption depends on the call depth, so it only happens on the error path, or only when an interrupt fires during a particular function, or only once the device has been running long enough to hit an unusual branch
- It is not reproducible, and it appears to be a different bug each time depending on which variable happened to be at that address
This is why embedded practice puts three defences in place that a hosted program does not need:
- Fill the stack region with a known pattern at startup and check at run time how far down the pattern has been overwritten. That converts "how much stack do I use" from a guess into a measurement, and it is a dozen lines of code
- Enable the MPU or a stack guard region where the hardware has one, so a stack overflow faults immediately at the point of the overflow rather than corrupting data silently
- Avoid dynamic allocation entirely.
mallocin a 20 KB device fragments, has no failure path anybody tests, and makes the worst-case memory use impossible to compute statically. Static allocation means the linker's map file is the truth
P2. (a) Motor current loop, 20 kHz, 2 microsecond jitter: MCU, bare metal or a small RTOS.
The deciding number: 2 microseconds of jitter on a 50 microsecond period, which is 4 percent.
An MCU running a timer-triggered interrupt has interrupt latency of a few hundred nanoseconds, deterministic and bounded, because there is no cache, no MMU, no scheduler and no other process. Jitter is dominated by whether a longer instruction is executing when the interrupt arrives: tens of nanoseconds.
What an MPU running Linux would do instead: even with PREEMPT_RT, interrupt latency is typically tens of microseconds with occasional excursions into hundreds, because of cache misses, TLB misses, and the scheduler. The occasional excursion is the killer: a control loop that meets its deadline 99.99 percent of the time and misses it by 200 microseconds once a second has a torque glitch once a second, which the motor turns into audible noise and the mechanism turns into wear.
(b) Vision, 10 fps object detection on 640 by 480: MPU, running Linux.
The deciding number: memory. A single 640 by 480 RGB frame is 921 600 bytes. A detector needs several frames plus intermediate tensors plus the model weights, so hundreds of megabytes. An MCU with 20 KB, or even 512 KB, cannot hold one frame, let alone run inference on it.
Beyond memory: the software stack. OpenCV, a neural network runtime, camera drivers and a filesystem are Linux software, and porting them to bare metal is not a project anyone should undertake.
What an MCU would do instead: nothing. This is not a performance trade-off, it is a capability boundary. The honest MCU version of this task is a much smaller model on a much smaller image, which is a different product.
(c) Battery management, 8 cells at 1 Hz, five years on a coin cell: MCU, deeply sleeping.
The deciding number: average current. A CR2032 holds about 220 mAh. Five years is 43 800 hours, so the average current must be under 220 / 43 800 = 5 microamps, and that has to include self-discharge, so the budget is really 2 to 3 microamps.
That is achievable only by being asleep essentially all the time: an MCU in deep sleep draws 1 to 2 microamps, wakes on a timer, reads eight channels in a few milliseconds, and sleeps again. The duty cycle is on the order of 0.1 percent.
What an MPU would do instead: an MPU running Linux idles at tens to hundreds of milliamps, cannot deep-sleep in the same sense (its DRAM must stay refreshed), and takes seconds to boot, so it cannot wake-work-sleep on a 1 second cycle. The coin cell would last hours.
The pattern across the three. The choice is never "MCU is simpler" or "MPU is more capable". It is decided by whichever of these three is binding:
- Determinism, measured in microseconds of jitter: MCU
- Memory and software ecosystem, measured in megabytes and in which libraries exist: MPU
- Average power, measured in microamps: MCU
And a real robot has all three requirements in different places, which is why it has both, and why the interesting engineering is the link between them.
P3.
| Item | At rest (power off) | At run time | Cost |
|---|---|---|---|
const lookup table, 4 KB |
Flash | Flash, read in place. Never copied to RAM | 4 KB of Flash, 0 bytes of RAM |
char buf[256] inside a function |
Nowhere. It does not exist until the function is entered | Stack, created on entry, destroyed on return | 256 B of stack, for the duration of that call and everything it calls |
int counter = 5 (global, initialised) |
Flash holds the value 5, in the initialisation image | RAM, in the .data section. Copied from Flash to RAM by the startup code before main |
4 bytes of RAM and 4 bytes of Flash |
| String literal passed to a log call | Flash | Flash, if the toolchain leaves it there and passes a pointer. On some architectures it is copied to RAM | Flash, and potentially RAM |
Two entries in that table are worth dwelling on.
The const table costs no RAM at all, because the processor can fetch from Flash directly. This is the single most useful fact in the model: on a device with 128 KB of Flash and 20 KB of RAM, Flash is six times more plentiful, so anything constant should be there. Drop the const and the table becomes a RAM variable that is initialised from Flash at startup, costing 4 KB of RAM and 4 KB of Flash: it consumes 20 percent of the device's RAM to gain nothing.
An initialised global costs both. The value has to be stored somewhere while the power is off, so it sits in Flash, and the startup code copies it into RAM before main runs. An uninitialised global (or one initialised to zero) goes in .bss, which costs RAM but no Flash, because the startup code just zeroes it.
The two habits that follow.
Habit 1: mark everything that does not change as const, and check the map file.
What it prevents: silent RAM consumption by data that never needed to be there. Lookup tables, string tables, configuration defaults, font data, calibration curves. On a constrained device this routinely recovers several kilobytes, which is a large fraction of the budget.
The check is the linker's map file, which reports the size of .text (code and constants in Flash), .data (initialised RAM, costing both) and .bss (zero-initialised RAM). Reading the map file after every build is the embedded equivalent of watching a test suite: it is the only place the truth about memory appears, and a 2 KB jump in .bss between two commits is a question worth asking on the day it happens rather than three months later.
Habit 2: allocate statically, never dynamically.
What it prevents: three distinct failures that all look like the same mysterious hang.
- Fragmentation. Repeated allocation and freeing of different sizes leaves a heap with enough total free space and no contiguous block large enough. On a device that runs for months, this is a matter of when, not whether
- An untested failure path.
mallocreturning null is a case almost nobody handles correctly, because it never happens in testing and always happens in the field - Unanalysable worst case. With static allocation, the linker map tells you the exact memory use, forever. With a heap, the worst case depends on the sequence of allocations, which depends on the sequence of events, which cannot be enumerated
The static version is a fixed pool sized at design time: static uint8_t buffers[N][SIZE]; with an index, or a fixed-capacity ring. It is less elegant and it is computable, and on a device with no display, no console and no operator, computable is worth more than elegant.
Both habits come from the same root: on an embedded device the memory is a fixed resource with no virtual memory, no swap and no operating system to arbitrate it. What you have is what the linker says you have, and every technique above exists to make that number knowable before the device ships.
EFeynman Exercise
Explain to a beginner, using the difference between a single-purpose appliance-brain and a bare computer CPU: (1) why an embedded system is a dedicated, constrained brain built for one job rather than a general computer, (2) why a microcontroller is a whole tiny self-contained brain on one chip while a microprocessor is a bare CPU that needs external parts to become a small computer, and (3) why you pick each for different kinds of jobs.
REVEAL MODEL ANSWER
An embedded system is best understood as a single-purpose appliance-brain rather than a general computer. Think of the brain inside a thermostat, a washing machine, or a robot's motor controller: it's dedicated to one job, it lives with tight limits (a little memory, a little power), it often has to react right on time (meet deadlines), and it's directly wired to the physical world (reading sensors, driving motors). That's completely different from a PC, which is a general machine with loads of memory and power for running anything, so you design an embedded brain for its constraints and its one job, not like a desktop. Now, there are two flavours of brain. A microcontroller (MCU) is a whole tiny self-contained brain on one chip: it has the thinking part (CPU), its memory, and its connections to the outside world (the I/O) all built in, ready to do its dedicated job straight away, sipping little power. It's like the complete little brain inside an appliance. A microprocessor (MPU) is more like a bare computer CPU, just the powerful thinking part, that needs external memory and storage bolted on to become a small general computer (like a Raspberry Pi running Linux). So you pick each for different jobs: the self-contained appliance-brain (MCU) for dedicated, real-time control that must be cheap, low-power, and reliable (the robot's motor and sensor controller); the small computer (MPU) for heavy, complex work like vision or networking (the robot's high-level brain). Often a robot has both. A powerful small computer for thinking and several tiny self-contained brains for the real-time muscle-and-sense control. A dedicated, constrained brain for one job, in two flavours (self-contained appliance-brain or bare-CPU-needing-support) chosen for the kind of work: that's the big picture of embedded systems.
FError Analysis Framework
- Treating an embedded system like a small PC. Why: it's a computer too. Recognise: you bring wrong assumptions (abundant resources, ignore timing/hardware). Avoid: design for the embedded mindset: dedicated, constrained, real-time, hardware-interfacing.
- Treating MCU and MPU as interchangeable. Why: both have a CPU. Recognise: you mis-pick. MCU is self-contained real-time control, MPU needs an OS for complex work. Avoid: choose MCU for dedicated low-power real-time control, MPU for complex computing.
- Ignoring the resource constraints when programming. Why: memory is cheap on a PC. Recognise: embedded RAM/Flash are tiny; wasteful code won't fit/run. Avoid: design for efficiency (small data structures, often static allocation).
- Treating Flash and RAM as the same. Why: it's all memory. Recognise: Flash is non-volatile program storage; RAM is volatile working data. Avoid: know the Flash/RAM model and their tight limits.
GMini Challenge
Frame the architecture for a robot that needs both real-time motor/sensor control and computer vision: decide MCU vs MPU for each part, explain the embedded characteristics that drive the choice, and note the memory-model considerations. Explaining how this framing shapes the design.
REVEAL MODEL ANSWER
Architecture decision (MCU + MPU: use both):
| Part of the robot | Processor | Why |
|---|---|---|
| Real-time motor/sensor control (encoders, control loops, safety) | MCU | Dedicated, real-time (hard deadlines), low-power, cheap, deterministic, self-contained |
| Computer vision / high-level brain (perception, planning, networking) | MPU (Linux) | Heavy computation, complex software, rich libraries. Needs a small general computer |
Embedded characteristics driving the choice:
- The control job is dedicated, real-time, resource-constrained, and hardware-interfacing: exactly the MCU's strengths (self-contained CPU+memory+peripherals, deterministic timing, low power).
- The vision job is complex, computation-heavy. Needing the MPU's power and OS (an MCU couldn't do it).
- So the robot uses both: the MPU thinks (vision/planning), the MCUs do the real-time muscle-and-sense control, communicating between them, each playing to its strength.
Memory-model considerations:
- On the MCU: distinguish Flash (holds the program/constants, non-volatile) from RAM (working data, volatile), both tightly limited, so design efficiently (small data structures, often static allocation, no reliance on a big heap), mindful of stack and code size (and the Harvard model for speed).
- On the MPU: external RAM/storage and an OS give far more memory, but it's not inherently real-time.
How this framing shapes the design: recognizing the robot has two kinds of job (dedicated real-time control vs complex computing) leads to the MCU+MPU split, which is how real robots are built; the embedded constraints (real-time, limited memory/power, hardware interfacing) dictate how you program the MCUs (efficiently, for determinism, around registers/peripherals); and the memory model (Flash vs RAM, tight limits) governs what fits and how you allocate. The big-picture framing (what an embedded system is, MCU vs MPU, the memory model) shapes every subsequent decision (peripherals, RTOS, memory management, low power), which is why it comes first.
Quiz Check
A quick auto-graded check, separate from the recall cards above. Your score is pooled with the recall cards into this module's Mastery score, and completing this lesson requires the quiz submitted with pooled mastery at 80% or above.