Operating Systems · Module 8 — File Systems, Storage & I/O
I/O: polling, interrupts, DMA, buffering and caching
Both keep data in RAM, and they are constantly confused.
Sign in to track your score
The backup copies matrix — 1.4 MB, 350 blocks — to the external disk.
Somebody has to carry 1.4 MB from RAM to the disk controller.
If the CPU does that byte by byte, it is doing nothing else for the entire transfer. On a machine with four cores and a compile running, that is a noticeable loss.
So the CPU does not do it.
Why & what
Three ways to run a device, in order of how much CPU they waste.
- Polling. The CPU repeatedly asks the device whether it is ready. It works, and it burns an entire core doing nothing. Only sensible when the wait is shorter than the cost of being interrupted.
- Interrupt-driven. The CPU starts the operation and goes to run something else. The device raises an interrupt when it is done — Topic 1.4's mechanism, used for its original purpose. Much better, but the CPU still copies each block, and 350 blocks means 350 interrupts.
- DMA — direct memory access. A separate controller does the copying. The CPU says "move 350 blocks from this address to that device" and walks away. The DMA controller transfers everything and raises one interrupt at the end.
DMA is the one that matters. With interrupt-driven I/O the CPU is still the delivery vehicle. With DMA it is not involved in the transfer at all.
Buffering and caching. Both keep data in RAM, and they are constantly confused.
- A buffer smooths a speed or size mismatch. The SSD delivers a whole 4 KB block; the editor asked for 100 bytes. The buffer holds the rest, so the next forty reads are satisfied without touching the disk. Data is passing through a buffer.
- A cache keeps a copy of something in case it is wanted again. The page cache keeps recently read disk blocks in spare RAM, so reading matrix.c a second time never reaches the SSD at all. Data in a cache has already arrived; the copy is speculative.
One sentence to keep them apart: a buffer holds data on its way through, a cache holds a copy in case it is wanted again.
Why writes are dangerous. The page cache also absorbs writes. When the compile job writes to build.log, the data goes into RAM and the call returns immediately — long before anything reaches the SSD. That is why writes feel instant.
It is also why pulling the power out corrupts files. The OS periodically flushes dirty blocks to disk, and a program that truly needs data on the platter has to ask explicitly.
How it works
The backup writing 350 blocks with DMA:
- The process makes a write system call. Trap into the kernel, as in Topic 1.3.
- The kernel sets up the DMA controller with a source address in RAM, a destination device, and a count.
- The process goes to Waiting and the scheduler runs something else — the compile job gets the core.
- The DMA controller transfers all 350 blocks on its own, without the CPU.
- It raises one interrupt when finished. The kernel marks the process Ready, and it resumes.

Common confusion
"Interrupt-driven I/O means the CPU is free during the transfer." It is free during the waiting, not during the transfer. Each block still passes through the CPU. DMA is what frees it from the transfer itself.
"DMA makes the disk faster." The disk runs at exactly the same speed. DMA frees the CPU to do other work while the disk takes as long as it always did.
"Polling is always wrong." For a device that responds in nanoseconds, setting up an interrupt costs more than the wait. Spinning is the right answer there for the same reason spinlocks were in Topic 4.3 — the comparison is against the cost of switching away, not against zero.
Interview angle
"Explain polling, interrupts and DMA." Order them by CPU cost and give the distinguishing line for each: polling wastes the core, interrupts free the wait but not the transfer, DMA frees both. "Difference between a buffer and a cache?" Buffers handle mismatched speeds or sizes for data in transit; caches hold copies of data already fetched in case it is needed again. Give the page cache as the example, since it also explains why a second read of a file is so much faster. "Why can a crash lose data that a program already wrote?" Because the write returned as soon as the data reached the page cache, not the disk. This question is asked more often than people expect, and the answer connects directly to why databases force flushes.
- 1.
With DMA, how many interrupts does a 350-block transfer generate?
- 2.
Which best describes the page cache?
- 3.
A program writes to build.log and the call returns immediately. Where is the data?