CPU performance – clock speed, cores, cache and system on a chip
Computer ScienceComputers & HardwareAges 16–17
Loading…
Sign in to playSet the clock speed, the number of cores and the L2 cache size, then run five kinds of task to see how long each takes. A chart of core activity separates the serial part, the parallel part and time spent waiting for memory, and an Amdahl's law graph shows why adding cores has a limit. The Memory hierarchy tab steps through reads via registers, L1, L2 and RAM to show cache hits, misses and locality; the System on a chip tab compares power with performance.
Lesson: CPU performance: clock speed, number of cores, cache size and pipelining; memory hierarchy; system on a chip (SoC)
What it shows
Three factors decide how fast a processor finishes a task. Clock speed sets how many cycles run each second; more cores help only with the part of a task that can run in parallel, which is Amdahl's law: speed-up = 1 / ((1 − p) + p/n). Cache is small, fast memory close to the cores: a cache hit avoids a slow trip to RAM, and programs with good locality of reference get far more hits. Pipelining overlaps the stages of several instructions. A system on a chip puts CPU, GPU, memory and radios on one die to save energy and space.
How to use
In Performance test, move Clock speed, Cores and L2 cache, choose a Task and Parallel part p, tick Pipelining, then press Run test and Add to table to compare results. In Memory hierarchy, choose a Program and press Step or Run to watch each read hit L1, hit L2 or go to RAM. In System on a chip, tap a block and switch Build to compare power and battery life.
Parameters you can change
- Starting tab Performance test, Memory hierarchy, System on a chip
- Clock speed 1–5 GHz
- Number of cores 1–16
- L2 cache size 0.5 MB, 1 MB, 2 MB, 4 MB, 8 MB, 16 MB, 32 MB
- Task Run a web-page script, Recalculate a spreadsheet, Apply a photo filter, Compile a large program, Export a video
- Parallel part of the task 0–100 %
- Pipelining
- Memory-read program Read an array in order, Re-use a small array (12 items) 5 times, Re-use a large array (40 items) 3 times, Random reads
Questions to explore
- Why does doubling the cores from 8 to 16 barely help the web-page script task?
- Why does reading an array in order give a 75% L1 hit rate even though every item is new?
- Why does raising the clock from 4 to 5 GHz cost much more power than from 1 to 2 GHz?