CPU performance – clock speed, cores, cache and system on a chip

Computer ScienceComputers & HardwareAges 16–17

Loading…

Use with my class ✨ Customize with AI Report a problem

Set the clock speed, the number of cores and the L2 cache size, then run five kinds of task to see how long each takes. A chart of core activity separates the serial part, the parallel part and time spent waiting for memory, and an Amdahl's law graph shows why adding cores has a limit. The Memory hierarchy tab steps through reads via registers, L1, L2 and RAM to show cache hits, misses and locality; the System on a chip tab compares power with performance.

Lesson: CPU performance: clock speed, number of cores, cache size and pipelining; memory hierarchy; system on a chip (SoC)

What it shows

Three factors decide how fast a processor finishes a task. Clock speed sets how many cycles run each second; more cores help only with the part of a task that can run in parallel, which is Amdahl's law: speed-up = 1 / ((1 − p) + p/n). Cache is small, fast memory close to the cores: a cache hit avoids a slow trip to RAM, and programs with good locality of reference get far more hits. Pipelining overlaps the stages of several instructions. A system on a chip puts CPU, GPU, memory and radios on one die to save energy and space.

How to use

In Performance test, move Clock speed, Cores and L2 cache, choose a Task and Parallel part p, tick Pipelining, then press Run test and Add to table to compare results. In Memory hierarchy, choose a Program and press Step or Run to watch each read hit L1, hit L2 or go to RAM. In System on a chip, tap a block and switch Build to compare power and battery life.

Parameters you can change

  • Starting tab Performance test, Memory hierarchy, System on a chip
  • Clock speed 1–5 GHz
  • Number of cores 1–16
  • L2 cache size 0.5 MB, 1 MB, 2 MB, 4 MB, 8 MB, 16 MB, 32 MB
  • Task Run a web-page script, Recalculate a spreadsheet, Apply a photo filter, Compile a large program, Export a video
  • Parallel part of the task 0–100 %
  • Pipelining
  • Memory-read program Read an array in order, Re-use a small array (12 items) 5 times, Re-use a large array (40 items) 3 times, Random reads

Questions to explore

  1. Why does doubling the cores from 8 to 16 barely help the web-page script task?
  2. Why does reading an array in order give a 75% L1 hit rate even though every item is new?
  3. Why does raising the clock from 4 to 5 GHz cost much more power than from 1 to 2 GHz?