Parallel and distributed computing – Amdahl's law, GPUs and pipelining
Computer ScienceComputers & HardwareAges 16–17
Loading…
Sign in to playSplit a job into tasks with a parallel fraction p, then run it sequentially on one core, in parallel on n cores, or distributed over networked computers that must first receive their data. A Gantt chart shows what each core does, the speed-up graph follows Amdahl's law S = 1/((1 − p) + p/n) and its limit 1/(1 − p), and a table records your runs. The GPU tab runs an image filter on a few fast CPU cores and many simple GPU cores; the Pipelining tab shows fetch, decode and execute stages overlapping, with a stall caused by a data hazard.
Lesson: Parallel and distributed computing: sequential and parallel tasks, multicore processors, Amdahl's law, speed-up and efficiency, communication overhead, GPUs, pipelining and data hazards
What it shows
Parallel computing runs parts of a program at the same time on several cores; distributed computing shares the work between computers on a network, which must send data to one another. Only the parallel fraction p of a job can be shared out, so Amdahl's law gives the speed-up S = 1/((1 − p) + p/n), which can never exceed 1/(1 − p). In distributed systems, communication time grows with the number of computers and can make extra computers slow the job down. A GPU has thousands of simple cores that suit identical operations on many data items. Pipelining overlaps the fetch, decode and execute stages of successive instructions.
How to use
In the Parallel and distributed tab, choose the Mode, set the Parallel fraction p, the Number of cores n, the Tasks and, for distributed mode, Send data per computer; click Run to animate the Gantt chart and Add to table to record the result. In the CPU and GPU tab, choose the Task and the number of GPU cores and click Run. In the Pipelining tab, choose a Program, tick or untick Pipelining and use Step or Play.
Parameters you can change
- Starting tab Parallel and distributed, CPU and GPU, Pipelining
- Mode Sequential (1 core), Parallel (n cores), Distributed (n networked computers)
- Parallel fraction p 0–100 %
- Number of cores or computers n 1–32
- Number of parallel tasks 4–48
- Time to send data to each computer 0–5 s
- Number of GPU cores 32, 64, 128, 256, 512
- Task in the GPU tab Black-and-white filter (each pixel independent), Running blur (each pixel needs the one before)
- Pipelining on
- Program in the Pipelining tab Independent instructions, With a data hazard, Hazard, instructions reordered
Questions to explore
- With p = 80%, why does going from 16 to 32 cores hardly make the job any faster?
- In distributed mode, why can adding more computers make the job slower?
- Why does the GPU win the black-and-white filter but lose the running blur to the CPU?