Professional Documents
Culture Documents
1 2
Adapted from UCB CS252 S01
3 4
5 6
1
Popular Flynn Categories for
Where is Supercomputing heading? Parallel Computers
1997, 500 fastest machines in the world: SISD (Single Instruction Single Data)
319 MPPs, 73 bus-based shared memory (SMP), 106 Uniprocessors
parallel vector processors (PVP) MISD (Multiple Instruction Single Data)
multiple processors on a single data stream
http://www.top500.org/
7 8
9 10
11 12
2
Shared Address/Memory
Shared-Memory Programming Examples Multiprocessor Model
struct alloc_t { int first; int last } alloc[MAX_THR];
pthread_t tid[MAX_THR]; Communicate via Load and Store
Oldest and most popular model
main()
{
… Based on timesharing: processes on multiple
processors vs. sharing single processor
for (int i=0; i<num_thr; i++) {
alloc[i].first = i*N/num_thr;
alloc[i].last = (i!=num_thr)?((i+1)*(N/num_thr)-1):N;
process: a virtual address space
and > 1 thread of control
pthread_create(&tid[i]/*thread id pointer*/,
NLLL/*detach method*/, dmm_func/*thread function*/,
(void *)&alloc[i]/*parameters*/);
ALL threads of a process share a process address
space
}
for (i=0; i<num_thr; i++) {
pthread_join(tid[i]/*thread id*/, NULL/*return value*/); Example: Pthread
Writes to shared address space by one thread
}
}
are visible to reads of other threads
dmm_func(struct alloc_t *alloc) {
for (int i=alloc->first; i<alloc->last; i++)
for (int k=0; k<N; k++)
for (int j=0; j<N; j++)
Z[i][j] += X[i][k]*Y[k][j];
} 13 14
15 16
Advantages of Message-Passing
Message Passing Model Communication
Whole computers (CPU, memory, I/O devices), explicit The hardware can be much simpler and is usually
send/receive as explicit I/O operations standard
Send specifies local buffer + receiving process on
remote computer Explicit communication => simpler to understand, help
make effort to reduce communication cost
Receive specifies sending process on remote computer +
local buffer to place data Synchronization is naturally associated with
sending/receiving messages
Send+receive => memory-memory copy, where each each
supplies local address Easier to use sender-initiated communication, which may
have some advantages in performance
17 18
3
Three Types of Parallel Applications Amdahl’s Law and Parallel Computers
Commercial Workload Amdahl’s Law: speedup is limited by the fraction
TPC-C, TPC-D, Altavista (Web search engine) of the portions that can be parallelized
4
An Example Snoopy Protocol Snoopy-Cache State Machine-I
CPU Read hit
Invalidation protocol, write-back cache State machine
for CPU requests
Each block of memory is in one state: for each CPU Read Shared
Clean in all caches and up-to-date in memory (Shared) cache block
Invalid
Place read miss
(read/only)
OR Dirty in exactly one cache (Exclusive) on bus
OR Not in any caches
Each cache block is in one state (track these): CPU Write
CPU read miss CPU Read miss
Shared : block can be read
Place Write
Write back block, Place read miss
Miss on bus
OR Exclusive : cache has only copy, its writeable, and Place read miss on bus
dirty on bus
OR Invalid : block contains no data CPU Write
Cache Block Place Write Miss on Bus
Read misses: cause all caches to snoop bus State Exclusive
Writes to clean line are treated as misses (read/write)
CPU read hit CPU Write Miss
CPU write hit Write back cache block
Place write miss on bus
25 26
Example Example
P2:
P2: Write
Write 20 to A1
to A1 P2:
P2: Write
Write 20 to A1
to A1
P2:
P2: Write
Write 40 to A2
to A2 P2:
P2: Write
Write 40 to A2
to A2
Assumes A1 and A2 map to same cache block, Assumes A1 and A2 map to same cache block
initial cache state is invalid
29 30
5
Example Example
Assumes A1 and A2 map to same cache block Assumes A1 and A2 map to same cache block
31 32
Example Example
Assumes A1 and A2 map to same cache block Assumes A1 and A2 map to same cache block,
but A1 != A2
33 34
6
Implementing Snooping Caches MESI Highlights
Bus serializes writes, getting bus ensures no one From local processor P’s viewpoint, for each cache
else can perform memory operation block
On a miss in a write back cache, may have the Modified: Only P has a copy and the copy has
desired copy and its dirty, so must reply been modifed; must respond to any read/write
request
Add extra state bit to cache to determine Exclusive-clean: Only P has a copy and the copy
shared or not is clear; no need to inform others about my
Add 4th state (MESI) changes
Modfied (private,!=Memory) Shared: Some other machines else may have
eXclusive (private,=Memory) copy; have to inform others about P’s changes
Shared (shared,=Memory) Invalid: The block has been invalidated (possibly
Invalid on the request of someone else)
37 38
MESI Highlights
Actions:
Have read misses on a block: send read
request onto bus
Have write misses on a block: send write
request onto bus
Receive bus read request: transit the
block to shared state
Receive bus write request: transit the
block to invalid state
Must write back data when transiting
from modified state
39