PSTSWM Paragon Communication Protocol Performance

Performance Studies using

PSTSWM


Intel Paragon Protocol Performance

(transpose LT experiment II-A2 / O(P) swap transpose algorithm)

Date/Person: October 21, 1994 / P. Worley
Platform: Intel Paragon at Sandia National Laboratory (acoma):
     1824 GP nodes (2 50-MHz iPSC/860 processors per node)
Environment: SUNMOS 1.6.1
f77/Paragon Paragon Version ???
Code Version: 3.2
Compilation Options: if77 -O4 -Mnodepchk -Knoieee
Math Library: none
Communication Library: SUNMOS
Parallel Algorithm: swtrans
Partition: 1x8, 1x16, or 1x32
Number of Timesteps: 12
Results:

1x16 Processors / Problem T21L2
Runtime Statistics
min(mean-min)/min(median-min)/min(max-min)/min
  2.1127e-01   0.09   0.06   1.11 
Three Fastest
Protocols
1st2nd3rd
  e3   e1   e2 
       Number of Proctocols With
Runtimes Within X% of Min
1%5%25%
  4   12   29 

1x32 Processors / Problem T42L1
Runtime Statistics
min(mean-min)/min(median-min)/min(max-min)/min
  2.8096e-01   0.11   0.11   0.24 
Three Fastest
Protocols
1st2nd3rd
  e1   a1   e0 
       Number of Proctocols With
Runtimes Within X% of Min
1%5%25%
  3   9   29 

1x8 Processors / Problem T42L2
Runtime Statistics
min(mean-min)/min(median-min)/min(max-min)/min
  1.2768e+00   0.09   0.08   0.11 
Three Fastest
Protocols
1st2nd3rd
  a0   b6   d2 
       Number of Proctocols With
Runtimes Within X% of Min
1%5%25%
  1   1   30 

1x16 Processors / Problem T85L1
Runtime Statistics
min(mean-min)/min(median-min)/min(max-min)/min
  1.4990e+00   0.18   0.19   0.21 
Three Fastest
Protocols
1st2nd3rd
  a0   b6   c2 
       Number of Proctocols With
Runtimes Within X% of Min
1%5%25%
  1   1   30 

1x32 Processors / Problem T85L2
Runtime Statistics
min(mean-min)/min(median-min)/min(max-min)/min
  1.8680e+00   0.01   0.01   0.07 
Three Fastest
Protocols
1st2nd3rd
  b0   b1   a1 
       Number of Proctocols With
Runtimes Within X% of Min
1%5%25%
  14   29   30 

1x8 Processors / Problem T85L4
Runtime Statistics
min(mean-min)/min(median-min)/min(max-min)/min
  1.2639e+01   0.21   0.22   0.25 
Three Fastest
Protocols
1st2nd3rd
  a0   e5   e4 
       Number of Proctocols With
Runtimes Within X% of Min
1%5%25%
  1   1   30 

DISCUSSION

Unlike the more recent Paragon OSF experiments, partitions of the Paragon processor grid were NOT used that match the processor subsets that the parallel algorithms would run on in a two dimensional data decomposition. Partitions WERE chosen to minimize the impact of other users on the timings.

Patrick H. Worley / ( worleyph@ornl.gov)
Last Modified Monday, 15-Jul-2002 10:23:09 EDT.