This gives a main memory bandwidth of roughly 5 MB/sec. A top-end single CPU system today has probably 25 GB/sec available to it, a factor of 5,000 more. Moreover, much of that is achieved through optimizing burst reads--actual sustained random access throughput is going to be much lower and the delta much less.
Given modern implementation techniques, the actual efficiency loss is probably on the order of 10x rather than 1000x. And much of it is the result of the memory wall, which has been driven by DRAM physics rather than micro-architecture. Doing a couple of memory lookups to support dynamic dispatch is a hell of a lot more expensive, relative to an ALU operation, these days than it was 30 years ago.
The Alto's main memory had a cycle time of about 850 nsec, and could transfer 2 16-bit words per cycle: http://www.computer-refuge.org/bitsavers/pdf/xerox/parc/tech....
This gives a main memory bandwidth of roughly 5 MB/sec. A top-end single CPU system today has probably 25 GB/sec available to it, a factor of 5,000 more. Moreover, much of that is achieved through optimizing burst reads--actual sustained random access throughput is going to be much lower and the delta much less.
Given modern implementation techniques, the actual efficiency loss is probably on the order of 10x rather than 1000x. And much of it is the result of the memory wall, which has been driven by DRAM physics rather than micro-architecture. Doing a couple of memory lookups to support dynamic dispatch is a hell of a lot more expensive, relative to an ALU operation, these days than it was 30 years ago.