CUDA: Why Thrust is so slow on uploading data to GPU? -
i'm new gpu world , installed cuda writing program. played thrust library find out slow when uploading data gpu. 35mb/s in host-to-device part on not-bad desktop. how come is?
environment: visual studio 2012, cuda 5.0, gtx760, intel-i7, windows 7 x64
gpu bandwidth test: 
it supposed have @ least 11gb/s of transfer speed host device or vice versa! didn't!
here's test program:
#include <iostream> #include <ctime> #include <thrust/device_vector.h> #include <thrust/host_vector.h> #define n 32<<22 int main(void) { using namespace std; cout<<"gpu bandwidth test via thrust, data size: "<< (sizeof(double)*n) / 1000000000.0 <<" gbytes"<<endl; cout<<"============program start=========="<<endl; int = time(0); cout<<"initializing h_vec..."; thrust::host_vector<double> h_vec(n,0.0f); cout<<"time spent: "<<time(0)-now<<"secs"<<endl; = time(0); cout<<"uploading data gpu..."; thrust::device_vector<double> d_vec = h_vec; cout<<"time spent: "<<time(0)-now<<"secs"<<endl; = time(0); cout<<"downloading data h_vec..."; thrust::copy(d_vec.begin(), d_vec.end(), h_vec.begin()); cout<<"time spent: "<<time(0)-now<<"secs"<<endl<<endl; system("pause"); return 0; } program out put: 
download speed: less 1 sec, pretty make sense compare nominal 11gb/s.
upload speed: 1.07374gb /32 secs 33.5 mb/s, doesn't make sense @ all.
does know reason? or way thrust is?
thanks!!
your comparison has several flaws, of covered in comments.
- you need eliminate allocation effects. can doing "warm-up" transfers first.
- you need eliminate "start-up" effects. can doing "warm-up" transfers first.
- when comparing data, remember
bandwidthtestusingpinnedmemory allocation, thrust not use. therefore thrust data transfer rate slower. typically contributes 2x factor (i.e. pinned memory transfers typically 2x faster pageable memory transfers. if want better comparisonbandwidthtestrun--memory=pageableswitch. - your choice of timing functions might not best. cudaevents pretty reliable timing cuda operations.
here code proper timing:
$ cat t213.cu #include <iostream> #include <thrust/device_vector.h> #include <thrust/host_vector.h> #include <thrust/copy.h> #include <thrust/fill.h> #define dsize ((1ul<<20)*32) int main(){ thrust::device_vector<int> d_data(dsize); thrust::host_vector<int> h_data(dsize); float et; cudaevent_t start, stop; cudaeventcreate(&start); cudaeventcreate(&stop); thrust::fill(h_data.begin(), h_data.end(), 1); thrust::copy(h_data.begin(), h_data.end(), d_data.begin()); std::cout<< "warm iteration " << d_data[0] << std::endl; thrust::fill(d_data.begin(), d_data.end(), 2); thrust::copy(d_data.begin(), d_data.end(), h_data.begin()); std::cout<< "warm iteration " << h_data[0] << std::endl; thrust::fill(h_data.begin(), h_data.end(), 3); cudaeventrecord(start); thrust::copy(h_data.begin(), h_data.end(), d_data.begin()); cudaeventrecord(stop); cudaeventsynchronize(stop); cudaeventelapsedtime(&et, start, stop); std::cout<<"host device iteration " << d_data[0] << " elapsed time: " << (et/(float)1000) << std::endl; std::cout<<"apparent bandwidth: " << (((dsize*sizeof(int))/(et/(float)1000))/((float)1048576)) << " mb/s" << std::endl; thrust::fill(d_data.begin(), d_data.end(), 4); cudaeventrecord(start); thrust::copy(d_data.begin(), d_data.end(), h_data.begin()); cudaeventrecord(stop); cudaeventsynchronize(stop); cudaeventelapsedtime(&et, start, stop); std::cout<<"device host iteration " << h_data[0] << " elapsed time: " << (et/(float)1000) << std::endl; std::cout<<"apparent bandwidth: " << (((dsize*sizeof(int))/(et/(float)1000))/((float)1048576)) << " mb/s" << std::endl; std::cout << "finished" << std::endl; return 0; } i compile (i have pcie gen2 system cc2.0 device)
$ nvcc -o3 -arch=sm_20 -o t213 t213.cu when run following results:
$ ./t213 warm iteration 1 warm iteration 2 host device iteration 3 elapsed time: 0.0476644 apparent bandwidth: 2685.44 mb/s device host iteration 4 elapsed time: 0.0500736 apparent bandwidth: 2556.24 mb/s finished $ this looks correct me because bandwidthtest on system report 6gb/s in either direction have pcie gen2 system. since thrust uses pageable, not pinned memory, half bandwidth, i.e. 3gb/s, , thrust reporting 2.5gb/s.
for comparison, here bandwidth test on system, using pageable memory:
$ /usr/local/cuda/samples/bin/linux/release/bandwidthtest --memory=pageable [cuda bandwidth test] - starting... running on... device 0: quadro 5000 quick mode host device bandwidth, 1 device(s) pageable memory transfers transfer size (bytes) bandwidth(mb/s) 33554432 2718.2 device host bandwidth, 1 device(s) pageable memory transfers transfer size (bytes) bandwidth(mb/s) 33554432 2428.2 device device bandwidth, 1 device(s) pageable memory transfers transfer size (bytes) bandwidth(mb/s) 33554432 99219.1 $
Comments
Post a Comment