CUDA: Why Thrust is so slow on uploading data to GPU? -


i'm new gpu world , installed cuda writing program. played thrust library find out slow when uploading data gpu. 35mb/s in host-to-device part on not-bad desktop. how come is?

environment: visual studio 2012, cuda 5.0, gtx760, intel-i7, windows 7 x64

gpu bandwidth test: enter image description here

it supposed have @ least 11gb/s of transfer speed host device or vice versa! didn't!

here's test program:

#include <iostream> #include <ctime> #include <thrust/device_vector.h> #include <thrust/host_vector.h>  #define n 32<<22  int main(void) {     using namespace std;      cout<<"gpu bandwidth test via thrust, data size: "<< (sizeof(double)*n) / 1000000000.0 <<" gbytes"<<endl;     cout<<"============program start=========="<<endl;      int = time(0);     cout<<"initializing h_vec...";     thrust::host_vector<double> h_vec(n,0.0f);     cout<<"time spent: "<<time(0)-now<<"secs"<<endl;      = time(0);     cout<<"uploading data gpu...";     thrust::device_vector<double> d_vec = h_vec;     cout<<"time spent: "<<time(0)-now<<"secs"<<endl;      = time(0);     cout<<"downloading data h_vec...";     thrust::copy(d_vec.begin(), d_vec.end(), h_vec.begin());     cout<<"time spent: "<<time(0)-now<<"secs"<<endl<<endl;      system("pause");     return 0; } 

program out put: enter image description here

  • download speed: less 1 sec, pretty make sense compare nominal 11gb/s.

  • upload speed: 1.07374gb /32 secs 33.5 mb/s, doesn't make sense @ all.

does know reason? or way thrust is?

thanks!!

your comparison has several flaws, of covered in comments.

  1. you need eliminate allocation effects. can doing "warm-up" transfers first.
  2. you need eliminate "start-up" effects. can doing "warm-up" transfers first.
  3. when comparing data, remember bandwidthtest using pinned memory allocation, thrust not use. therefore thrust data transfer rate slower. typically contributes 2x factor (i.e. pinned memory transfers typically 2x faster pageable memory transfers. if want better comparison bandwidthtest run --memory=pageable switch.
  4. your choice of timing functions might not best. cudaevents pretty reliable timing cuda operations.

here code proper timing:

$ cat t213.cu #include <iostream> #include <thrust/device_vector.h> #include <thrust/host_vector.h> #include <thrust/copy.h> #include <thrust/fill.h>  #define dsize ((1ul<<20)*32)  int main(){    thrust::device_vector<int> d_data(dsize);   thrust::host_vector<int> h_data(dsize);   float et;   cudaevent_t start, stop;   cudaeventcreate(&start);   cudaeventcreate(&stop);    thrust::fill(h_data.begin(), h_data.end(), 1);   thrust::copy(h_data.begin(), h_data.end(), d_data.begin());    std::cout<< "warm iteration " << d_data[0] << std::endl;   thrust::fill(d_data.begin(), d_data.end(), 2);   thrust::copy(d_data.begin(), d_data.end(), h_data.begin());   std::cout<< "warm iteration " << h_data[0] << std::endl;   thrust::fill(h_data.begin(), h_data.end(), 3);   cudaeventrecord(start);   thrust::copy(h_data.begin(), h_data.end(), d_data.begin());   cudaeventrecord(stop);   cudaeventsynchronize(stop);   cudaeventelapsedtime(&et, start, stop);   std::cout<<"host device iteration " << d_data[0] << " elapsed time: " << (et/(float)1000) << std::endl;   std::cout<<"apparent bandwidth: " << (((dsize*sizeof(int))/(et/(float)1000))/((float)1048576)) << " mb/s" << std::endl;   thrust::fill(d_data.begin(), d_data.end(), 4);   cudaeventrecord(start);   thrust::copy(d_data.begin(), d_data.end(), h_data.begin());   cudaeventrecord(stop);   cudaeventsynchronize(stop);   cudaeventelapsedtime(&et, start, stop);   std::cout<<"device host iteration " << h_data[0] << " elapsed time: " << (et/(float)1000) << std::endl;   std::cout<<"apparent bandwidth: " << (((dsize*sizeof(int))/(et/(float)1000))/((float)1048576)) << " mb/s" << std::endl;    std::cout << "finished" << std::endl;   return 0; } 

i compile (i have pcie gen2 system cc2.0 device)

$ nvcc -o3 -arch=sm_20 -o t213 t213.cu 

when run following results:

$ ./t213 warm iteration 1 warm iteration 2 host device iteration 3 elapsed time: 0.0476644 apparent bandwidth: 2685.44 mb/s device host iteration 4 elapsed time: 0.0500736 apparent bandwidth: 2556.24 mb/s finished $ 

this looks correct me because bandwidthtest on system report 6gb/s in either direction have pcie gen2 system. since thrust uses pageable, not pinned memory, half bandwidth, i.e. 3gb/s, , thrust reporting 2.5gb/s.

for comparison, here bandwidth test on system, using pageable memory:

$ /usr/local/cuda/samples/bin/linux/release/bandwidthtest --memory=pageable [cuda bandwidth test] - starting... running on...   device 0: quadro 5000  quick mode   host device bandwidth, 1 device(s)  pageable memory transfers    transfer size (bytes)        bandwidth(mb/s)    33554432                     2718.2   device host bandwidth, 1 device(s)  pageable memory transfers    transfer size (bytes)        bandwidth(mb/s)    33554432                     2428.2   device device bandwidth, 1 device(s)  pageable memory transfers    transfer size (bytes)        bandwidth(mb/s)    33554432                     99219.1  $ 

Comments

Popular posts from this blog

basic authentication with http post params android -

vb.net - Virtual Keyboard commands -

android - Inheriting from Theme.AppCompat* -