forked from jiuyuan/InfiniTensor
![]() * "add softmax.cu,.cc,.h" * Modify cuda softmax * "modified the introduction of softmax.cu" * "add format of cuda_softmax.h" * "modified where.cc(.cu,.h) and softmax.cu" * "modified format" * Fix cpu softmax kernel * "modified the // introduction of softmax.cu" * "modified softmax.cu and use 1D block" * "modified softmax.cu,format, and use 1D block" * "introduce share mem to speed softmax" * "reduce the input of function" * modified the format * remodify 2D block softmax * remodify 1D block softmax * modified the share memory * add warp reduce * conflict solve two * remove extra space line * solve comment --------- Co-authored-by: Haojie Wang <haojie0429@gmail.com> Co-authored-by: panzezhong <panzezhong@qiyuanlab.com> |
||
---|---|---|
.. | ||
cuda_clip.h | ||
cuda_common.h | ||
cuda_element_wise.h | ||
cuda_expand.h | ||
cuda_kernel_wihtout_config.h | ||
cuda_pad_slice.h | ||
cuda_runtime.h | ||
cuda_softmax.h | ||
cuda_split_concat.h | ||
cuda_transpose.h | ||
cuda_unary.h | ||
cuda_utility.h | ||
cuda_where.h | ||
gather.h | ||
gbmm_g2bmm.cuh | ||
gbmm_g2bmm.h | ||
nccl_communicator.h | ||
operator_timer.h | ||
resize.cuh |