InfiniTensor

History

constroy Li 4c321c8a91 tensor parallel for transformer (#125 ) * add cmake bits about NCCL * move example to examples/NNmodel * impl NCCL communicator * add comm related function to Runtime * export runtime interface * add launch.py * use unique name to distingush the the NCCL ID file * add timeout to communicator init * expose communicator obj from runtime obj, add unit test for nccl communicator * reformat files * Add allReduce operator and cuda nccl allReduce kernel * impl model parallel for resnet * add allGather nccl kernel and operator * Add allreduce allgather operator tests, change allgather kernel to output list of tensor, fix shape infer, handle nullptr output * fix format of onnx.py * use concat following AllGather * get tensor parallel for resnet * fix format of graph_handler.cc * change BUILD_DIST default to OFF * polish code of communicator * update .gitignore * export min/max to python * fix MatMul * modify launch.py to run opt * hack to treat ReduceSum as AllReduceSum * throw exception in cuda error * fix parallel_opt.py * improve the error prompt and cuda error check * fix GatherObj::GatherObj member init * fix size calculation for scalar (rank = 0) tensor * MatMul supports bias * fix add bias for row parallel gemm * add --gen_std to launch.py * fix AllReduceNCCL * update launch.py * less log * update parallel_opt * update launch.py * add __eq__ for Placement sub-classes * less benchmark run * fix placement infer for matmul * fix vacabuary size * fix Exception * Add shard tensor with group to support gpt2 * Add find successor function to find split op at different depth * recover CommunicatorObj * improve error mesasge * optimize parallel_opt.py * optimize launch.py * recover docs for all_reduce and all_gather * Fix API * fix format --------- Co-authored-by: panzezhong <panzezhong@qiyuanlab.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com>		2023-09-14 14:19:45 +08:00
..
broadcast_shape.h	"modified where" (#131 )	2023-09-14 10:45:57 +08:00
data_convert.h	support mixed dtype (#102 )	2023-08-16 21:49:43 +08:00
data_generator.h	支持fp16 dtype (#96 )	2023-08-02 16:38:16 +08:00
dataloader.h	Json perfrecord (#32 )	2022-09-22 15:34:34 +08:00
exception.h	tensor parallel for transformer (#125 )	2023-09-14 14:19:45 +08:00
operator_utils.h	support mixed dtype (#102 )	2023-08-16 21:49:43 +08:00
small_array.h	tensor parallel for transformer (#125 )	2023-09-14 14:19:45 +08:00
validation.h	Verify wanghailu 0902 (#22 )	2022-09-05 15:45:52 +08:00