forked from jiuyuan/InfiniTensor
f60767a770
* add cmake bits about NCCL * move example to examples/NNmodel * impl NCCL communicator * add comm related function to Runtime * export runtime interface * add launch.py * use unique name to distingush the the NCCL ID file * add timeout to communicator init * expose communicator obj from runtime obj, add unit test for nccl communicator * reformat files * Add allReduce operator and cuda nccl allReduce kernel * impl model parallel for resnet * add allGather nccl kernel and operator * Add allreduce allgather operator tests, change allgather kernel to output list of tensor, fix shape infer, handle nullptr output * fix format of onnx.py * use concat following AllGather * get tensor parallel for resnet * fix format of graph_handler.cc * change BUILD_DIST default to OFF * polish code of communicator * update .gitignore * Add broadcast operator and cuda kernel * Add comments for operators * remove const of class member * move communicator to CudaRuntimeObj * Add an empty line at EOF. --------- Co-authored-by: panzezhong <panzezhong@qiyuanlab.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com> |
||
---|---|---|
.. | ||
G2BMM.h | ||
GBMM.h | ||
activation_backward.h | ||
all_gather.h | ||
all_reduce.h | ||
batch_norm.h | ||
broadcast.h | ||
concat.h | ||
conv.h | ||
det.h | ||
dropout.h | ||
element_wise.h | ||
expand.h | ||
extend.h | ||
gather.h | ||
matmul.h | ||
membound.h | ||
pad.h | ||
pooling.h | ||
reduce_mean.h | ||
reshape.h | ||
resize.h | ||
slice.h | ||
softmax.h | ||
split.h | ||
transpose.h | ||
unary.h | ||
where.h |