InfiniTensor

Commit Graph

Author	SHA1	Message	Date
xgqdut2016	a6c919b61d	stream kernel	2024-03-07 09:01:00 +00:00
xgqdut2016	d4721cb40c	modified the memory allocattion	2024-03-06 02:48:45 +00:00
xgqdut2016	6ace4d8ae2	modified test cnnl and bang time	2024-03-05 09:00:49 +00:00
xgqdut2016	9aaf313c6f	modified test_bang_softmax.cc	2024-02-29 07:19:36 +00:00
xgqdut2016	1ed4b36db2	add bangSoftmax , compare cnnl and bang C	2024-02-28 03:14:58 +00:00
xgqdut2016	186a6f37f2	modified the recip	2024-02-27 08:59:01 +00:00
xgqdut2016	920b23cad8	Modify memory allocation method	2024-02-26 02:13:04 +00:00
xgqdut2016	e5d5085e6a	modified tensor.h	2024-02-23 02:19:30 +00:00
xgqdut2016	66d98a3f04	modified kernel, add nDim parameter	2024-02-22 03:11:58 +00:00
xgqdut2016	ddf90fb19e	Merge branch 'master' into bang-softmax	2024-02-22 10:36:39 +08:00
xgqdut2016	c41ad9120d	add bang softmax,output error	2024-02-22 02:35:20 +00:00
baominghelly	b51ccae3b2	fix broken link in docs (#216 ) Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2024-02-21 14:03:20 +08:00
xiaonans	1c08ba200c	[feature] add cudagraph support (#215 ) * [feature] add cudagraph support * modify code to pass the cuda_all_reduce test	2024-02-21 14:00:25 +08:00
xiaonans	900d8e58e3	Rope and silu (#214 ) 添加silu和rotary embedding算子的支持。	2024-02-04 11:05:27 +08:00
xiaonans	b0876a13ce	Merge branch 'master' into rope_and_silu	2024-02-04 10:57:36 +08:00
xiaonans	ae9f61de5a	add comment for rope operator	2024-02-04 10:57:01 +08:00
xiaonans	9a3c0f11f6	add test for rotary embedding cuda kernel	2024-02-04 10:24:20 +08:00
zhangyunze	67b2bcb7d5	fix mlu some kernel registration & gather op (#210 ) * fix: fix bang build/kernel registration \| test_onnx * delete assert float * fix gather * fix CMakeLists and Reshape * fix cncl ops * add hardsigmoid/hardswish * fix * add invalid datatype exception * fix gather * fix gather indices type * fix gather/prelu/hardsigmoid on mlu * fix format * fix --------- Co-authored-by: Bolun Zhang <48948016+Chamberlain0w0@users.noreply.github.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com> Co-authored-by: Zhang Bolun <Chamberlain0w0@gmail.com>	2024-02-01 15:02:02 +08:00
xiaonans	956ce37458	add unittest of silu kernel	2024-01-30 10:40:13 +08:00
zhangyunze	4813204a36	feat: add reshape/identity/squeeze/flatten/unsqueeze op cpu kernel (#213 )	2024-01-30 10:29:59 +08:00
xiaonans	030e5ca9c1	Merge branch 'master' of github.com:InfiniTensor/InfiniTensor into rope_and_silu	2024-01-26 10:16:18 +08:00
xiaonans	e8d111ef5d	add rope and silu support	2024-01-26 10:01:27 +08:00
xiaonans	d1a90ba3e2	[feature] support kvcache with static graph (#209 ) * [feature] support kvcache with static graph * use workspace to optimize kvcache attention --------- Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2024-01-25 14:20:43 +08:00
xiaonans	afed5d3c3d	use workspace to optimize kvcache attention	2024-01-25 10:33:01 +08:00
Haojie Wang	a5062f3f89	Update README.md	2024-01-24 22:16:48 +08:00
Hardy	09b2ecf98a	support more data type on mlu (#211 ) * support more data type * clang format * fix little bug * fix cncl datatype * fix format --------- Co-authored-by: wanghailu <wanghailu0717@163.com> Co-authored-by: Zhang Bolun <Chamberlain0w0@gmail.com>	2024-01-24 13:33:33 +08:00
xiaonans	6a1bfd6c45	[feature] support kvcache with static graph	2024-01-17 11:38:44 +08:00
Chenjie Duan	51086d2b8d	Modify kernel registration & support fp16 (#205 ) * - Remove dataType from the kernel registration. * - support fp16 for conv * - cpu kernel: adapt the new registration mechanism * modified all register kernel * add where fp16 * add layernorm fp16 * add split_concat fp16 * - element_wise support fp16 * feat: support transpose fp16 * feat: support sliceOp fp16 * - unary support fp16 * - feat: support reduceOp fp16 * feat: support matmulOp/expandOp fp16 * feat: support powOp int8 * add cuda cast & support half-precision for gather * style: fix style * feat:support int8 for gather * style:fix style * modified test_cuda_conv_transposed * fix: fix dist code to support fp16 * fix(graph.cc): fix topo_sort * fix: fix recv and send kernel registration * feat: add field tensors for stub * refactor(frontend): 先排序后构图 Signed-off-by: YdrMaster <ydrml@hotmail.com> * fix: 为中间结果提供tensor到node的mapping * fix (slice): add guard for area out of range * fix: fix matmul fp16 * fix: fix re-dataMalloc for weight tensor and use of naive allocator * feat: add dataType filter for cuda kernel * feat: bang kernel adapt the new registration mechanism * fix: fix some error on mlu * feat: intelcpu kernel adapt the new registration mechanism * feat: modify kernel registration on kunlun * fix intelcpu compiler bug * feat: bang reshape support all dataType * fix: fix bang reduce * fix(all_reduce.cc): fix as reviewer suggessted * fix: fix style and restore unary test codes --------- Signed-off-by: YdrMaster <ydrml@hotmail.com> Co-authored-by: xgqdut2016 <kenan_gewei@163.com> Co-authored-by: xgqdut2016 <140036308+xgqdut2016@users.noreply.github.com> Co-authored-by: zhangyunze <z13785159769@163.com> Co-authored-by: OdinaryWord <sx-hz@163.com> Co-authored-by: YdrMaster <ydrml@hotmail.com> Co-authored-by: panzezhong <panzezhong@qiyuanlab.com>	2024-01-15 11:02:13 +08:00
zhangyunze	58993d4339	解除前端对onnx infershape功能的依赖 (#206 ) * feat: SqueezeOp lift the dependency of onnx infershape. * feat: UnsqueezeOp lift the dependency of onnx infershape. * feat: lift the dependency of onnx infershape * fix: fix Makefile off nccl	2024-01-12 14:54:27 +08:00
PanZezhong1725	46e61a5bd4	修正Slice内存越界问题 (#204 ) fix (slice): add guard for area out of range Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2024-01-05 09:19:50 +08:00
zhangyunze	b15c4979fa	fix Issue-189 question 1-15 (#195 ) * fix: fix nativecpu elementwise only support 4d tensor * fix format --------- Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2024-01-05 08:40:18 +08:00
Hardy	42032356fb	Bang cncl (#163 ) * MLU CNCL base * add FindCNCL.cmake, not find -lcncl * bangPrintFloat not find * docker:make sucessful, test error * delete net file and onnxtest.py * init * fix cncl * format * fix * format * fix cncl * run dist gpt2 on mlu * format * fix import error on mlu docker * run llama single card * run distributed llama2 * add test for slice/reduce on mlu * fix cncl related test * fix format * format * delete comments * change GPU to MLU * MLU CNCL base * add FindCNCL.cmake, not find -lcncl * bangPrintFloat not find * docker:make sucessful, test error * delete net file and onnxtest.py * init * fix cncl * format * fix * format * fix cncl * run dist gpt2 on mlu * format * fix import error on mlu docker * run llama single card * run distributed llama2 * add test for slice/reduce on mlu * fix cncl related test * fix format * format * delete comments * change GPU to MLU * modify launch script * fix name * fix format * fix gather * format python script --------- Co-authored-by: xgqdut2016 <kenan_gewei@163.com> Co-authored-by: Bolun <chamberlain0w0@gmail.com> Co-authored-by: Bolun Zhang <48948016+Chamberlain0w0@users.noreply.github.com>	2024-01-03 13:28:03 +08:00
Chenjie Duan	83f1de93d0	add frontend resize kernel (#194 ) * - add frontend resize kernel * - fix resize test * - fix bug - add onnx test for resize * fix: modify codes as reviewer suggested --------- Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-12-29 13:32:56 +08:00
zhangyunze	3967b437c8	fix Issue 187 split infershape wrong (#197 ) * fix: fix splitOp to support unequal portions * fix: fix as review comment --------- Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-12-28 21:39:24 +08:00
Chenjie Duan	6e7bd6ca0c	fix(perf.py): change NNmodel commit to fix perf.py (#203 )	2023-12-28 21:31:39 +08:00
Hardy	5ac0ab442f	Fix bang (#198 ) * fix bang batchnorm * fix pooling test bang * add test batchnorm * HIGH PRECISION ACTIVATION * fix pooling * fix matmul * fix test * add layernorm * fix softmax * fix * better code * fix * fix worlflow * fix workflow * fix * fix * fxi matmul * add LRN * fix lrn * fix lrn --------- Co-authored-by: wanghailu <wanghailu0717@163.com> Co-authored-by: Baoming Li <1508269885@qq.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-12-28 13:44:10 +08:00
Chenjie Duan	3f34372012	- modify error info when kernel not found (#191 ) * - modify error info when kernel not found * - modify code as reviewer suggested --------- Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-12-27 09:43:57 +08:00
learner2468	9a9587556c	Add examples: inference of Paddle models (#192 ) * Add paddle model and infer with InfiniTensor * Remove unused import --------- Co-authored-by: kilinchange <44265800+kilinchange@users.noreply.github.com> 【Hackathon No.106】Add paddle model and infer with InfiniTensor	2023-12-14 19:42:43 +08:00
xgqdut2016	a3929c25f8	Add send and recv operators based on NCCL (#182 ) * baseline sendrecv, bug * success sendrecv * get rank from comm * set output shape * successful:set output shape equal to input shape * shape as attribute * success:shape as attribute * success send recv, output 0 * add onnx test * split send and recv * success split send and recv * test-onnx bug * success test-onnx * modified onnx.py * solve review	2023-12-14 16:38:03 +08:00
Derui Yang	c143eebdf7	不依赖 onnx models 的模型存储 (#196 ) Signed-off-by: YdrMaster <ydrml@hotmail.com>	2023-12-11 10:44:06 +08:00
Hardy	67974aee8a	Fix https://github.com/InfiniTensor/InfiniTensor/pull/160 (#185 ) Co-authored-by: wanghailu <wanghailu0717@163.com>	2023-11-27 14:18:12 +08:00
Hardy	3ead20a23a	Fix workspace & bang conv (#183 ) * fix bang workspace * fix convbpdata * fix code * add code * fix * fix * fix conv * fix test conv --------- Co-authored-by: wanghailu <wanghailu0717@163.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-11-24 15:16:25 +08:00
xgqdut2016	a7293c12ba	Add layer normalization (#181 ) * - add layernorm kernel * success:add layernorm kernel and test * fix: remove unusalble comments * fix: modify code as reviewer suggested * debug,modified .cu and test * optional bias support * overloading function * fix bug after merging; remove time constrain in conv test --------- Co-authored-by: kilinchange <kilinchange@163.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-11-24 15:15:14 +08:00
PanZezhong1725	6ece3f4a77	Add ReduceSum op and kernel (#160 ) * Add reduceSum op and kernel * fix merge and format * Reduce: reuse cat macro, add doc string --------- Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-11-24 09:29:58 +08:00
xgqdut2016	595a9906d2	add infer index function (#175 ) Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-11-24 09:24:25 +08:00
zhangyunze	331f7ab2b8	support Dynamic tensor infer shape and fix memory pool (#176 ) * feat: support dynamic tensor part1 * feat: support dynamic-tensor part2 * feat: support dynamic tensor part 3 * fix: fix some .. * - add kvcache example * feat: support concat to identity kernel * add a simple mempory pool for allocator * fix: rebase to master * fix bug after merging * - remove outdated script * fix: fix as review --------- Co-authored-by: kilinchange <kilinchange@163.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-11-23 13:11:50 +08:00
xiaonans	965df4e294	[feature] add fused attention_kvcache operator support (#179 ) * [feature] add fused attention_kvcache operator support * add test to attention_kvcache op * Add space line at EOF --------- Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-11-14 23:44:22 +08:00
Hardy	f22fa2766e	add reduce_mean and gather on bang (#167 ) * add code * fix reduce_mean * add softmax on BANG * fix gather * fix boradcast on ele kernel when dim size is zero * add where kernel and fix softmax kernel * fix convbpdata bug * fix format --------- Co-authored-by: wanghailu <wanghailu@qiyuanlab.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-11-10 18:02:44 +08:00
Hardy	50862df765	[Kunlun & CUDA & BANG] add depth2space operator (#178 ) * add depth2space operator * fix format * add depth2space on cambricon bang * add depth2space on gpu --------- Co-authored-by: wanghailu <wanghailu0717@163.com> Co-authored-by: wanghailu <wanghailu@qiyuanlab.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-11-10 17:58:26 +08:00
Hardy	1ea450882b	add reduce_mean and gather on kunlun (#169 ) * add reduce_mean and gather * fix format * fix gather * fix * fix xpu, add where operation, fix element-wise operation * fix format --------- Co-authored-by: wanghailu <wanghailu0717@163.com> Co-authored-by: wanghailu <wanghailu@qiyuanlab.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com>	2023-11-10 17:52:09 +08:00

1 2 3 4 5

244 Commits All Branches Search

244 Commits

All Branches