Go to file

PanZezhong1725 7f6aec6c17 针对bert和gpt2模型分布式推理的优化 (#221 ) * fix(dist): 改善分布式脚本，只打印绝对误差 * feat(dist): 增加可导出onnx的pytorch运行脚本 * feat(front): 增加对Y值为-inf的where算子的图优化 * feat(kernel): 对b为常数的pow和div算子进行特判优化 * fix(front): 消除前端对global output形状信息的依赖，分布式脚本删除不必要的shape infer * feat(kernel): 针对matmul中bias为行向量时的expand操作的特化优化 * fix(kernel): 删除div pow const中不必要的同步 * Update expand.cu * fix: fix comments --------- Co-authored-by: Haojie Wang <haojie0429@gmail.com> Co-authored-by: Derui Yang <ydrml@hotmail.com>		2024-04-01 14:04:28 +08:00
.github/workflows	不依赖 onnx models 的模型存储 (#196 )	2023-12-11 10:44:06 +08:00
3rd-party	test: enhance ci (#62 )	2023-02-12 00:01:36 +08:00
cmake	XCCL support (#171 )	2024-02-29 11:48:35 +08:00
docs	fix broken link in docs (#216 )	2024-02-21 14:03:20 +08:00
examples	针对bert和gpt2模型分布式推理的优化 (#221 )	2024-04-01 14:04:28 +08:00
include	针对bert和gpt2模型分布式推理的优化 (#221 )	2024-04-01 14:04:28 +08:00
proto	Tensor serialization (#25 )	2022-09-13 11:27:41 +08:00
pyinfinitensor	针对bert和gpt2模型分布式推理的优化 (#221 )	2024-04-01 14:04:28 +08:00
python	NNET supports TVM backend and kernels (#78 )	2023-04-18 00:26:36 +08:00
scripts	Xpu (#82 )	2023-10-16 10:57:08 +08:00
src	针对bert和gpt2模型分布式推理的优化 (#221 )	2024-04-01 14:04:28 +08:00
test	XCCL support (#171 )	2024-02-29 11:48:35 +08:00
.clang-format	Add: graph, tensor, and operator	2022-07-31 21:44:03 +08:00
.cmake-format.json	Add: graph, tensor, and operator	2022-07-31 21:44:03 +08:00
.gitignore	impl distributed launch with NCCL (#106 )	2023-09-05 09:47:35 +08:00
.gitmodules	impl distributed launch with NCCL (#106 )	2023-09-05 09:47:35 +08:00
CHANGELOG.md	Update docs (#92 )	2023-07-10 02:31:45 +08:00
CMakeLists.txt	XCCL support (#171 )	2024-02-29 11:48:35 +08:00
LICENSE	Initial commit	2022-07-27 22:40:23 +08:00
Makefile	XCCL support (#171 )	2024-02-29 11:48:35 +08:00
README.md	Update README.md	2024-01-24 22:16:48 +08:00
README_CN.md	Update docs (#92 )	2023-07-10 02:31:45 +08:00
env.sh	Xpu (#82 )	2023-10-16 10:57:08 +08:00

README.md

InfiniTensor

中文项目简介 | Documentation | 中文文档

InfiniTensor is a high-performance inference engine tailored for GPUs and AI accelerators. Its design focuses on effective deployment and swift academic validation.

Get started

Make Commands

make/make build: Builds the project;
make install-python: Builds the project then install the python frontend;
make test-cpp: Builds the project then run cpp unit tests;
make test-onnx: Run python unit tests;

Sets env: TEST=OFF to accelerate compiling.

Sets env: CUDA=ON to enable cuda.

Sets env: BANG=ON to enable bang.

CMake Options

There are several configurable CMake options, see the CMakeLists.txt file.

If USE_BACKTRACE is ON, libdw-dev have to be installed. See the README of backward-cpp for details.
If USE_PROTOBUF is ON, protobuf have to be installed. See the README of protobuf for details.
If USE_CUDA is ON, cuda have to be installed.

Roadmap

RefactorGraph is a newly designed AI framework that is set to replace the current main branch.
EinNet is going to be merged into the main branch.
Integration of PET, a tensor program optimizer supporting partially equivalent transformations.
Supported hardware
- ✔ NVIDIA GPU
- ✔ Cambricon MLU
- ✔ Kunlunxin XPU
- ⬜ Ascend NPU

Contributor Guide

InfiniTensor development is based on the pull request on Github. Before requesting for merging, a PR should satisfy the following requirements

Pass all tests.
1. Now CI on Github will test everything that can be tested in the ci environment, including code format. So, script test/script/clang_format_inplace.sh is for formatting all code.
2. Contributors should run ctest manually and copy its output to the PR. Use fenced code blocks (triple backquotes, i.e., ```) to avoid referencing in Github. Otherwise, # in the output is interpreted as a Github reference. Do not directly paste the ctest output in commit messages either for the same reason.
Receive at least one approval from reviewers.
PR title should be concise since it is going to be the commit message in the main branch after merging and squashing.

Reference

Please cite EinNet or PET in your publications if it helps your research:

@article{zheng2023einnet,
  title={EINNET: Optimizing Tensor Programs with Derivation-Based Transformations},
  author={Zheng, Liyan and Wang, Haojie and Zhai, Jidong and Hu, Muyan and Ma, Zixuan and Wang, Tuowei and Huang, Shuhong and Miao, Xupeng and Tang, Shizhi and Huang, Kezhao and Jia, Zhihao},
  booktitle={17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23)},
  pages={739--755},
  year={2023}
}

@inproceedings{wang2021pet,
  title={PET: Optimizing tensor programs with partially equivalent transformations and automated corrections},
  author={Wang, Haojie and Zhai, Jidong and Gao, Mingyu and Ma, Zixuan and Tang, Shizhi and Zheng, Liyan and Li, Yuanzhi and Rong, Kaiyuan and Chen, Yuanyong and Jia, Zhihao},
  booktitle={15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21)},
  pages={37--54},
  year={2021}
}