InfiniTensor/include/cuda/cuda_layernorm.h

#pragma once
#include "operators/unary.h"

namespace infini {
void LaynormKernel(const float *input, const float *scale, const float eps,
                   int size, int scaleSize, const int dimsize, const int stride,
                   float *output, const float *bias, int biasSize);
void LaynormKernel(const float *input, const float *scale, const float eps,
                   int size, int scaleSize, const int dimsize, const int stride,
                   float *output);
void LaynormKernel(const half *input, const half *scale, const half eps,
                   int size, int scaleSize, const int dimsize, const int stride,
                   half *output, const half *bias, int biasSize);
void LaynormKernel(const half *input, const half *scale, const half eps,
                   int size, int scaleSize, const int dimsize, const int stride,
                   half *output);
}; // namespace infini
Add layer normalization (#181) * - add layernorm kernel * success:add layernorm kernel and test * fix: remove unusalble comments * fix: modify code as reviewer suggested * debug,modified .cu and test * optional bias support * overloading function * fix bug after merging; remove time constrain in conv test --------- Co-authored-by: kilinchange <kilinchange@163.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com> 2023-11-24 15:15:14 +08:00			`#pragma once`
			`#include "operators/unary.h"`

			`namespace infini {`
			`void LaynormKernel(const float input, const float scale, const float eps,`
			`int size, int scaleSize, const int dimsize, const int stride,`
			`float output, const float bias, int biasSize);`
			`void LaynormKernel(const float input, const float scale, const float eps,`
			`int size, int scaleSize, const int dimsize, const int stride,`
			`float *output);`
Modify kernel registration & support fp16 (#205) * - Remove dataType from the kernel registration. * - support fp16 for conv * - cpu kernel: adapt the new registration mechanism * modified all register kernel * add where fp16 * add layernorm fp16 * add split_concat fp16 * - element_wise support fp16 * feat: support transpose fp16 * feat: support sliceOp fp16 * - unary support fp16 * - feat: support reduceOp fp16 * feat: support matmulOp/expandOp fp16 * feat: support powOp int8 * add cuda cast & support half-precision for gather * style: fix style * feat:support int8 for gather * style:fix style * modified test_cuda_conv_transposed * fix: fix dist code to support fp16 * fix(graph.cc): fix topo_sort * fix: fix recv and send kernel registration * feat: add field tensors for stub * refactor(frontend): 先排序后构图 Signed-off-by: YdrMaster <ydrml@hotmail.com> * fix: 为中间结果提供tensor到node的mapping * fix (slice): add guard for area out of range * fix: fix matmul fp16 * fix: fix re-dataMalloc for weight tensor and use of naive allocator * feat: add dataType filter for cuda kernel * feat: bang kernel adapt the new registration mechanism * fix: fix some error on mlu * feat: intelcpu kernel adapt the new registration mechanism * feat: modify kernel registration on kunlun * fix intelcpu compiler bug * feat: bang reshape support all dataType * fix: fix bang reduce * fix(all_reduce.cc): fix as reviewer suggessted * fix: fix style and restore unary test codes --------- Signed-off-by: YdrMaster <ydrml@hotmail.com> Co-authored-by: xgqdut2016 <kenan_gewei@163.com> Co-authored-by: xgqdut2016 <140036308+xgqdut2016@users.noreply.github.com> Co-authored-by: zhangyunze <z13785159769@163.com> Co-authored-by: OdinaryWord <sx-hz@163.com> Co-authored-by: YdrMaster <ydrml@hotmail.com> Co-authored-by: panzezhong <panzezhong@qiyuanlab.com> 2024-01-15 11:02:13 +08:00			`void LaynormKernel(const half input, const half scale, const half eps,`
			`int size, int scaleSize, const int dimsize, const int stride,`
			`half output, const half bias, int biasSize);`
			`void LaynormKernel(const half input, const half scale, const half eps,`
			`int size, int scaleSize, const int dimsize, const int stride,`
			`half *output);`
Add layer normalization (#181) * - add layernorm kernel * success:add layernorm kernel and test * fix: remove unusalble comments * fix: modify code as reviewer suggested * debug,modified .cu and test * optional bias support * overloading function * fix bug after merging; remove time constrain in conv test --------- Co-authored-by: kilinchange <kilinchange@163.com> Co-authored-by: Haojie Wang <haojie0429@gmail.com> 2023-11-24 15:15:14 +08:00			`}; // namespace infini`