YoctoHan

Beijing

Pinned Repositories

aiXcoder-7B
official repository of aiXcoder-7B Code Large Language Model
Language:Python2.3k 20 34359
AixVllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Language:Python0 0 00
TensorRT-LLM
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
Language:C++10k 113 2.3k1.3k
aix_infer_trt
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
Language:C++0 0 00
FasterTransformer
Transformer related optimization, including BERT, GPT
Language:C++0 0 00
lmdeploy
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Language:C++0 0 00

YoctoHan/aix_infer_trt
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
Language:C++0 0 00
YoctoHan/FasterTransformer
Transformer related optimization, including BERT, GPT
Language:C++0 0 00
YoctoHan/lmdeploy
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
Language:C++0 0 00