Implement custom kernel for LLaMA rotary embedding #14

WoosukKwon · 2023-03-30T10:27:36Z

This PR implements a custom CUDA kernel for rotary embedding, which is used in LLaMA. The kernel is responsible for the entire process of applying rotary embedding to query and key, and is thus much more efficient than the PyTorch implementation.

Tested models:

LLaMA-7B
LLaMA-13B

Tested GPUs:

A100

zhuohan123

LGTM!

csrc/pos_encoding_kernels.cu

Install NNCF

* remove JambaConfig and use official one from transformers * changes in Jamba modeling file to align with official HF format

enable fused topK_softmax kernel for hip path

0612 kernel of FP8 on A100

Summary: Add benchmarking scripts and utils. Things to note : - All files are stored in `neuralmagic` folder. - neuralmagic/benchmarks/scripts/* : Actual benchmarking scripts that interact with vllm engine. - neuralmagic/benchmarks/configs/* : JSON config files that define what benchmark commands to run. - neuralmagic/benchmarks/run_*.py : Scripts that consume some config file and run the benchmark scripts. - neuralmagic/tools : Add tools Testing: Local testing --------- Co-authored-by: Varun Sundar Rabindranath <[email protected]> Co-authored-by: rsnm2 <[email protected]>

a fix follow up [MRotaryEmbedding change](vllm-project@bf3b79e#diff-6bc44986c91bf0876240dec03d56c748403691c7fcd90f7a22e7affff7b033ecR839) Signed-off-by: z00897138 <[email protected]> Co-authored-by: z00897138 <[email protected]>

ISSUE: The USE_CUTLASS_MOE environment variable support (CLAUDE.md entry vllm-project#14) was lost during a previous merge, removing critical debugging/compatibility control. ROOT CAUSE: Upstream changes overwrote the Mantle modification that added environment variable control for CUTLASS MoE implementations. SOLUTION: Restored the missing environment variable logic: - Added `import os` to imports - Restored `default_use_cutlass` calculation with original conditions - Restored `USE_CUTLASS_MOE` environment variable with smart defaults: * USE_CUTLASS_MOE=1 forces CUTLASS MoE on (default when conditions met) * USE_CUTLASS_MOE=0 disables CUTLASS MoE, fallback to other implementations - Maintains backward compatibility with automatic detection CODE CHANGES: - File: `vllm/model_executor/layers/quantization/compressed_tensors/compressed_tensors_moe.py` - Lines: 5 (import), 547-556 (environment variable logic) - Annotation: Added comprehensive Mantle modification comments for future merge guidance TESTING: Verified import functionality and environment variable integration. This fix enables debugging and compatibility control for CUTLASS MoE implementations as documented in CLAUDE.md registry entry vllm-project#14. Signed-off-by: Pradyun Ramadorai <[email protected]>

setup sparse attention backend

WoosukKwon added 7 commits March 30, 2023 07:01

Minor

512c7bf

Add test code for rotary embedding

56674f4

Minor

3533de0

Minor

8e0e6a4

Add rotary embedding kernel

3b6652a

Add test code for rotary embedding kernel

b29eb16

Implement Llama attention layer

7392665

WoosukKwon requested a review from zhuohan123 March 30, 2023 10:29

Minor fix in comment

eef11ba

WoosukKwon changed the title ~~Add custom kernel for rotary embedding~~ Implement custom kernel for LLaMA rotary embedding Mar 30, 2023

zhuohan123 approved these changes Mar 30, 2023

View reviewed changes

csrc/pos_encoding_kernels.cu Show resolved Hide resolved

Test more head sizes

1a26188

WoosukKwon merged commit 88c0268 into main Mar 30, 2023

WoosukKwon deleted the rotary-embedding branch March 30, 2023 18:04

bigPYJ1151 added a commit to bigPYJ1151/vllm that referenced this pull request Sep 12, 2023

Add multi-attention op. (vllm-project#14)

824dfc9

shanshanpt mentioned this pull request Nov 17, 2023

Run long conetxt error : CUDA error: an illegal memory access was encountered #1700

Closed

junior-zsy mentioned this pull request Nov 20, 2023

Error with 32k Long Text in chatglm2-6b-32k Model #1725

Closed

hongxiayang pushed a commit to hongxiayang/vllm that referenced this pull request Feb 13, 2024

Implement custom kernel for LLaMA rotary embedding (vllm-project#14)

cb12020

luo-cheng2021 pushed a commit to luo-cheng2021/vllm that referenced this pull request Mar 25, 2024

Merge pull request vllm-project#14 from ilya-lavrenov/install-nncf

05b9161

Install NNCF

mzusman pushed a commit to mzusman/vllm that referenced this pull request May 6, 2024

Jamba official hf (vllm-project#14)

988718e

* remove JambaConfig and use official one from transformers * changes in Jamba modeling file to align with official HF format

fxmarty pushed a commit to fxmarty/vllm-public that referenced this pull request May 31, 2024

Merge pull request vllm-project#14 from ROCm/fused_topK_softmax

e3ae076

enable fused topK_softmax kernel for hip path

yuhuixu1993 mentioned this pull request Jun 2, 2024

[Bug]: loading squeezellm model #5190

Closed

ykim362 pushed a commit to ykim362/vllm that referenced this pull request Jun 17, 2024

Merge pull request vllm-project#14 from wenxcs/wenxh/fp8-on-a100-v5-pr

b28848e

0612 kernel of FP8 on A100

alixiaodi mentioned this pull request Aug 2, 2024

[Bug]: #7072

Closed

SpaceHunterInf mentioned this pull request Sep 30, 2024

[Bug]: Bus error (core dumped) #8974

Closed

1 task

hao-cold mentioned this pull request May 13, 2025

[Bug]: CUDA error: an illegal instruction was encountered #18045

Closed

1 task

markmc mentioned this pull request May 21, 2025

[Bug][Failing Test]: Distributed Comm Ops - distributed/test_shm_broadcast.py #18492

Closed

1 task

zerosurplus mentioned this pull request Jun 16, 2025

[Bug]: torch.distributed.DistNetworkError: The client socket has timed out after 600000ms while trying to connect to (172.17.0.9, 46229). #19670

Open

1 task

xiaomofang mentioned this pull request Jul 31, 2025

[Bug]: There is an issue with speculative inference in Eagle mode, where the context length of vLLM inference is constrained by the draft model. #21986

Open

1 task

heheda12345 pushed a commit to heheda12345/vllm that referenced this pull request Sep 29, 2025

Merge pull request vllm-project#14 from vllm-model-0920/mla_backend

446c0de

setup sparse attention backend

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Uh oh!

Implement custom kernel for LLaMA rotary embedding #14

Implement custom kernel for LLaMA rotary embedding #14

Uh oh!

WoosukKwon commented Mar 30, 2023 •

edited

Loading

Uh oh!

zhuohan123 left a comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

Uh oh!

Implement custom kernel for LLaMA rotary embedding #14

Implement custom kernel for LLaMA rotary embedding #14

Uh oh!

Conversation

WoosukKwon commented Mar 30, 2023 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

zhuohan123 left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

WoosukKwon commented Mar 30, 2023 •

edited

Loading